The bot responds in Hindi and offers either a human-agent connection or continued help, showing it can understand a non-English support request.
What was measured
Multilingual Understanding
Correctly understands support requests written in more than one language.
decisive for this rankingtransformation
If the bot cannot understand customer requests in the languages it is expected to serve, it cannot automate support well. (3 of 3 judges)
What was given, what came back
Test input: Customer-support handoff request · text
Input — what we sent
The exact prompt
I received the wrong item in my order. I've already checked the order details and this is clearly a mistake on your end. I don't want any more back and forth — can you please connect me to a customer support agent or raise a ticket for this?
A frustrated support-escalation request after receiving the wrong item, asking to be connected to a support agent or have a ticket raised.
Why this input is hard
- · Human handoff reliability
- · Ticket creation workflow
- · Escalation contextual awareness
- · Professional tone under frustration
Output — unretouched

Also checked on this input — same tool, 3 other criteria
Conversational Quality◐ MixedThe bot acknowledges the request to speak to a real person and offers a human-agent connection, but it also tries to keep the conversation with the bot first.Conversational Quality✓ WorkedThe bot validates the preference for human support and offers to connect the user with an agent.Conversational Safety Behavior◐ MixedThe bot remains calm about the fraud claim and offers a way to verify whether the double charge is a temporary hold, but it does not immediately escalate the fraud concern.
Provenance
- Observation
- ed7e36a6-1531-4f98-9f1f-8a4b3dce0040
- Evidence run
- 2645dc92-49df-478a-b809-21dfd09f06a7
- Study
- Automate customer support using an AI chatbot
- Research task
- 86b9jm3ev
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "fin"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 4 other tools
measured on Multilingual Understanding
Freshdesk✓ WorkedThe bot recognizes a Spanish request for customer service, states that it is trained for English, French, and Urdu, and still assigns the chat to support.FS Agent✓ WorkedUnderstands and answers a Spanish escalation request in Spanish, requesting the user's email and a brief problem description to create a support ticket.JotForm✓ WorkedIt correctly understood 2 escalation requests written in Hindi and Spanish, responded in the same language, and asked for an email address for follow-up.Wonderchat✓ WorkedIt understands the Hindi support request and answers fully in Hindi, preserving the support-channel options and response-time details.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com