It fulfills a direct request for human help by listing multiple escalation channels — email, phone, and live chat — with response-time guidance instead of forcing self-service.
What was measured
Response Completeness
Addresses all parts of the user’s support request instead of leaving key questions unanswered.
decisive for this rankingtransformation
A support bot that leaves key questions unanswered is not solving the customer’s issue. (3 of 3 judges)
What was given, what came back
Test input: Customer-support handoff request · text
Input — what we sent
The exact prompt
I received the wrong item in my order. I've already checked the order details and this is clearly a mistake on your end. I don't want any more back and forth — can you please connect me to a customer support agent or raise a ticket for this?
A frustrated support-escalation request after receiving the wrong item, asking to be connected to a support agent or have a ticket raised.
Why this input is hard
- · Human handoff reliability
- · Ticket creation workflow
- · Escalation contextual awareness
- · Professional tone under frustration
Output — unretouched

Also checked on this input — same tool, 3 other criteria
Multilingual Understanding✓ WorkedIt understands the Hindi support request and answers fully in Hindi, preserving the support-channel options and response-time details.Multilingual Understanding✓ WorkedIt understands the Spanish support request and replies fully in Spanish with the same multi-channel escalation options and timing information.Policy Accuracy✓ WorkedIt follows the billing workflow for a suspected double charge by asking whether the entries are pending or posted, requesting transaction references, and reserving escalation for a genuine duplicate posted charge.
Provenance
- Observation
- f705712c-5bcc-4515-9831-fcde7696aefe
- Evidence run
- 2645dc92-49df-478a-b809-21dfd09f06a7
- Study
- Automate customer support using an AI chatbot
- Research task
- 86b9jm3ev
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "wonderchat"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 5 other tools
measured on Response Completeness
Chatbase✓ WorkedFor a user who does not want to talk to a bot, it asks only for the email address and a brief issue description to create the support request.Freshdesk◐ MixedOn a double-charge fraud claim, the bot does not escalate immediately; it first frames the issue as a possible authorization hold and asks for transaction details, so the handoff is delayed rather than direct.FS Agent✓ WorkedCompletes a 'no bot' escalation by routing the user to support and requesting the email address and brief issue description needed to proceed.Respond✓ WorkedWhen the user said they did not want to talk to a bot, the bot handed the conversation to support and confirmed the connection.Zendesk✓ WorkedThe bot escalates a fraud-like duplicate-charge complaint immediately and does not attempt to troubleshoot the payment itself.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com