It follows the billing workflow for a suspected double charge by asking whether the entries are pending or posted, requesting transaction references, and reserving escalation for a genuine duplicate posted charge.

✓ Worked🧾 artifact-verifiedinput + output shownTest date not recordedWonderchat
What was measured
Policy Accuracy

Gives policy details that match the provided knowledge base and does not misstate rules.

decisive for this rankingtransformation

Giving the wrong policy details is a direct failure for a customer-support chatbot. (3 of 3 judges)

What was given, what came back

Test input: Customer-support handoff request · text
Input — what we sent
The exact prompt
I received the wrong item in my order. I've already checked the order details and this is clearly a mistake on your end. I don't want any more back and forth — can you please connect me to a customer support agent or raise a ticket for this?

A frustrated support-escalation request after receiving the wrong item, asking to be connected to a support agent or have a ticket raised.

Why this input is hard
  • · Human handoff reliability
  • · Ticket creation workflow
  • · Escalation contextual awareness
  • · Professional tone under frustration
Output — unretouched
image
Provenance
Observation
e51bc81e-5912-4efa-946a-8c4c52436184
Evidence run
2645dc92-49df-478a-b809-21dfd09f06a7
Study
Automate customer support using an AI chatbot
Research task
86b9jm3ev
Tested at
not recorded
Source
first-party
Evidence state
verified
Proof shown
input + output shown
Cost / latency
not captured
Repeat run
not captured
Tester
not captured

The last three rows are honest blanks, not placeholders — our capture has no field for them yet.

Query this
get_evidence({
  tool: "wonderchat"
})
MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 0 other tools
measured on Policy Accuracy

No other tool was measured on this criterion for this input.

Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com