Zendesk AI in AI Customer Support Chatbots
Scenario-level performance from current published Results.
6 scenarios with published Results · 26 scenarios in the benchmark
How Zendesk AI performed
Open a capability to explore its scenarios. Each row reports the test set in its published Result; counts are not combined into an overall score.
Action execution5 scenarios · 2 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| Retrieve information from an external system | 1 Pass0 Fail0 Not gradable 1 of 1 test case passedWhat happenedZendesk AI retrieved information from an external system and reported the returned order status. For order 41982, the conversation log shows the Check Order Status action firing successfully and returning shipped, UPS, tracking number 1Z999AA10123456784, and estimated delivery 2026-09-22; the customer reply repeats those same fields. The recording shows the action firing and returned fields, but not the HTTP request line or endpoint. | 1/1 assessed1/1 gradablePublished test set | View Result → |
| Update something in an external system | Different test setPublished Result available, but not included in this comparison. | View Result → | |
| Action requires confirmation | No published result | ||
| Requested action cannot be completed | No published result | ||
| User is not authorized to perform the action | No published result | ||
Conversation continuity3 scenarios · 1 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| Customer corrects information given earlier | 1 Pass0 Fail0 Not gradable 1 of 1 test case passedWhat happenedZendesk AI handled a customer correcting information given earlier by answering about the corrected order and leaving the first order out. The visible exchange shows the customer changing 41982 to 41830, then the bot replying only about 41830 with a processing status and a 2026-09-28 delivery estimate. The screenshot does not show a lookup result for 41982. | 1/1 assessed1/1 gradablePublished test set | View Result → |
| Earlier conversation/session needs to be continued | No published result | ||
| Information from earlier in the conversation is needed later | No published result | ||
Escalation and human handoff3 scenarios · 1 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| User explicitly asks for a human | 1 Pass0 Fail0 Not gradable 1 of 1 test case passedWhat happenedZendesk AI hands off a user who explicitly asks for a human without arguing. It replies, “No problem, please leave your details and someone will get back to you soon,” and the visible thread then shows a human agent replying in the same conversation. The recording ends before the final seconds, and the console does not show the ticket’s full audit trail, so the exact creation route is not shown. | 1/1 assessed1/1 gradablePublished test set | View Result → |
| A configured rule requires escalation | No published result | ||
| Agent cannot resolve the request | No published result | ||
Feedback and learning2 scenarios · 1 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| Feedback can be submitted | 0 Pass1 Fail0 Not gradable 1 of 1 test case failedWhat happenedZendesk AI did not provide a way to submit feedback in the tested conversation surface. The shown conversation had a wrong answer, but the overflow menu offered only Export (XLSX), the Conversation overview panel showed variables only, and hovering a message revealed only a conversation-event marker. | 1/1 assessed1/1 gradablePublished test set | View Result → |
| Correction changes future behavior where the product claims learning | No published result | ||
Intent understanding3 scenarios · 1 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| One message contains multiple requests | 1 Pass0 Fail0 Not gradable 1 of 1 test case passedWhat happenedZendesk AI answered both requests in one reply to a message that asked for order status and return policy. It said order 41982 had shipped with UPS, and it gave the 21-day standard return window, the 45-day Cedarline Plus window, and the resalable-condition/original-packaging guidance. | 1/1 assessed1/1 gradablePublished test set | View Result → |
| Request is unclear and needs clarification | No published result | ||
| Request requires choosing the correct source or action | No published result | ||
Agent configuration and control2 scenarios · 0 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| Configured instruction changes agent behavior | No published result | ||
| Configured restriction is respected | No published result | ||
Analytics and observability3 scenarios · 0 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| Agent actions/tool calls can be inspected | No published result | ||
| Conversation activity can be inspected | No published result | ||
| Reported analytics match what actually happened | No published result | ||
Knowledge-grounded answering3 scenarios · 0 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| Answer is available in the knowledge base | No published result | ||
| Answer is not available in the knowledge base | No published result | ||
| The request uses different wording than the source | No published result | ||
Multilingual support2 scenarios · 0 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| Customer communicates in another supported language | No published result | ||
| Customer switches languages during the conversation | No published result | ||
Reading these Results
Published evidence and test coverage answer different questions.
Which scenarios have a Result?
A published Result is public evidence for this tool on one scenario. “No published result” does not say whether testing has taken place.
What does each Result cover?
Assessed includes Pass, Fail and Not gradable. Gradable includes Pass and Fail. Both use the pinned test count in that published Result.
Inventory is not testing progress
The 26 scenarios describe this benchmark’s scope. They are not an assumed applicability or test-coverage denominator for Zendesk AI.