AI Customer Support Chatbots
This benchmark evaluates AI support agents for businesses that need customer questions answered and customer requests handled in connected systems.
Benchmark overview
What is included and excluded
Evaluates AI support agents on grounded answers, intent understanding, conversation continuity, human handoff, connected-system actions, multilingual support, configuration and control, analytics and observability, and feedback-driven learning.
A reader learns how a product handles the core support capabilities buyers care about: grounded answers, intent understanding, continuity across turns and sessions, human handoff, connected actions, multilingual use, configuration, observability, and feedback-driven learning.
In scope
- AI support agents for customer support.
- Grounded answers from the business’s own knowledge base.
- Intent understanding before answering or acting.
- Conversation continuity across turns and sessions.
- Escalation and human handoff.
- Actions in connected systems.
- Multilingual support.
- Agent configuration and control.
- Analytics and observability.
- Feedback-driven learning.
Out of scope
- General RAG chatbots with no support workflow.
- Live chat tools with no AI.
- Voice support agents.
- Prompt-injection resistance.
Participating tools
Tools in this benchmark’s public roster. Publication availability is not a performance ranking.
| Tool | Published Results | Explore |
|---|---|---|
| Botpress | 8 scenarios with published Results | View tool in this benchmark → |
| Freshdesk Freddy | 7 scenarios with published Results | View tool in this benchmark → |
| Kommunicate | 1 scenario with a published Result | View tool in this benchmark → |
| Zendesk AI | 6 scenarios with published Results | View tool in this benchmark → |
| Ada | No published result | View tool in this benchmark → |
| Chatbase | No published result | View tool in this benchmark → |
| CustomGPT.ai | No published result | View tool in this benchmark → |
| Decagon | No published result | View tool in this benchmark → |
| FS Agent (DIY control) | No published result | View tool in this benchmark → |
| Gorgias | No published result | View tool in this benchmark → |
| Help Scout | No published result | View tool in this benchmark → |
| Intercom Fin | No published result | View tool in this benchmark → |
| Tidio Lyro | No published result | View tool in this benchmark → |
Capabilities & scenarios
26 scenarios grouped by 9 capabilities. Open a group to explore its scenarios in this benchmark.
Knowledge-grounded answering3 scenarios
The agent answers customer questions from the business’s own knowledge base rather than guessing from model assumptions.
Capability in this benchmark → · Global definition →
- Answer is available in the knowledge baseNo published results
- Answer is not available in the knowledge baseNo published results
- The request uses different wording than the sourceNo published results
Intent understanding3 scenarios
The agent works out what the customer actually wants before it answers or acts.
Capability in this benchmark → · Global definition →
- Request requires choosing the correct source or action1 tool with a published Result
- One message contains multiple requests3 tools with published Results
- Request is unclear and needs clarificationNo published results
Conversation continuity3 scenarios
The agent keeps track of earlier context across turns and sessions when later replies depend on it.
Capability in this benchmark → · Global definition →
- Information from earlier in the conversation is needed later1 tool with a published Result
- Customer corrects information given earlier2 tools with published Results
- Earlier conversation/session needs to be continued1 tool with a published Result
Escalation and human handoff3 scenarios
The agent knows when a conversation should go to a human and hands it over properly.
Capability in this benchmark → · Global definition →
- User explicitly asks for a human2 tools with published Results
- A configured rule requires escalationNo published results
- Agent cannot resolve the request1 tool with a published Result
Action execution5 scenarios
The agent can look things up and make changes in connected systems such as orders, subscriptions, or tickets.
Capability in this benchmark → · Global definition →
- Retrieve information from an external system1 tool with a published Result
- Update something in an external system1 tool with a published Result
- Requested action cannot be completed1 tool with a published Result
- Action requires confirmation3 tools with published Results
- User is not authorized to perform the action1 tool with a published Result
Multilingual support2 scenarios
The agent handles customers who do not write in English and can follow language switches during a conversation.
Capability in this benchmark → · Global definition →
- Customer communicates in another supported languageNo published results
- Customer switches languages during the conversationNo published results
Agent configuration and control2 scenarios
The agent follows configured instructions and respects configured restrictions.
Capability in this benchmark → · Global definition →
- Configured instruction changes agent behavior1 tool with a published Result
- Configured restriction is respected1 tool with a published Result
Analytics and observability3 scenarios
The operator can inspect conversations and agent actions, and the reported numbers match what actually happened.
Capability in this benchmark → · Global definition →
- Conversation activity can be inspected1 tool with a published Result
- Agent actions/tool calls can be inspectedNo published results
- Reported analytics match what actually happenedNo published results
Feedback and learning2 scenarios
The agent accepts corrections and, where the product claims learning, changes future answers.
Capability in this benchmark → · Global definition →
- Feedback can be submitted1 tool with a published Result
- Correction changes future behavior where the product claims learningNo published results
Results overview
Current published evidence in this benchmark.
Publication availability is separate from test coverage and unpublished research progress.
How the benchmark works
A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.
Scenarios describe conditions, not inputs
A scenario is a condition that makes the capability hard, not an input archetype or a smaller capability.
Simplest direct verification
The first test case under a scenario is the simplest, most natural test that directly verifies it.
Add harder variants from execution
Harder variants are added later from observed product differences; depth is earned from execution, not designed upfront.
Resources and fixtures
The registered material and systems that create a consistent test environment for this benchmark.
Cedarline commerce database
A shared PostgreSQL 16 commerce database used for benchmark evaluations.
Cedarline knowledge base
Versioned customer-support help-centre corpus used as the benchmark KB fixture.
Cedarline order system API
A customer-support order-system API for looking up and changing order and account state.