Benchmark · Version 1

AI Customer Support Chatbots

This benchmark evaluates AI support agents for businesses that need customer questions answered and customer requests handled in connected systems.

Benchmark overview

9Capabilities
26Scenarios
13Participating tools
What is included and excluded

Evaluates AI support agents on grounded answers, intent understanding, conversation continuity, human handoff, connected-system actions, multilingual support, configuration and control, analytics and observability, and feedback-driven learning.

A reader learns how a product handles the core support capabilities buyers care about: grounded answers, intent understanding, continuity across turns and sessions, human handoff, connected actions, multilingual use, configuration, observability, and feedback-driven learning.

In scope

  • AI support agents for customer support.
  • Grounded answers from the business’s own knowledge base.
  • Intent understanding before answering or acting.
  • Conversation continuity across turns and sessions.
  • Escalation and human handoff.
  • Actions in connected systems.
  • Multilingual support.
  • Agent configuration and control.
  • Analytics and observability.
  • Feedback-driven learning.

Out of scope

  • General RAG chatbots with no support workflow.
  • Live chat tools with no AI.
  • Voice support agents.
  • Prompt-injection resistance.

Participating tools

Tools in this benchmark’s public roster. Publication availability is not a performance ranking.

ToolPublished ResultsExplore
Botpress8 scenarios with published ResultsView tool in this benchmark →
Freshdesk Freddy7 scenarios with published ResultsView tool in this benchmark →
Kommunicate1 scenario with a published ResultView tool in this benchmark →
Zendesk AI6 scenarios with published ResultsView tool in this benchmark →
AdaNo published resultView tool in this benchmark →
ChatbaseNo published resultView tool in this benchmark →
CustomGPT.aiNo published resultView tool in this benchmark →
DecagonNo published resultView tool in this benchmark →
FS Agent (DIY control)No published resultView tool in this benchmark →
GorgiasNo published resultView tool in this benchmark →
Help ScoutNo published resultView tool in this benchmark →
Intercom FinNo published resultView tool in this benchmark →
Tidio LyroNo published resultView tool in this benchmark →

Capabilities & scenarios

26 scenarios grouped by 9 capabilities. Open a group to explore its scenarios in this benchmark.

Knowledge-grounded answering3 scenarios

The agent answers customer questions from the business’s own knowledge base rather than guessing from model assumptions.

Capability in this benchmark → · Global definition →

  1. Answer is available in the knowledge baseNo published results
  2. Answer is not available in the knowledge baseNo published results
  3. The request uses different wording than the sourceNo published results
Intent understanding3 scenarios

The agent works out what the customer actually wants before it answers or acts.

Capability in this benchmark → · Global definition →

  1. Request requires choosing the correct source or action1 tool with a published Result
  2. One message contains multiple requests3 tools with published Results
  3. Request is unclear and needs clarificationNo published results
Conversation continuity3 scenarios

The agent keeps track of earlier context across turns and sessions when later replies depend on it.

Capability in this benchmark → · Global definition →

  1. Information from earlier in the conversation is needed later1 tool with a published Result
  2. Customer corrects information given earlier2 tools with published Results
  3. Earlier conversation/session needs to be continued1 tool with a published Result
Escalation and human handoff3 scenarios

The agent knows when a conversation should go to a human and hands it over properly.

Capability in this benchmark → · Global definition →

  1. User explicitly asks for a human2 tools with published Results
  2. A configured rule requires escalationNo published results
  3. Agent cannot resolve the request1 tool with a published Result
Action execution5 scenarios

The agent can look things up and make changes in connected systems such as orders, subscriptions, or tickets.

Capability in this benchmark → · Global definition →

  1. Retrieve information from an external system1 tool with a published Result
  2. Update something in an external system1 tool with a published Result
  3. Requested action cannot be completed1 tool with a published Result
  4. Action requires confirmation3 tools with published Results
  5. User is not authorized to perform the action1 tool with a published Result
Multilingual support2 scenarios

The agent handles customers who do not write in English and can follow language switches during a conversation.

Capability in this benchmark → · Global definition →

  1. Customer communicates in another supported languageNo published results
  2. Customer switches languages during the conversationNo published results
Agent configuration and control2 scenarios

The agent follows configured instructions and respects configured restrictions.

Capability in this benchmark → · Global definition →

  1. Configured instruction changes agent behavior1 tool with a published Result
  2. Configured restriction is respected1 tool with a published Result
Analytics and observability3 scenarios

The operator can inspect conversations and agent actions, and the reported numbers match what actually happened.

Capability in this benchmark → · Global definition →

  1. Conversation activity can be inspected1 tool with a published Result
  2. Agent actions/tool calls can be inspectedNo published results
  3. Reported analytics match what actually happenedNo published results
Feedback and learning2 scenarios

The agent accepts corrections and, where the product claims learning, changes future answers.

Capability in this benchmark → · Global definition →

  1. Feedback can be submitted1 tool with a published Result
  2. Correction changes future behavior where the product claims learningNo published results

Results overview

Current published evidence in this benchmark.

22Current published Results
16Scenarios with published Results
4Tools with published Results

Publication availability is separate from test coverage and unpublished research progress.

Explore scenarios →

How the benchmark works

A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.

Scenario meaning

Scenarios describe conditions, not inputs

A scenario is a condition that makes the capability hard, not an input archetype or a smaller capability.

Initial test case

Simplest direct verification

The first test case under a scenario is the simplest, most natural test that directly verifies it.

Later depth

Add harder variants from execution

Harder variants are added later from observed product differences; depth is earned from execution, not designed upfront.

Full benchmark methodology →

Resources and fixtures

The registered material and systems that create a consistent test environment for this benchmark.

Database

Cedarline commerce database

A shared PostgreSQL 16 commerce database used for benchmark evaluations.

AI Customer Support Chatbots — Benchmark definition | AI Demos