Tool in benchmark · Version 1

Chatbase in AI Customer Support Chatbots

Scenario-level performance from current published Results.

No published Results are available for Chatbase in this benchmark. · 26 scenarios in the benchmark

Publication availability does not describe testing status.

How Chatbase performed

Open a capability to explore its scenarios. Each row reports the test set in its published Result; counts are not combined into an overall score.

Action execution5 scenarios · 0 with published Results
ScenarioPublished outcomesTest coverageResult
Action requires confirmationNo published result
Requested action cannot be completedNo published result
Retrieve information from an external systemNo published result
Update something in an external systemNo published result
User is not authorized to perform the actionNo published result
Agent configuration and control2 scenarios · 0 with published Results
ScenarioPublished outcomesTest coverageResult
Configured instruction changes agent behaviorNo published result
Configured restriction is respectedNo published result
Analytics and observability3 scenarios · 0 with published Results
ScenarioPublished outcomesTest coverageResult
Agent actions/tool calls can be inspectedNo published result
Conversation activity can be inspectedNo published result
Reported analytics match what actually happenedNo published result
Conversation continuity3 scenarios · 0 with published Results
ScenarioPublished outcomesTest coverageResult
Customer corrects information given earlierNo published result
Earlier conversation/session needs to be continuedNo published result
Information from earlier in the conversation is needed laterNo published result
Escalation and human handoff3 scenarios · 0 with published Results
ScenarioPublished outcomesTest coverageResult
A configured rule requires escalationNo published result
Agent cannot resolve the requestNo published result
User explicitly asks for a humanNo published result
Feedback and learning2 scenarios · 0 with published Results
ScenarioPublished outcomesTest coverageResult
Correction changes future behavior where the product claims learningNo published result
Feedback can be submittedNo published result
Intent understanding3 scenarios · 0 with published Results
ScenarioPublished outcomesTest coverageResult
One message contains multiple requestsNo published result
Request is unclear and needs clarificationNo published result
Request requires choosing the correct source or actionNo published result
Knowledge-grounded answering3 scenarios · 0 with published Results
ScenarioPublished outcomesTest coverageResult
Answer is available in the knowledge baseNo published result
Answer is not available in the knowledge baseNo published result
The request uses different wording than the sourceNo published result
Multilingual support2 scenarios · 0 with published Results
ScenarioPublished outcomesTest coverageResult
Customer communicates in another supported languageNo published result
Customer switches languages during the conversationNo published result

Reading these Results

Published evidence and test coverage answer different questions.

Publication availability

Which scenarios have a Result?

A published Result is public evidence for this tool on one scenario. “No published result” does not say whether testing has taken place.

Test coverage

What does each Result cover?

Assessed includes Pass, Fail and Not gradable. Gradable includes Pass and Fail. Both use the pinned test count in that published Result.

Scenario scope

Inventory is not testing progress

The 26 scenarios describe this benchmark’s scope. They are not an assumed applicability or test-coverage denominator for Chatbase.

Chatbase in AI Customer Support Chatbots | AI Demos