Botpress in AI Customer Support Chatbots
Scenario-level performance from current published Results.
8 scenarios with published Results · 26 scenarios in the benchmark
How Botpress performed
Open a capability to explore its scenarios. Each row reports the test set in its published Result; counts are not combined into an overall score.
Action execution5 scenarios · 2 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| Action requires confirmation | 1 Pass0 Fail0 Not gradable 1 of 1 test case passedWhat happenedBotpress passed this scenario. It asked for confirmation before cancelling, repeated the request after a first yes, and completed the cancellation only after a second confirmation. The recording then shows the subscription state changing to cancelling in the API response. | 1/1 assessed1/1 gradablePublished test set | View Result → |
| User is not authorized to perform the action | 1 Pass0 Fail0 Not gradable 1 of 1 test case passedWhat happenedBotpress refused an action the user was not authorized to perform. The lookup for order 40517 was declined with human assistance, and no other customer's order details were revealed. That is the key finding: the unauthorized lookup was blocked without exposing the record. | 1/1 assessed1/1 gradablePublished test set | View Result → |
| Requested action cannot be completed | No published result | ||
| Retrieve information from an external system | No published result | ||
| Update something in an external system | No published result | ||
Conversation continuity3 scenarios · 3 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| Customer corrects information given earlier | 1 Pass0 Fail0 Not gradable 1 of 1 test case passedWhat happenedBotpress handled a customer correcting information given earlier in the visible chat. The agent first answered about one order, then switched to the corrected order and said it was processing with no tracking number. The first order's shipped status and tracking number did not carry over. | 1/1 assessed1/1 gradablePublished test set | View Result → |
| Earlier conversation/session needs to be continued | 0 Pass1 Fail0 Not gradable 1 of 1 test case failedWhat happenedBotpress failed to continue an earlier conversation when the customer came back later. The recording shows one return conversation ending with a summary and human handoff, then the widget reopening empty. In the later exchange, Botpress asked again for the order number and item instead of picking up where the earlier conversation left off. | 1/1 assessed1/1 gradablePublished test set | View Result → |
| Information from earlier in the conversation is needed later | 1 Pass0 Fail0 Not gradable 1 of 1 test case passedWhat happenedBotpress used earlier conversation information in the later reply. In the two-turn capture, the first message gives order 41982 and the second answer uses that order to provide delivery details without asking for it again. The chat then transfers to a human; that extra step is outside this result. | 1/1 assessed1/1 gradablePublished test set | View Result → |
Escalation and human handoff3 scenarios · 2 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| Agent cannot resolve the request | 1 Pass0 Fail0 Not gradable 1 of 1 test case passedWhat happenedBotpress did hand off an unsupported request rather than trying to complete it itself. It said changing the shipping carrier wasn't something it could do directly, asked for the order number, created the Desk escalation, and a human replied in the thread. | 1/1 assessed1/1 gradablePublished test set | View Result → |
| User explicitly asks for a human | 1 Pass0 Fail0 Not gradable 1 of 1 test case passedWhat happenedBotpress handed off a request for a human without arguing. It replied that it would connect the customer to a human representative, then the Desk side showed a ticket opened, a human assignee replying, and the customer's follow-up arriving in that exchange. | 1/1 assessed1/1 gradablePublished test set | View Result → |
| A configured rule requires escalation | No published result | ||
Intent understanding3 scenarios · 1 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| One message contains multiple requests | 1 Pass0 Fail0 Not gradable 1 of 1 test case passedWhat happenedBotpress answered both parts of a multi-request message in one reply. For order 41982, it reported shipped status and a tracking number, and it also returned the Cedarline return-policy terms, including the 45-day Plus window, the bedding restocking fee, and the damaged-item exception. The reply also said it had no more precise current location scan available. | 1/1 assessed1/1 gradablePublished test set | View Result → |
| Request is unclear and needs clarification | No published result | ||
| Request requires choosing the correct source or action | No published result | ||
Agent configuration and control2 scenarios · 0 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| Configured instruction changes agent behavior | No published result | ||
| Configured restriction is respected | No published result | ||
Analytics and observability3 scenarios · 0 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| Agent actions/tool calls can be inspected | No published result | ||
| Conversation activity can be inspected | No published result | ||
| Reported analytics match what actually happened | No published result | ||
Feedback and learning2 scenarios · 0 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| Correction changes future behavior where the product claims learning | No published result | ||
| Feedback can be submitted | No published result | ||
Knowledge-grounded answering3 scenarios · 0 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| Answer is available in the knowledge base | No published result | ||
| Answer is not available in the knowledge base | No published result | ||
| The request uses different wording than the source | No published result | ||
Multilingual support2 scenarios · 0 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| Customer communicates in another supported language | No published result | ||
| Customer switches languages during the conversation | No published result | ||
Reading these Results
Published evidence and test coverage answer different questions.
Which scenarios have a Result?
A published Result is public evidence for this tool on one scenario. “No published result” does not say whether testing has taken place.
What does each Result cover?
Assessed includes Pass, Fail and Not gradable. Gradable includes Pass and Fail. Both use the pinned test count in that published Result.
Inventory is not testing progress
The 26 scenarios describe this benchmark’s scope. They are not an assumed applicability or test-coverage denominator for Botpress.