Intent understanding
The agent works out what the customer actually wants before it answers or acts.
3 scenarios · 4 current published Results · 3 tools with published Results
How the tools performed
Explore current published Results across this capability’s scenarios. Open a scenario to compare its tools, or a Result to inspect the evidence.
Each cell reports its own published test set. There is no combined capability score or claim that different Results used identical tests.
Scenarios in this capability
| Participating tool | Request requires choosing the correct source or action → | One message contains multiple requests → | Request is unclear and needs clarification → |
|---|---|---|---|
| Botpress | Request requires choosing the correct source or action →No published result | One message contains multiple requests → 1 Pass 1/1 assessed1/1 gradableView Result → | Request is unclear and needs clarification →No published result |
| Freshdesk Freddy | Request requires choosing the correct source or action → 1 Pass 1/1 assessed1/1 gradableView Result → | One message contains multiple requests → 1 Pass 1/1 assessed1/1 gradableView Result → | Request is unclear and needs clarification →No published result |
| Zendesk AI | Request requires choosing the correct source or action →No published result | One message contains multiple requests → 1 Pass 1/1 assessed1/1 gradableView Result → | Request is unclear and needs clarification →No published result |
| Ada | No published result across these scenarios | ||
| Chatbase | No published result across these scenarios | ||
| CustomGPT.ai | No published result across these scenarios | ||
| Decagon | No published result across these scenarios | ||
| FS Agent (DIY control) | No published result across these scenarios | ||
| Gorgias | No published result across these scenarios | ||
| Help Scout | No published result across these scenarios | ||
| Intercom Fin | No published result across these scenarios | ||
| Kommunicate | No published result across these scenarios | ||
| Tidio Lyro | No published result across these scenarios | ||
Reading the comparison
Publication availability and test coverage describe different things.
Publication is not testing progress
“No published result” says only that there is no current public Result. It does not indicate whether a tool has been tested, passed, failed or is applicable.
Read each Result’s test set
Assessed includes Pass, Fail and Not gradable. Gradable includes Pass and Fail. Each fraction uses the pinned tests in that same published Result.
Choose the view you need
A scenario compares tools on one situation. A tool page follows one tool across this benchmark. A Result explains the outcome and shows the evidence.