Capability in benchmark · Version 1

Action execution

The agent can look things up and make changes in connected systems such as orders, subscriptions, or tickets.

5 scenarios · 7 current published Results · 4 tools with published Results

How the tools performed

Explore current published Results across this capability’s scenarios. Open a scenario to compare its tools, or a Result to inspect the evidence.

Each cell reports its own published test set. There is no combined capability score or claim that different Results used identical tests.

Scenarios in this capability
Tools with published Results appear first, alphabetically. Remaining participants stay visible below.
Participating toolRetrieve information from an external system →Update something in an external system →Requested action cannot be completed →Action requires confirmation →User is not authorized to perform the action →
BotpressRetrieve information from an external system →No published resultUpdate something in an external system →No published resultRequested action cannot be completed →No published resultAction requires confirmation →
1 Pass
1/1 assessed1/1 gradableView Result →
User is not authorized to perform the action →
1 Pass
1/1 assessed1/1 gradableView Result →
Freshdesk FreddyRetrieve information from an external system →No published resultUpdate something in an external system →No published resultRequested action cannot be completed →
1 Pass
1/1 assessed1/1 gradableView Result →
Action requires confirmation →
1 Pass
1/1 assessed1/1 gradableView Result →
User is not authorized to perform the action →No published result
KommunicateRetrieve information from an external system →No published resultUpdate something in an external system →No published resultRequested action cannot be completed →No published resultAction requires confirmation →
1 Pass
1/1 assessed1/1 gradableView Result →
User is not authorized to perform the action →No published result
Zendesk AIRetrieve information from an external system →
1 Pass
1/1 assessed1/1 gradableView Result →
Update something in an external system →Different test setPublished Result available, but not included in this comparison.View Result →Requested action cannot be completed →No published resultAction requires confirmation →No published resultUser is not authorized to perform the action →No published result
AdaNo published result across these scenarios
ChatbaseNo published result across these scenarios
CustomGPT.aiNo published result across these scenarios
DecagonNo published result across these scenarios
FS Agent (DIY control)No published result across these scenarios
GorgiasNo published result across these scenarios
Help ScoutNo published result across these scenarios
Intercom FinNo published result across these scenarios
Tidio LyroNo published result across these scenarios

Reading the comparison

Publication availability and test coverage describe different things.

Published evidence

Publication is not testing progress

“No published result” says only that there is no current public Result. It does not indicate whether a tool has been tested, passed, failed or is applicable.

Test coverage

Read each Result’s test set

Assessed includes Pass, Fail and Not gradable. Gradable includes Pass and Fail. Each fraction uses the pinned tests in that same published Result.

Go deeper

Choose the view you need

A scenario compares tools on one situation. A tool page follows one tool across this benchmark. A Result explains the outcome and shows the evidence.

Action execution in AI Customer Support Chatbots | AI Demos