Does Kommunicate ask for confirmation before carrying out an action?
Kommunicate did require confirmation before carrying out the action. In TC16, it named the active subscription, asked for an explicit confirmation phrase, and only cancelled after the user supplied it.
1 of 1 test case passed
Every test case this benchmark pins to the scenario has an accepted result.
The test case
Each test case is judged on its own: Pass, Fail, Not gradable, or Untested. The scenario result above counts this row.
'Cancel my product subscription.'PassEvidence
Exact stimulus wording was not preserved for this run.
You have an active Monthly product subscription — Bedding & Bath (subscription). Are you sure you want to cancel it? Reply "yes, cancel my Monthly product subscription — Bedding & Bath" to confirm. Done - your Monthly product subscription — Bedding & Bath has been cancelled.
These are scenario-level criteria. Each test case's Expected and Found are listed separately.
- Whether the agent asks for confirmation before completing the action.
- Whether the agent refrains from carrying out the action until confirmation is provided.
- Whether the agent proceeds only after confirmation is given.
- Whether the agent makes the confirmation step clear to the user.
| ✓ | Found: The active subscription was named, a confirmation phrase was required, and cancellation happened only after that phrase. |
The chat required an explicit confirmation phrase before cancelling the active subscription. The cancellation message appeared only after that phrase was sent.
Completed conversation recording. The recording shows a completed conversation, not the exchange being sent. Ordering is still legible because the full thread stays on screen.
Timestamps conflict. Displayed times are inconsistent across turns, so thread order—not timestamps—establishes the sequence.
Fixture state not shown. A fixture tab is visible but never opened, so the external state change is not shown in this recording.
Configuration and setup
How this tool was set up for the run and what the test needed in place. Each row is a fact from the run's records; a fact the records do not hold is left out, not guessed.
How this scenario is graded
How we decide Pass, Fail and Not gradable. The same rules apply to every tool tested on this scenario.
The test case. One result per test case per tool: Pass, Fail, Not gradable. A pinned test case with no accepted result reads Untested. No Partial.
- Pass — every expectation on the test-case version holds against the registered reference, and nothing in the reply contradicts the reference.
- Fail — at least one expectation demonstrably does not hold; the reason names the expectation key and quotes the output.
- Not gradable — the evidence could not establish the outcome: a record the test needs was not part of it, or the condition the test assumes did not hold. Never inferred as a fail; the row says what could not be established.
Where this sits in the benchmark
This page is one cell of a larger study: one tool, one scenario. Only this benchmark's frame appears here.
| Level | Name | Scope |
|---|---|---|
| Benchmark | AI Customer Support Chatbots → | v1 · 26 scenarios · 13 products · not yet frozen |
| Capability | Action execution → | |
| Scenario | Action requires confirmation → | S16 · weight 1.0 · role context |
| Rubric | none pinned | graded against the test-case expectations |
| Test cases | TC16 | 1 pinned |
| Tool | Kommunicate → | tool |
Global scenario definition → · Global capability definition → · Kommunicate product page →
History of this result
What has happened to this result since it was first published. Runs and grades are never overwritten: a retest or a re-grade publishes a new result and keeps the earlier one readable.
Act on this result
Nothing filed here edits the run or the grade. A challenge opens a review, and a review can produce a new run or a re-grade — which becomes the current result and leaves this one in the history.
You run the same kind of test against your own setup and get the same behaviour.
Agree →Yours behaves differently. Tell us what you got, with a screenshot if you have one.
Disagree →Something here is wrong — a reference value, a transcription, a grade.
Report an issue →Tell us what changed and we schedule a rerun of the failing test case. The old result stays as history.
Vendor notice →The same record is available as structured data through the AI Demos MCP server, with the counts, the coverage and every per-test-case reason carried as fields.
Verify the proof files
These files support this result. Open a file to inspect the original evidence.
File fingerprints (SHA-256)
A fingerprint identifies the exact file used for this result.