The operator needs to find the questions the agent is failing on
A reviewer can locate failed or unanswered questions, and they are surfaced as failures rather than ordinary successful answers.
What this scenario means
This scenario checks whether a system keeps failure states visible in its observability records. A good agent does not hide unanswered or failed questions inside ordinary answer logs. It lets the person reviewing runs find those cases, inspect them, and see that they are classified as failures.
What we evaluate
- Whether failed or unanswered questions can be found from the recorded activity.
- Whether those cases are surfaced as failures rather than as ordinary answered questions.
- Whether the system makes the failure state distinguishable from a successful answer.
Capabilities this scenario exercises
A scenario may exercise one or more capabilities.
Analytics & Observability
Lets whoever runs the agent see what people actually asked and what actually happened — and reports numbers that match reality. ON TRIAL: aggregate analytics are S57 and S24; individual inspection is S58 and S59. Demotes if the scenarios cannot discriminate. Deliberately NOT the same capability as the Customer Support benchmark's 'Analytics and observability' (C8) — there the operator is a different person from the user; here the buyer IS the user and sees every answer (rule C-10).
Benchmarks that use this scenario
A scenario has global identity and may be reused across benchmarks.
AI Database Agents
Which AI database agent answers business questions about a live relational database correctly — in the company's own terms, and honestly when the data cannot answer?