Analytics & Observability
Lets whoever runs the agent see what people actually asked and what actually happened — and reports numbers that match reality. ON TRIAL: aggregate analytics are S57 and S24; individual inspection is S58 and S59. Demotes if the scenarios cannot discriminate. Deliberately NOT the same capability as the Customer Support benchmark's 'Analytics and observability' (C8) — there the operator is a different person from the user; here the buyer IS the user and sees every answer (rule C-10).
What this capability means
Lets whoever runs the agent see what people actually asked and what actually happened, and whether the reported numbers match reality.
Boundary: Does not judge a separate operator/user support workflow where the operator is a different person from the user.
Scenarios that test this capability
A scenario is a real-world situation used to test a capability.
Benchmarks that include this capability
A capability has global identity and may be used by more than one benchmark.
AI Database Agents
Which AI database agent answers business questions about a live relational database correctly — in the company's own terms, and honestly when the data cannot answer?