Scenario definition · S57

Someone needs to see what people have been asking

A usage view shows what people asked and attributes those questions correctly.

What this scenario means

This scenario checks whether the product can surface a usage view that reflects real demand, not just raw logs. A good agent shows the questions people asked, keeps them attributable, and accounts for the full set without dropping, merging, or mislabelling entries. That matters because oversight depends on seeing actual usage clearly.

What we evaluate

  • Whether the usage view lists the questions that were asked.
  • Whether each question is attributed to the source that asked it.
  • Whether the view accounts for the full set of questions, not just a subset.
  • Whether the totals or counts shown by the view match the underlying activity.
◌Test-case detail. Exact wording, inputs, fixture state, expected output and detailed grading remain at the test-case level and may be withheld while the benchmark version is active. The scenario and its evaluation intent are public.

Capabilities this scenario exercises

A scenario may exercise one or more capabilities.

Capability

Analytics & Observability

Lets whoever runs the agent see what people actually asked and what actually happened — and reports numbers that match reality. ON TRIAL: aggregate analytics are S57 and S24; individual inspection is S58 and S59. Demotes if the scenarios cannot discriminate. Deliberately NOT the same capability as the Customer Support benchmark's 'Analytics and observability' (C8) — there the operator is a different person from the user; here the buyer IS the user and sees every answer (rule C-10).

Benchmarks that use this scenario

A scenario has global identity and may be reused across benchmarks.

Someone needs to see what people have been asking — Scenario definition | AI Demos