Benchmark methodology
AI Customer Support Chatbots
The public evaluation method approved for this benchmark definition.
How the benchmark works
A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.
Scenarios describe conditions, not inputs
A scenario is a condition that makes the capability hard, not an input archetype or a smaller capability.
Simplest direct verification
The first test case under a scenario is the simplest, most natural test that directly verifies it.
Add harder variants from execution
Harder variants are added later from observed product differences; depth is earned from execution, not designed upfront.
Methodology still open
Capability weights, decisive/context status, eligibility rules, hard-failure caps, and scenario rubrics are still open.
D4 methodology, weighting, ranking, eligibility, caps, and rubrics are still open.