AI Database Agents
The public evaluation method approved for this benchmark definition.
How the benchmark works
A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.
Answer, not SQL
The benchmark grades the answer, never the SQL; SQL is captured only as instrumentation.
Scenario is a condition
A scenario is a condition, not a stimulus; changing the stimulus creates a test case, not a new scenario.
Specifications first
Test cases are specifications, not runnable tests, and concrete literals are bound later at fixture implementation.
Simplest direct verification
Each scenario starts with the simplest, most obvious, natural test that directly verifies it.
Still open
Weights, decisive versus context, eligibility, hard-fail caps, and scenario rubrics are still open.
Analytics split under test
Analytics & Observability is tested through aggregate and individual groups so execution can show whether they are really separate capabilities.
Methodology, weighting and ranking decisions are still open in D4.