Benchmark methodology

AI Database Agents

The public evaluation method approved for this benchmark definition.

How the benchmark works

A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.

Graded artefact

Answer, not SQL

The benchmark grades the answer, never the SQL; SQL is captured only as instrumentation.

Evaluation unit

Scenario is a condition

A scenario is a condition, not a stimulus; changing the stimulus creates a test case, not a new scenario.

Test form

Specifications first

Test cases are specifications, not runnable tests, and concrete literals are bound later at fixture implementation.

Test design

Simplest direct verification

Each scenario starts with the simplest, most obvious, natural test that directly verifies it.

Open methodology

Still open

Weights, decisive versus context, eligibility, hard-fail caps, and scenario rubrics are still open.

On-trial grouping

Analytics split under test

Analytics & Observability is tested through aggregate and individual groups so execution can show whether they are really separate capabilities.

Methodology, weighting and ranking decisions are still open in D4.

AI Database Agents methodology | AI Demos