Managed RAG
The public evaluation method approved for this benchmark definition.
How the benchmark works
A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.
Unchanged vendor default
Every tool is tested on its own default setup, unchanged.
Build-it-yourself control
A build-it-yourself pipeline runs as a subject alongside the managed services.
Judge retrieval against the documents
Retrieval is judged against the documents: the supporting passage must come back when it exists.
Judge generation against retrieved material
Grounded generation is judged against what the tool actually retrieved, not against the ideal answer.
Do not score one miss twice
A retrieval miss is not charged again in generation; declining is correct when the retrieved material does not support an answer.
Judge citations by support
Citations are judged by whether the cited passage supports the claim, not by whether the answer is correct.
Correct citation, wrong answer
A correct citation can point to the source of a wrong answer; a citation attached to an invented claim is aggravating.
Unsupported capabilities score zero
A capability a tool does not support scores 0, and the 0 stays in the denominator.
End-to-end answer correctness is separate
End-to-end answer correctness is reported separately as a composite outcome.
Follow the documented refresh path
Each Knowledge Sync test follows the tool's own documented update path to completion; time is recorded, and non-completion counts as non-convergence.
Scoring is still open
Scoring, weighting, critical failures, the eligibility contract, and corpus scale are still open methodology work.
V1 is accepted, but scoring and scenario rubrics are still open.