Web Search for AI
The public evaluation method approved for this benchmark definition.
How the benchmark works
A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.
Single-call search layer
This benchmark tests the single-call search layer, not a multi-step research endpoint.
SERP-data APIs count
SERP-data APIs are first-class subjects, and their titles, snippets and answer boxes are treated as source material.
Web Page → Markdown stays separate
Web Page → Markdown is a different boundary where the URL is already known, and structural fidelity is not graded here.
Source-limited answers
When a product writes the answer, it should say only what retrieved sources support, decline when support is absent, flag source disagreement, and attribute claims.
Probe before and after
Freshness runs use a before probe and an after probe; if the before probe comes out the wrong way, the run is void, not failed.
Ranking and scoring remain undesigned
Ranking and scoring are in scope but remain undesigned.
Ranking and scoring are in scope but remain undesigned.