Benchmark methodology

Web Search for AI

The public evaluation method approved for this benchmark definition.

How the benchmark works

A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.

Endpoint boundary

Single-call search layer

This benchmark tests the single-call search layer, not a multi-step research endpoint.

Subject class

SERP-data APIs count

SERP-data APIs are first-class subjects, and their titles, snippets and answer boxes are treated as source material.

Separate boundary

Web Page → Markdown stays separate

Web Page → Markdown is a different boundary where the URL is already known, and structural fidelity is not graded here.

Answer grounding

Source-limited answers

When a product writes the answer, it should say only what retrieved sources support, decline when support is absent, flag source disagreement, and attribute claims.

Freshness method

Probe before and after

Freshness runs use a before probe and an after probe; if the before probe comes out the wrong way, the run is void, not failed.

Method status

Ranking and scoring remain undesigned

Ranking and scoring are in scope but remain undesigned.

Ranking and scoring are in scope but remain undesigned.

Web Search for AI methodology | AI Demos