Benchmark · Version 1

Web Search for AI

This benchmark looks at live-web search and answer products used by AI applications.

Benchmark overview

3Capabilities
10Scenarios
12Participating tools
What is included and excluded

Evaluates whether live-web search and answer APIs find supporting material, notice when it is absent, keep up with web changes, and stay grounded in their own sources when they answer.

A reader learns how these products behave when support exists, when it is missing, when pages change, and when sources disagree. It is about the buyer question, not about a particular corpus or interface.

In scope

  • Live-web search and answer APIs used by AI applications.
  • Products that must find the URL rather than be given it.
  • SERP-data APIs as first-class subjects.
  • The single-call search layer.
  • Live-web source material that includes titles, snippets and answer boxes.

Out of scope

  • Web Page → Markdown.
  • Managed RAG.
  • Multi-step research endpoints.
  • Structural fidelity grading.

Participating tools

Tools in this benchmark’s public roster. Publication availability is not a performance ranking.

ToolPublished ResultsExplore
Brave Search APINo published resultView tool in this benchmark →
ExaNo published resultView tool in this benchmark →
FirecrawlNo published resultView tool in this benchmark →
GPT-5.6 terraNo published resultView tool in this benchmark →
LinkupNo published resultView tool in this benchmark →
OpenAI web searchNo published resultView tool in this benchmark →
Perplexity SonarNo published resultView tool in this benchmark →
SerpAPINo published resultView tool in this benchmark →
SerperNo published resultView tool in this benchmark →
TavilyNo published resultView tool in this benchmark →
ValyuNo published resultView tool in this benchmark →
You.comNo published resultView tool in this benchmark →

Capabilities & scenarios

10 scenarios grouped by 3 capabilities. Open a group to explore its scenarios in this benchmark.

Web Retrieval4 scenarios

Web Retrieval means the product can find live-web material that supports an answer when it exists, and can tell when it does not.

Capability in this benchmark → · Global definition →

  1. The request uses different wording than the sourceNo published results
  2. A page on the web answers the questionNo published results
  3. Something more famous has the same nameNo published results
  4. Search only these websitesNo published results
Freshness3 scenarios

Freshness means the product's results match the web as it is now, including new pages, changed pages, and removed pages.

Capability in this benchmark → · Global definition →

  1. Content is removed from the sourceNo published results
  2. A new page has just gone liveNo published results
  3. A page the tool already finds has changedNo published results
Answer Generation3 scenarios

Answer Generation means that, where the product writes the answer, it says only what retrieved sources support, declines when support is absent, flags disagreement, and attributes claims.

Capability in this benchmark → · Global definition →

  1. The retrieved material supports an answerNo published results
  2. The retrieved material doesn't support an answerNo published results
  3. The retrieved sources disagree with each otherNo published results

Results overview

Current published evidence in this benchmark.

0Current published Results
0Scenarios with published Results
0Tools with published Results

No published results for this benchmark yet.

Publication availability is separate from test coverage and unpublished research progress.

Explore scenarios →

How the benchmark works

A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.

Endpoint boundary

Single-call search layer

This benchmark tests the single-call search layer, not a multi-step research endpoint.

Subject class

SERP-data APIs count

SERP-data APIs are first-class subjects, and their titles, snippets and answer boxes are treated as source material.

Separate boundary

Web Page → Markdown stays separate

Web Page → Markdown is a different boundary where the URL is already known, and structural fidelity is not graded here.

Answer grounding

Source-limited answers

When a product writes the answer, it should say only what retrieved sources support, decline when support is absent, flag source disagreement, and attribute claims.

Freshness method

Probe before and after

Freshness runs use a before probe and an after probe; if the before probe comes out the wrong way, the run is void, not failed.

Full benchmark methodology →

Web Search for AI — Benchmark definition | AI Demos