Web Search for AI
This benchmark looks at live-web search and answer products used by AI applications.
Benchmark overview
What is included and excluded
Evaluates whether live-web search and answer APIs find supporting material, notice when it is absent, keep up with web changes, and stay grounded in their own sources when they answer.
A reader learns how these products behave when support exists, when it is missing, when pages change, and when sources disagree. It is about the buyer question, not about a particular corpus or interface.
In scope
- Live-web search and answer APIs used by AI applications.
- Products that must find the URL rather than be given it.
- SERP-data APIs as first-class subjects.
- The single-call search layer.
- Live-web source material that includes titles, snippets and answer boxes.
Out of scope
- Web Page → Markdown.
- Managed RAG.
- Multi-step research endpoints.
- Structural fidelity grading.
Participating tools
Tools in this benchmark’s public roster. Publication availability is not a performance ranking.
| Tool | Published Results | Explore |
|---|---|---|
| Brave Search API | No published result | View tool in this benchmark → |
| Exa | No published result | View tool in this benchmark → |
| Firecrawl | No published result | View tool in this benchmark → |
| GPT-5.6 terra | No published result | View tool in this benchmark → |
| Linkup | No published result | View tool in this benchmark → |
| OpenAI web search | No published result | View tool in this benchmark → |
| Perplexity Sonar | No published result | View tool in this benchmark → |
| SerpAPI | No published result | View tool in this benchmark → |
| Serper | No published result | View tool in this benchmark → |
| Tavily | No published result | View tool in this benchmark → |
| Valyu | No published result | View tool in this benchmark → |
| You.com | No published result | View tool in this benchmark → |
Capabilities & scenarios
10 scenarios grouped by 3 capabilities. Open a group to explore its scenarios in this benchmark.
Web Retrieval4 scenarios
Web Retrieval means the product can find live-web material that supports an answer when it exists, and can tell when it does not.
Capability in this benchmark → · Global definition →
- The request uses different wording than the sourceNo published results
- A page on the web answers the questionNo published results
- Something more famous has the same nameNo published results
- Search only these websitesNo published results
Freshness3 scenarios
Freshness means the product's results match the web as it is now, including new pages, changed pages, and removed pages.
Capability in this benchmark → · Global definition →
- Content is removed from the sourceNo published results
- A new page has just gone liveNo published results
- A page the tool already finds has changedNo published results
Answer Generation3 scenarios
Answer Generation means that, where the product writes the answer, it says only what retrieved sources support, declines when support is absent, flags disagreement, and attributes claims.
Capability in this benchmark → · Global definition →
- The retrieved material supports an answerNo published results
- The retrieved material doesn't support an answerNo published results
- The retrieved sources disagree with each otherNo published results
Results overview
Current published evidence in this benchmark.
No published results for this benchmark yet.
Publication availability is separate from test coverage and unpublished research progress.
How the benchmark works
A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.
Single-call search layer
This benchmark tests the single-call search layer, not a multi-step research endpoint.
SERP-data APIs count
SERP-data APIs are first-class subjects, and their titles, snippets and answer boxes are treated as source material.
Web Page → Markdown stays separate
Web Page → Markdown is a different boundary where the URL is already known, and structural fidelity is not graded here.
Source-limited answers
When a product writes the answer, it should say only what retrieved sources support, decline when support is absent, flag source disagreement, and attribute claims.
Probe before and after
Freshness runs use a before probe and an after probe; if the before probe comes out the wrong way, the run is void, not failed.