Benchmark · Version 1

Managed RAG

This benchmark covers managed RAG and RAG-as-a-Service products that let a buyer upload or connect private documents and get a query API back.

Benchmark overview

5Capabilities
15Scenarios
7Participating tools
What is included and excluded

Managed RAG evaluates services that turn private documents into a query API, so buyers can compare buying a managed service with building the pipeline themselves across grounding, citation and refresh behaviour.

The buying question is build versus buy: whether a managed service is better than assembling and running the retrieval pipeline yourself. Readers learn whether the service makes private knowledge reachable, stays grounded in what it retrieved, cites the supporting source, and catches up when the source changes.

In scope

  • Services where you upload or connect documents and get a query API back.
  • Services that only retrieve and services that also generate answers.
  • Connectors and automatic updates.
  • Retrieval machinery sold as a service, including hybrid search, reranking and metadata filters.

Out of scope

  • Raw vector databases.
  • Website chat widgets.
  • Web search.
  • Document parsing on its own.
  • Multi-tenant isolation and per-user permissions.
  • Structured Document Extraction's querying over already extracted records; here Retrieval works on document text.

Participating tools

Tools in this benchmark’s public roster. Publication availability is not a performance ranking.

ToolPublished ResultsExplore
Dify1 scenario with a published ResultView tool in this benchmark →
Denser AINo published resultView tool in this benchmark →
FutureSmart Document IntelligenceNo published resultView tool in this benchmark →
Llama CloudNo published resultView tool in this benchmark →
OnyxNo published resultView tool in this benchmark →
VectaraNo published resultView tool in this benchmark →
Vectorize.ioNo published resultView tool in this benchmark →

Capabilities & scenarios

15 scenarios grouped by 5 capabilities. Open a group to explore its scenarios in this benchmark.

Data Ingestion4 scenarios

The handed-over files become reachable, and the service reports when ingestion did not work.

Capability in this benchmark → · Global definition →

  1. An ordinary digital text documentNo published results
  2. A simple, clearly formatted tableNo published results
  3. A cleanly scanned document1 tool with a published Result
  4. A bad file in the uploadNo published results
Retrieval4 scenarios

The passage that answers the question comes back from the documents when it exists, and a missing answer is discoverable when it does not.

Capability in this benchmark → · Global definition →

  1. Answer is available in the knowledge baseNo published results
  2. Answer is not available in the knowledge baseNo published results
  3. The request uses different wording than the sourceNo published results
  4. Retrieval is scoped to part of the documentsNo published results
Grounded Generation3 scenarios

The answer stays faithful to the passages the system itself retrieved, including declining when those passages do not support an answer.

Capability in this benchmark → · Global definition →

  1. The retrieved material supports an answerNo published results
  2. The retrieved material doesn't support an answerNo published results
  3. The retrieved sources disagree with each otherNo published results
Citations1 scenario

The cited passage genuinely supports the claim the answer makes, even when the answer itself is wrong.

Capability in this benchmark → · Global definition →

  1. A claim in the answer traced back to its sourceNo published results
Knowledge Sync3 scenarios

The indexed state catches up with the source after content is added, changed, or removed.

Capability in this benchmark → · Global definition →

  1. New content is added to the sourceNo published results
  2. Existing content is changedNo published results
  3. Content is removed from the sourceNo published results

Results overview

Current published evidence in this benchmark.

1Current published Results
1Scenarios with published Results
1Tools with published Results

Publication availability is separate from test coverage and unpublished research progress.

Explore scenarios →

How the benchmark works

A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.

Default setup

Unchanged vendor default

Every tool is tested on its own default setup, unchanged.

Build-vs-buy

Build-it-yourself control

A build-it-yourself pipeline runs as a subject alongside the managed services.

Retrieval unit

Judge retrieval against the documents

Retrieval is judged against the documents: the supporting passage must come back when it exists.

Generation unit

Judge generation against retrieved material

Grounded generation is judged against what the tool actually retrieved, not against the ideal answer.

No double charge

Do not score one miss twice

A retrieval miss is not charged again in generation; declining is correct when the retrieved material does not support an answer.

Citation unit

Judge citations by support

Citations are judged by whether the cited passage supports the claim, not by whether the answer is correct.

Wrong answer can still cite

Correct citation, wrong answer

A correct citation can point to the source of a wrong answer; a citation attached to an invented claim is aggravating.

Zero support

Unsupported capabilities score zero

A capability a tool does not support scores 0, and the 0 stays in the denominator.

Separate outcome

End-to-end answer correctness is separate

End-to-end answer correctness is reported separately as a composite outcome.

Sync completion

Follow the documented refresh path

Each Knowledge Sync test follows the tool's own documented update path to completion; time is recorded, and non-completion counts as non-convergence.

Full benchmark methodology →

Resources and fixtures

The registered material and systems that create a consistent test environment for this benchmark.

Managed RAG — Benchmark definition | AI Demos