Managed RAG
This benchmark covers managed RAG and RAG-as-a-Service products that let a buyer upload or connect private documents and get a query API back.
Benchmark overview
What is included and excluded
Managed RAG evaluates services that turn private documents into a query API, so buyers can compare buying a managed service with building the pipeline themselves across grounding, citation and refresh behaviour.
The buying question is build versus buy: whether a managed service is better than assembling and running the retrieval pipeline yourself. Readers learn whether the service makes private knowledge reachable, stays grounded in what it retrieved, cites the supporting source, and catches up when the source changes.
In scope
- Services where you upload or connect documents and get a query API back.
- Services that only retrieve and services that also generate answers.
- Connectors and automatic updates.
- Retrieval machinery sold as a service, including hybrid search, reranking and metadata filters.
Out of scope
- Raw vector databases.
- Website chat widgets.
- Web search.
- Document parsing on its own.
- Multi-tenant isolation and per-user permissions.
- Structured Document Extraction's querying over already extracted records; here Retrieval works on document text.
Participating tools
Tools in this benchmark’s public roster. Publication availability is not a performance ranking.
| Tool | Published Results | Explore |
|---|---|---|
| Dify | 1 scenario with a published Result | View tool in this benchmark → |
| Denser AI | No published result | View tool in this benchmark → |
| FutureSmart Document Intelligence | No published result | View tool in this benchmark → |
| Llama Cloud | No published result | View tool in this benchmark → |
| Onyx | No published result | View tool in this benchmark → |
| Vectara | No published result | View tool in this benchmark → |
| Vectorize.io | No published result | View tool in this benchmark → |
Capabilities & scenarios
15 scenarios grouped by 5 capabilities. Open a group to explore its scenarios in this benchmark.
Data Ingestion4 scenarios
The handed-over files become reachable, and the service reports when ingestion did not work.
Capability in this benchmark → · Global definition →
- An ordinary digital text documentNo published results
- A simple, clearly formatted tableNo published results
- A cleanly scanned document1 tool with a published Result
- A bad file in the uploadNo published results
Retrieval4 scenarios
The passage that answers the question comes back from the documents when it exists, and a missing answer is discoverable when it does not.
Capability in this benchmark → · Global definition →
- Answer is available in the knowledge baseNo published results
- Answer is not available in the knowledge baseNo published results
- The request uses different wording than the sourceNo published results
- Retrieval is scoped to part of the documentsNo published results
Grounded Generation3 scenarios
The answer stays faithful to the passages the system itself retrieved, including declining when those passages do not support an answer.
Capability in this benchmark → · Global definition →
- The retrieved material supports an answerNo published results
- The retrieved material doesn't support an answerNo published results
- The retrieved sources disagree with each otherNo published results
Citations1 scenario
The cited passage genuinely supports the claim the answer makes, even when the answer itself is wrong.
Capability in this benchmark → · Global definition →
- A claim in the answer traced back to its sourceNo published results
Knowledge Sync3 scenarios
The indexed state catches up with the source after content is added, changed, or removed.
Capability in this benchmark → · Global definition →
- New content is added to the sourceNo published results
- Existing content is changedNo published results
- Content is removed from the sourceNo published results
Results overview
Current published evidence in this benchmark.
Publication availability is separate from test coverage and unpublished research progress.
How the benchmark works
A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.
Unchanged vendor default
Every tool is tested on its own default setup, unchanged.
Build-it-yourself control
A build-it-yourself pipeline runs as a subject alongside the managed services.
Judge retrieval against the documents
Retrieval is judged against the documents: the supporting passage must come back when it exists.
Judge generation against retrieved material
Grounded generation is judged against what the tool actually retrieved, not against the ideal answer.
Do not score one miss twice
A retrieval miss is not charged again in generation; declining is correct when the retrieved material does not support an answer.
Judge citations by support
Citations are judged by whether the cited passage supports the claim, not by whether the answer is correct.
Correct citation, wrong answer
A correct citation can point to the source of a wrong answer; a citation attached to an invented claim is aggravating.
Unsupported capabilities score zero
A capability a tool does not support scores 0, and the 0 stays in the denominator.
End-to-end answer correctness is separate
End-to-end answer correctness is reported separately as a composite outcome.
Follow the documented refresh path
Each Knowledge Sync test follows the tool's own documented update path to completion; time is recorded, and non-completion counts as non-convergence.
Resources and fixtures
The registered material and systems that create a consistent test environment for this benchmark.
Grounded QA fixture corpora — Fenwake documents for the Managed RAG ranking
A fictional document corpus used to evaluate managed RAG systems.