Structured Document Extraction
This benchmark evaluates document extraction platforms that take a user-defined schema and return structured records.
Benchmark overview
What is included and excluded
Evaluates whether a document extraction platform can turn a user-defined schema into correct structured records across layouts, scans and repeated sets, while staying honest about what it could not find.
A buyer learns how far the platform can go in turning documents into usable structured output, and whether it clearly separates extracted facts from missing ones or values found elsewhere in the document.
In scope
- User-defined schema to structured records from many documents
- Scalar fields and repeated or nested records
- Digital, scanned and mixed documents
- Training from labelled examples
- Correcting returned data
- Working with extracted records: filtering, totalling and asking questions
Out of scope
- Document classification
- Packet splitting
- Routing document types to different schemas
- Transcription quality
- Non-English documents
Participating tools
Tools in this benchmark’s public roster. Publication availability is not a performance ranking.
| Tool | Published Results | Explore |
|---|---|---|
| Extend AI | 3 scenarios with published Results | View tool in this benchmark → |
| FutureSmart Document Intelligence | 9 scenarios with published Results | View tool in this benchmark → |
| Landing AI | 12 scenarios with published Results | View tool in this benchmark → |
| Nanonets | 6 scenarios with published Results | View tool in this benchmark → |
| Datalab | No published result | View tool in this benchmark → |
| Docsumo | No published result | View tool in this benchmark → |
| LlamaParse | No published result | View tool in this benchmark → |
| Reducto | No published result | View tool in this benchmark → |
| Retab | No published result | View tool in this benchmark → |
| Unstract | No published result | View tool in this benchmark → |
Capabilities & scenarios
19 scenarios grouped by 7 capabilities. Open a group to explore its scenarios in this benchmark.
Field Extraction5 scenarios
Returns the scalar fields the schema asks for — present, interpreted, derived, absent or differently laid out — correctly typed
Capability in this benchmark → · Global definition →
- The value is directly available1 tool with a published Result
- The value needs a supplied definition1 tool with a published Result
- The value must be derived3 tools with published Results
- The same field across layouts1 tool with a published Result
- The field is absent1 tool with a published Result
Nested & Repeated Fields4 scenarios
Returns repeated sets of records completely, with each value under the right record: none dropped, none invented, none cross-bound
Capability in this benchmark → · Global definition →
- A repeated set of records2 tools with published Results
- The set continues across a page break3 tools with published Results
- The set contains a non-record1 tool with a published Result
- The set is legitimately empty3 tools with published Results
Training1 scenario
Learns from labeled examples — documents paired with their expected output — and applies what it learned to a document it hasn't seen
Capability in this benchmark → · Global definition →
- A field the tool can only get right from examplesNo published results
OCR2 scenarios
The same schema returns the same values when the document is an image, and a mixed document is handled page by page
Capability in this benchmark → · Global definition →
- A cleanly scanned document2 tools with published Results
- A document mixing digital and scanned pages4 tools with published Results
Source Grounding2 scenarios
Points to where a returned value came from — and attaches no pointer to a value that wasn't there
Capability in this benchmark → · Global definition →
- Trace a value to its location2 tools with published Results
- The field is absent1 tool with a published Result
Review and Correction3 scenarios
A human fix reaches the stored record and the systems downstream of it
Capability in this benchmark → · Global definition →
- Correct a field valueNo published results
- Repair a recordNo published results
- Corrected data goes downstream1 tool with a published Result
Querying3 scenarios
Operates over the extracted records — filters, totals, answers questions — and is clear when an answer didn't come from them
Capability in this benchmark → · Global definition →
- Filter records across documents2 tools with published Results
- Aggregate across documents2 tools with published Results
- The answer was never in the schema1 tool with a published Result
Results overview
Current published evidence in this benchmark.
Publication availability is separate from test coverage and unpublished research progress.
How the benchmark works
A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.
Unsupported capabilities
A capability a tool doesn't support scores 0, and the 0 stays in the denominator. No tool ranks higher for having less product.
Resources and fixtures
The registered material and systems that create a consistent test environment for this benchmark.
Structured Document Extraction fixture corpus — round 1
A fixed round-1 fixture corpus of 19 fictional PDFs for Structured Document Extraction (R5 v1).