Benchmark · Version 1

Structured Document Extraction

This benchmark evaluates document extraction platforms that take a user-defined schema and return structured records.

Benchmark overview

7Capabilities
19Scenarios
10Participating tools
What is included and excluded

Evaluates whether a document extraction platform can turn a user-defined schema into correct structured records across layouts, scans and repeated sets, while staying honest about what it could not find.

A buyer learns how far the platform can go in turning documents into usable structured output, and whether it clearly separates extracted facts from missing ones or values found elsewhere in the document.

In scope

  • User-defined schema to structured records from many documents
  • Scalar fields and repeated or nested records
  • Digital, scanned and mixed documents
  • Training from labelled examples
  • Correcting returned data
  • Working with extracted records: filtering, totalling and asking questions

Out of scope

  • Document classification
  • Packet splitting
  • Routing document types to different schemas
  • Transcription quality
  • Non-English documents

Participating tools

Tools in this benchmark’s public roster. Publication availability is not a performance ranking.

ToolPublished ResultsExplore
Extend AI3 scenarios with published ResultsView tool in this benchmark →
FutureSmart Document Intelligence9 scenarios with published ResultsView tool in this benchmark →
Landing AI12 scenarios with published ResultsView tool in this benchmark →
Nanonets6 scenarios with published ResultsView tool in this benchmark →
DatalabNo published resultView tool in this benchmark →
DocsumoNo published resultView tool in this benchmark →
LlamaParseNo published resultView tool in this benchmark →
ReductoNo published resultView tool in this benchmark →
RetabNo published resultView tool in this benchmark →
UnstractNo published resultView tool in this benchmark →

Capabilities & scenarios

19 scenarios grouped by 7 capabilities. Open a group to explore its scenarios in this benchmark.

Field Extraction5 scenarios

Returns the scalar fields the schema asks for — present, interpreted, derived, absent or differently laid out — correctly typed

Capability in this benchmark → · Global definition →

  1. The value is directly available1 tool with a published Result
  2. The value needs a supplied definition1 tool with a published Result
  3. The value must be derived3 tools with published Results
  4. The same field across layouts1 tool with a published Result
  5. The field is absent1 tool with a published Result
Nested & Repeated Fields4 scenarios

Returns repeated sets of records completely, with each value under the right record: none dropped, none invented, none cross-bound

Capability in this benchmark → · Global definition →

  1. A repeated set of records2 tools with published Results
  2. The set continues across a page break3 tools with published Results
  3. The set contains a non-record1 tool with a published Result
  4. The set is legitimately empty3 tools with published Results
Training1 scenario

Learns from labeled examples — documents paired with their expected output — and applies what it learned to a document it hasn't seen

Capability in this benchmark → · Global definition →

  1. A field the tool can only get right from examplesNo published results
OCR2 scenarios

The same schema returns the same values when the document is an image, and a mixed document is handled page by page

Capability in this benchmark → · Global definition →

  1. A cleanly scanned document2 tools with published Results
  2. A document mixing digital and scanned pages4 tools with published Results
Source Grounding2 scenarios

Points to where a returned value came from — and attaches no pointer to a value that wasn't there

Capability in this benchmark → · Global definition →

  1. Trace a value to its location2 tools with published Results
  2. The field is absent1 tool with a published Result
Review and Correction3 scenarios

A human fix reaches the stored record and the systems downstream of it

Capability in this benchmark → · Global definition →

  1. Correct a field valueNo published results
  2. Repair a recordNo published results
  3. Corrected data goes downstream1 tool with a published Result
Querying3 scenarios

Operates over the extracted records — filters, totals, answers questions — and is clear when an answer didn't come from them

Capability in this benchmark → · Global definition →

  1. Filter records across documents2 tools with published Results
  2. Aggregate across documents2 tools with published Results
  3. The answer was never in the schema1 tool with a published Result

Results overview

Current published evidence in this benchmark.

30Current published Results
16Scenarios with published Results
4Tools with published Results

Publication availability is separate from test coverage and unpublished research progress.

Explore scenarios →

How the benchmark works

A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.

Scoring

Unsupported capabilities

A capability a tool doesn't support scores 0, and the 0 stays in the denominator. No tool ranks higher for having less product.

Full benchmark methodology →

Resources and fixtures

The registered material and systems that create a consistent test environment for this benchmark.

Structured Document Extraction — Benchmark definition | AI Demos