Capability in benchmark · Version 1

Querying

Operates over the extracted records — filters, totals, answers questions — and is clear when an answer didn't come from them

3 scenarios · 5 current published Results · 3 tools with published Results

How the tools performed

Explore current published Results across this capability’s scenarios. Open a scenario to compare its tools, or a Result to inspect the evidence.

Each cell reports its own published test set. There is no combined capability score or claim that different Results used identical tests.

Scenarios in this capability
Tools with published Results appear first, alphabetically. Remaining participants stay visible below.
Participating toolFilter records across documents →Aggregate across documents →The answer was never in the schema →
FutureSmart Document IntelligenceFilter records across documents →No published resultAggregate across documents →
1 Pass
1/1 assessed1/1 gradableView Result →
The answer was never in the schema →No published result
Landing AIFilter records across documents →
1 Fail
1/1 assessed1/1 gradableView Result →
Aggregate across documents →
1 Fail
1/1 assessed1/1 gradableView Result →
The answer was never in the schema →
1 Pass
1/1 assessed1/1 gradableView Result →
NanonetsFilter records across documents →
1 Pass
1/1 assessed1/1 gradableView Result →
Aggregate across documents →No published resultThe answer was never in the schema →No published result
DatalabNo published result across these scenarios
DocsumoNo published result across these scenarios
Extend AINo published result across these scenarios
LlamaParseNo published result across these scenarios
ReductoNo published result across these scenarios
RetabNo published result across these scenarios
UnstractNo published result across these scenarios

Reading the comparison

Publication availability and test coverage describe different things.

Published evidence

Publication is not testing progress

“No published result” says only that there is no current public Result. It does not indicate whether a tool has been tested, passed, failed or is applicable.

Test coverage

Read each Result’s test set

Assessed includes Pass, Fail and Not gradable. Gradable includes Pass and Fail. Each fraction uses the pinned tests in that same published Result.

Go deeper

Choose the view you need

A scenario compares tools on one situation. A tool page follows one tool across this benchmark. A Result explains the outcome and shows the evidence.

Querying in Structured Document Extraction | AI Demos