Scenario in benchmark · Version 1
Aggregate across documents
A request asks for a total over the extracted records from all source documents.
How the tools performed
Every in-scope tool is visible. Outcomes come from current published Results for this scenario and benchmark version.
Publication availability: 2 tools have a current published Result.
No published result does not tell you whether a tool has been tested. Test coverage is shown only for comparable published Results.
| Tool | Published outcomes | Test coverage | Result |
|---|---|---|---|
| Published Results · alphabetical, not ranked | |||
| FutureSmart Document Intelligence | 1 Pass0 Fail0 Not gradable 1 of 1 test case passed | 1 of 1 assessed1 of 1 gradable | View Result → |
| Landing AI | 0 Pass1 Fail0 Not gradable 1 of 1 test case failed | 1 of 1 assessed1 of 1 gradable | View Result → |
| No published result · alphabetical | |||
| Datalab | — | — | No published result |
| Docsumo | — | — | No published result |
| Extend AI | — | — | No published result |
| LlamaParse | — | — | No published result |
| Nanonets | — | — | No published result |
| Reducto | — | — | No published result |
| Retab | — | — | No published result |
| Unstract | — | — | No published result |
Assessed = Pass + Fail + Not gradable. Gradable = Pass + Fail. Both use the published Result’s pinned-test denominator. — means not publicly available.
Test design
- Pinned test cases
- 1
- Disclosure
- 0 public · 1 withheld
- Capabilities exercised here
- Querying
What this scenario evaluates
- Whether it returns a total built from the extracted records from all source documents.
- Whether the total depends on every document contributing to the aggregate.
- Whether it avoids including any value that was not among the extracted records.
- Whether it returns a refusal instead of the total.
Specific test-case inputs and grading live on the related test-case pages.
How the results are graded
- Pass: the test-case expectations hold.
- Fail: an expectation demonstrably does not hold.
- Not gradable: the evidence cannot establish the outcome.
Version 1 uses test-case expectations; no scenario rubric is pinned.
Benchmark methodology →