Scenario in benchmark · Version 1
A table that continues across a page break
A table continues across a page break and is still treated as one table.
How the tools performed
Every in-scope tool is visible. Outcomes come from current published Results for this scenario and benchmark version.
Publication availability: 1 tool has a current published Result.
No published result does not tell you whether a tool has been tested. Test coverage is shown only for comparable published Results.
| Tool | Published outcomes | Test coverage | Result |
|---|---|---|---|
| Published Results · alphabetical, not ranked | |||
| Mistral AI | Different test setPublished Result available, but not included in this comparison. | — | View Result → |
| No published result · alphabetical | |||
| Adobe PDF Extract API | — | — | No published result |
| Extend AI | — | — | No published result |
| Firecrawl | — | — | No published result |
| GPT-5.6 terra | — | — | No published result |
| Landing AI | — | — | No published result |
| LlamaParse | — | — | No published result |
| Nutrient.io | — | — | No published result |
| PDF Vector | — | — | No published result |
| Reducto | — | — | No published result |
| Tensorlake | — | — | No published result |
| Unstructured | — | — | No published result |
| Upstage AI | — | — | No published result |
Assessed = Pass + Fail + Not gradable. Gradable = Pass + Fail. Both use the published Result’s pinned-test denominator. — means not publicly available.
Test design
- Pinned test cases
- 2
- Disclosure
- 1 public · 1 withheld
- Capabilities exercised here
- Table Extraction
What this scenario evaluates
- Whether the table is recognised as one continuous table rather than split into separate tables.
- Whether the continuation keeps the table header context intact across the page break.
- Whether rows and columns remain aligned as a single structure across both pages.
Exact wording, inputs, fixture state, expected output and detailed grading remain at the test-case level and may be withheld while the benchmark version is active. The scenario and its evaluation intent are public.
How the results are graded
- Pass: the test-case expectations hold.
- Fail: an expectation demonstrably does not hold.
- Not gradable: the evidence cannot establish the outcome.
Version 1 pins a scenario rubric for this scenario.
Benchmark methodology →