Scenario in benchmark · Version 1
A document with a data chart
A document includes a data chart that must be converted as part of the page.
How the tools performed
Every in-scope tool is visible. Outcomes come from current published Results for this scenario and benchmark version.
No published results for this scenario yet.
No published result does not tell you whether a tool has been tested. Test coverage is shown only for comparable published Results.
| Tool | Published outcomes | Test coverage | Result |
|---|---|---|---|
| No published result · alphabetical | |||
| Adobe PDF Extract API | — | — | No published result |
| Extend AI | — | — | No published result |
| Firecrawl | — | — | No published result |
| GPT-5.6 terra | — | — | No published result |
| Landing AI | — | — | No published result |
| LlamaParse | — | — | No published result |
| Mistral AI | — | — | No published result |
| Nutrient.io | — | — | No published result |
| PDF Vector | — | — | No published result |
| Reducto | — | — | No published result |
| Tensorlake | — | — | No published result |
| Unstructured | — | — | No published result |
| Upstage AI | — | — | No published result |
Assessed = Pass + Fail + Not gradable. Gradable = Pass + Fail. Both use the published Result’s pinned-test denominator. — means not publicly available.
Test design
- Pinned test cases
- 1
- Disclosure
- 1 public · 0 withheld
- Capabilities exercised here
- Figures & Charts
What this scenario evaluates
- Whether the chart remains present as a visual asset in the converted output.
- Whether the chart stays associated with its original position and its title or caption.
- Whether any extra textual chart details are treated as optional enrichment, not as the minimum bar.
- Whether the visual chart is not replaced by prose alone.
Exact wording, inputs, fixture state, expected output and detailed grading remain at the test-case level and may be withheld while the benchmark version is active. The scenario and its evaluation intent are public.
How the results are graded
- Pass: the test-case expectations hold.
- Fail: an expectation demonstrably does not hold.
- Not gradable: the evidence cannot establish the outcome.
Version 1 uses test-case expectations; no scenario rubric is pinned.
Benchmark methodology →