Keeps the table values intact but reconstructs the headers incorrectly, leaving a grouped-column table with inconsistent structure.
What was measured
Table Preservation
Preserves complex table structures, including rows, columns, multi-row headers, and merged-cell relationships in markdown.
decisive for this rankingtransformation
Accurate Markdown conversion of complex PDFs depends on keeping table structure intact, not flattening it into plain text. (3 of 3 judges)
What was given, what came back
Test input: Scanned Research Paper · pdf · group: scanned-research-paper
Input — what we sent
An image-only scanned research paper used to stress OCR and layout recovery in a multi-column academic document with figures, charts, tables, captions, and references.
Why this input is hard
- · OCR on scanned pages
- · Multi-column reading order
- · Figure and chart handling
- · Table reconstruction from scans
- · Caption association
- · Reference extraction
- · Overall document structure retention
Output — unretouched

Also checked on this input — same tool, 3 other criteria
Advanced Features✓ WorkedSeparately extracts Figure 3 into year-by-year values for 1972–1981 and tags it as an image asset, showing chart-specific extraction beyond plain OCR.Reading Order & Structure✗ FailedBreaks paragraph-level segmentation in the scanned two-column page, so the extracted text no longer follows the source column order cleanly.Text & OCR Completeness✓ WorkedOCRs dense scanned prose successfully, capturing the ABSTRACT heading and multiple paragraphs of body text rather than only captions or labels.
Provenance
- Observation
- 6db07c2d-fbc7-464c-ba3c-822a7495c3d5
- Evidence run
- 6e3160de-fe46-4b45-b071-72560b5c5d0e
- Study
- Convert a Complex PDF into Clean Markdown with an API
- Research task
- 86b9h7t37
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "upstage-ai",
scenario: "scanned-research-paper"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 8 other tools
measured on Table Preservation
Adobe API✗ FailedBreaks a grouped-column table when intervening text appears, fragmenting the 1979–1981 layout and corrupting the extracted alignment.Extend AI⚠ StruggledBreaks multirow header relationships in a scanned table, so grouped headers and header-level structure are not reliably preserved.Landing AI✓ WorkedPreserves a complex diameter-class table with before-cut, trees-cut-per-acre, and after-cut relationships across the treatment rows and check area.LlamaParse✓ WorkedPreserves a nested treatment table with the 7-inch, 10-inch, 12-inch, 100-leave-tree, and clearcut columns and the acres/live-lodgepole rows.Mistral AI✓ WorkedThe multicolumn table is reconstructed without losing its overall layout logic, so the table structure remains readable in the parsed output.Nutrient.io✓ WorkedLargely preserves grouped-column tables, keeping their internal organization intact in the extracted output.Reducto◐ MixedPartially reconstructs Table 1: most of the roughly 90 numeric values are exact, but literal 0 values in the 12-inch column become blanks, one mean cell picks up stray digits (33.0 830000), and a row-label-only section header is broadcast across all six columns in one instance.Tensorlake✗ FailedStruggles with hierarchical scanned tables, misplacing column headers and producing unreliable reconstructions on both the multicolumn table and the denser complex table.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com
