Reconstructs a multi-column research page so the 'STUDY AREA' heading, its paragraphs, and the following 'STAND PRESCRIPTIONS' section stay in order.
What was measured
Reading Order & Structure
Maintains headings, sections, column order, document hierarchy, and overall flow.
decisive for this rankingtransformation
Clean Markdown requires the original reading order, headings, sections, and hierarchy to be maintained correctly. (3 of 3 judges)
What was given, what came back
Test input: Scanned Research Paper · pdf · group: scanned-research-paper
Input — what we sent
An image-only scanned research paper used to stress OCR and layout recovery in a multi-column academic document with figures, charts, tables, captions, and references.
Why this input is hard
- · OCR on scanned pages
- · Multi-column reading order
- · Figure and chart handling
- · Table reconstruction from scans
- · Caption association
- · Reference extraction
- · Overall document structure retention
Output — unretouched

Also checked on this input — same tool, 6 other criteria
Complex Document Handling✓ WorkedHandles the scanned research paper end-to-end, including multi-column prose, tables, charts, and handwritten marginalia, while still producing parsed markdown.Table Preservation✓ WorkedRebuilds multi-row tables with the row/column grid intact, preserving treatment rows and year columns in the scanned table extraction.Table Preservation◐ MixedKeeps the numeric table structure, but drops between-column annotations and leaves stray OCR characters in some numeric cells, so contextual information is partially lost.Table Preservation⚠ StruggledBreaks multirow header relationships in a scanned table, so grouped headers and header-level structure are not reliably preserved.Text & OCR Completeness◐ MixedDetects faint handwritten margin text, but only partially; the transcription shows 'USDA Semaine' and the remainder is treated as illegible.Visual Content Retention◐ MixedExtracts chart values into a captioned figure block, but the report says the mortality chart's trend visualization is not fully retained.
Provenance
- Observation
- 4556b641-0861-4cec-9988-0748cdfa1664
- Evidence run
- 6e3160de-fe46-4b45-b071-72560b5c5d0e
- Study
- Convert a Complex PDF into Clean Markdown with an API
- Research task
- 86b9h7t37
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "extend-ai",
scenario: "scanned-research-paper"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 8 other tools
measured on Reading Order & Structure
Adobe API✗ FailedDumps the scanned title page as a dense OCR block without section boundaries or other structural cues, so the document hierarchy is lost.Landing AI✗ FailedMisinterprets the opening page's structure, breaking the relationship between the title and surrounding content on the scanned paper.LlamaParse✓ WorkedReflows a two-column scanned page into a single coherent reading order while preserving section-to-body flow.Mistral AI✗ FailedThe opening page loses the distinction between the document title and the abstract, flattening the semantic organization of the first page.Nutrient.io✓ WorkedThe tool preserves section hierarchy and column order in a dense multi-column scanned section, keeping headings aligned with the correct body text.Reducto⚠ StruggledDisplaces the byline by a full column: the author line that sits above the two-column split in the source is emitted only after the entire left column, and its footnote markers are rendered inconsistently.Tensorlake✓ WorkedRetains section-level reading order in a scanned multi-column paper, with headings continuing to guide the flow across columns and into the next section.Upstage AI✗ FailedBreaks paragraph-level segmentation in the scanned two-column page, so the extracted text no longer follows the source column order cleanly.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com
