Recovers dense readable prose from a scanned page-image source, including the section heading and multiple long paragraphs.
What was measured
Text & OCR Completeness
Extracts all readable content, including scanned pages, with accurate OCR and minimal omissions.
decisive for this rankingtransformation
If the tool misses readable text or fails on scanned pages, it has not actually converted the PDF faithfully into Markdown. (3 of 3 judges)
What was given, what came back
Test input: Scanned Research Paper · pdf · group: scanned-research-paper
Input — what we sent
An image-only scanned research paper used to stress OCR and layout recovery in a multi-column academic document with figures, charts, tables, captions, and references.
Why this input is hard
- · OCR on scanned pages
- · Multi-column reading order
- · Figure and chart handling
- · Table reconstruction from scans
- · Caption association
- · Reference extraction
- · Overall document structure retention
Output — unretouched

Also checked on this input — same tool, 6 other criteria
Complex Document Handling✓ WorkedProcesses a 12-page scanned paper end-to-end and reaches SUCCESS after extracting the page content and figure/table outputs.Reading Order & Structure✓ WorkedReflows a two-column scanned page into a single coherent reading order while preserving section-to-body flow.Table Preservation✓ WorkedPreserves the harvest-diameter table's rows and aligned columns in the extracted output.Table Preservation✗ FailedFails to faithfully reconstruct grouped headers in a multi-level stand-effects table, making parent-child column relationships ambiguous.Table Preservation✓ WorkedPreserves a nested treatment table with the 7-inch, 10-inch, 12-inch, 100-leave-tree, and clearcut columns and the acres/live-lodgepole rows.Visual Content Retention◐ MixedConverts a bar chart into a structured table, preserving the legend/value mapping but not keeping the chart as a visual chart.
Provenance
- Observation
- 04bda6ca-e954-4d0b-93a9-4220eb0b980a
- Evidence run
- 6e3160de-fe46-4b45-b071-72560b5c5d0e
- Study
- Convert a Complex PDF into Clean Markdown with an API
- Research task
- 86b9h7t37
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "llamaparse",
scenario: "scanned-research-paper"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 8 other tools
measured on Text & OCR Completeness
Adobe API✓ WorkedRecovers the visible title, abstract, keywords, and opening paragraphs from a scanned USDA forestry report as dense OCR text.Extend AI◐ MixedDetects faint handwritten margin text, but only partially; the transcription shows 'USDA Semaine' and the remainder is treated as illegible.Landing AI◐ MixedOCRs most of the scanned cover-page text, including the agency header, report number/date, title, authors, abstract, and keywords, but inserts noisy tokens such as '186153' and 'USA/-' into the title line.Mistral AI◐ MixedThe report says the parser recovers much of the underlying text from the scanned paper, but it does not present a measured completeness rate and the first-page hierarchy is still lossy.Nutrient.io✓ WorkedRecovers readable text from scanned pages, including the abstract, keywords, author affiliations, title, and opening paragraphs on the first page.PDFVector✓ WorkedSuccessfully parsed a 12-page scanned research paper and produced a long extracted-text preview in 20.6 s using 48 credits.Reducto✓ WorkedConverts all 12 scanned pages with no gaps; the page-11-to-page-12 handoff is preserved verbatim, and a separate page-marker output shows pages 1 through 12 present with no missing markers.Upstage AI✓ WorkedOCRs dense scanned prose successfully, capturing the ABSTRACT heading and multiple paragraphs of body text rather than only captions or labels.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com
