Recovers dense readable prose from a scanned page-image source, including the section heading and multiple long paragraphs.

✓ Worked🧾 artifact-verifiedinput + output shownTest date not recordedLlamaParse
What was measured
Text & OCR Completeness

Extracts all readable content, including scanned pages, with accurate OCR and minimal omissions.

decisive for this rankingtransformation

If the tool misses readable text or fails on scanned pages, it has not actually converted the PDF faithfully into Markdown. (3 of 3 judges)

What was given, what came back

Test input: Scanned Research Paper · pdf · group: scanned-research-paper
Input — what we sent
Input file 1 — as supplied
Input file 1 — as supplied
Input file 2 — as supplied
Scanned Research PDF.pdf
Scanned Research Paper

An image-only scanned research paper used to stress OCR and layout recovery in a multi-column academic document with figures, charts, tables, captions, and references.

Why this input is hard
  • · OCR on scanned pages
  • · Multi-column reading order
  • · Figure and chart handling
  • · Table reconstruction from scans
  • · Caption association
  • · Reference extraction
  • · Overall document structure retention
Output — unretouched
image
Provenance
Observation
04bda6ca-e954-4d0b-93a9-4220eb0b980a
Evidence run
6e3160de-fe46-4b45-b071-72560b5c5d0e
Study
Convert a Complex PDF into Clean Markdown with an API
Research task
86b9h7t37
Tested at
not recorded
Source
first-party
Evidence state
verified
Proof shown
input + output shown
Cost / latency
not captured
Repeat run
not captured
Tester
not captured

The last three rows are honest blanks, not placeholders — our capture has no field for them yet.

Query this
get_evidence({
  tool: "llamaparse",
  scenario: "scanned-research-paper"
})
MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 8 other tools
measured on Text & OCR Completeness
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com