Detects faint handwritten margin text, but only partially; the transcription shows 'USDA Semaine' and the remainder is treated as illegible.
What was measured
Text & OCR Completeness
Extracts all readable content, including scanned pages, with accurate OCR and minimal omissions.
decisive for this rankingtransformation
If the tool misses readable text or fails on scanned pages, it has not actually converted the PDF faithfully into Markdown. (3 of 3 judges)
What was given, what came back
Test input: Scanned Research Paper · pdf · group: scanned-research-paper
Input — what we sent
An image-only scanned research paper used to stress OCR and layout recovery in a multi-column academic document with figures, charts, tables, captions, and references.
Why this input is hard
- · OCR on scanned pages
- · Multi-column reading order
- · Figure and chart handling
- · Table reconstruction from scans
- · Caption association
- · Reference extraction
- · Overall document structure retention
Output — unretouched

Also checked on this input — same tool, 6 other criteria
Complex Document Handling✓ WorkedHandles the scanned research paper end-to-end, including multi-column prose, tables, charts, and handwritten marginalia, while still producing parsed markdown.Reading Order & Structure✓ WorkedReconstructs a multi-column research page so the 'STUDY AREA' heading, its paragraphs, and the following 'STAND PRESCRIPTIONS' section stay in order.Table Preservation✓ WorkedRebuilds multi-row tables with the row/column grid intact, preserving treatment rows and year columns in the scanned table extraction.Table Preservation◐ MixedKeeps the numeric table structure, but drops between-column annotations and leaves stray OCR characters in some numeric cells, so contextual information is partially lost.Table Preservation⚠ StruggledBreaks multirow header relationships in a scanned table, so grouped headers and header-level structure are not reliably preserved.Visual Content Retention◐ MixedExtracts chart values into a captioned figure block, but the report says the mortality chart's trend visualization is not fully retained.
Provenance
- Observation
- 6125882c-9dbc-46e0-99d4-07315648ef3c
- Evidence run
- 6e3160de-fe46-4b45-b071-72560b5c5d0e
- Study
- Convert a Complex PDF into Clean Markdown with an API
- Research task
- 86b9h7t37
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "extend-ai",
scenario: "scanned-research-paper"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 8 other tools
measured on Text & OCR Completeness
Adobe API✓ WorkedRecovers the visible title, abstract, keywords, and opening paragraphs from a scanned USDA forestry report as dense OCR text.Landing AI◐ MixedOCRs most of the scanned cover-page text, including the agency header, report number/date, title, authors, abstract, and keywords, but inserts noisy tokens such as '186153' and 'USA/-' into the title line.LlamaParse✓ WorkedRecovers dense readable prose from a scanned page-image source, including the section heading and multiple long paragraphs.Mistral AI◐ MixedThe report says the parser recovers much of the underlying text from the scanned paper, but it does not present a measured completeness rate and the first-page hierarchy is still lossy.Nutrient.io✓ WorkedRecovers readable text from scanned pages, including the abstract, keywords, author affiliations, title, and opening paragraphs on the first page.PDFVector✓ WorkedSuccessfully parsed a 12-page scanned research paper and produced a long extracted-text preview in 20.6 s using 48 credits.Reducto✓ WorkedConverts all 12 scanned pages with no gaps; the page-11-to-page-12 handoff is preserved verbatim, and a separate page-marker output shows pages 1 through 12 present with no missing markers.Upstage AI✓ WorkedOCRs dense scanned prose successfully, capturing the ABSTRACT heading and multiple paragraphs of body text rather than only captions or labels.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com
