Maintains report hierarchy and reading flow in an 84-page annual report, keeping section text, bullets, and business narrative aligned instead of flattening them.
What was measured
Reading Order & Structure
Maintains headings, sections, column order, document hierarchy, and overall flow.
decisive for this rankingtransformation
Clean Markdown requires the original reading order, headings, sections, and hierarchy to be maintained correctly. (3 of 3 judges)
What was given, what came back
Test input: Hybrid Earnings Report · pdf · group: hybrid-earnings-report
Input — what we sent

Hybrid Earnings Report
A long, real-world hybrid annual report used to benchmark end-to-end PDF-to-markdown conversion. It combines native text, financial tables, charts/graphics, and scanned signature/stamp regions, stressing preservation of document structure, reading order, and embedded visual content across a multi-section report.
Why this input is hard
- · Native digital text extraction
- · Complex financial table preservation
- · Chart and graphic retention
- · Image-based signatures and stamps
- · Reading order across a long multi-section report
- · Markdown quality and consistency
Output — unretouched


Also checked on this input — same tool, 4 other criteria
Complex Document Handling✓ WorkedProcesses an 84-page hybrid annual report end-to-end across native text, tables, charts, and scan-only signature regions without manual cleanup.Table Preservation✓ WorkedPreserves a 5-year financial summary table with row labels and year columns aligned, keeping the annual financial grid intact in markdown.Text & OCR Completeness◐ MixedReads low-clarity signer and auditor markings, but the stamp OCR is not perfect: 'LLP' is misread as '1LP'.Visual Content Retention◐ MixedRetains charts and logos as standalone figure blocks with captions, but they are extracted out of page position rather than embedded inline.
Provenance
- Observation
- b69aa8f8-8162-47d0-9487-6b903f8ffe84
- Evidence run
- 6e3160de-fe46-4b45-b071-72560b5c5d0e
- Study
- Convert a Complex PDF into Clean Markdown with an API
- Research task
- 86b9h7t37
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "extend-ai",
scenario: "hybrid-earnings-report"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 7 other tools
measured on Reading Order & Structure
Landing AI✓ WorkedPreserves section flow and heading-to-body relationships in a scanned report section, keeping the '19. Commitments and Contingencies' heading and the 'Data Breach' subsection attached to the following paragraphs.LlamaParse✓ WorkedKeeps the report's heading-and-paragraph sequence intact instead of flattening the page into disconnected text blocks.Mistral AI✓ WorkedThe parser preserves document hierarchy across most sections of the long report, keeping headings and supporting text structurally aligned through the majority of the output.Nutrient.io✓ WorkedPreserves the section hierarchy and bullet-point reading order on the sample page, keeping the 'A Growth Story Again' heading tied to its paragraph and 4 bullets in sequence.Reducto⚠ StruggledDrops heading structure on most of the filing: the FORM 10-K opening area receives no # or ## markup, and the report says 80 of 84 pages stay as plain paragraphs rather than navigable section headings.Tensorlake✓ WorkedPreserves heading order and section relationships in an 84-page hybrid annual report, keeping the narrative, bullets, and figure placement aligned with the source flow.Upstage AI✗ FailedBreaks the two-column reading order on the annual-report page, so headings and body text no longer follow the source navigation cleanly.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com