The export is usable markdown rather than a flat text dump: it includes an overall markdown file plus page-wise markdown files delivered in a downloadable ZIP.
What was measured
Markdown Quality
Produces clean, well-structured, usable markdown rather than a flat text dump.
decisive for this rankingtransformation
The ranking is specifically about producing clean Markdown, so the quality and usability of the Markdown output directly measures success. (3 of 3 judges)
What was given, what came back
Test input: Hybrid Earnings Report · pdf · group: hybrid-earnings-report
Input — what we sent
07829e6dcb414445886e618cd047efb2.pdf
Hybrid Earnings Report
A long, real-world hybrid annual report used to benchmark end-to-end PDF-to-markdown conversion. It combines native text, financial tables, charts/graphics, and scanned signature/stamp regions, stressing preservation of document structure, reading order, and embedded visual content across a multi-section report.
Why this input is hard
- · Native digital text extraction
- · Complex financial table preservation
- · Chart and graphic retention
- · Image-based signatures and stamps
- · Reading order across a long multi-section report
- · Markdown quality and consistency
Also checked on this input — same tool, 5 other criteria
Complex Document Handling✓ WorkedThe tool handles an 84-page mixed-content annual report end-to-end, including tables, charts, and scanned signature/stamp regions, and returns usable markdown output without manual correction.Reading Order & Structure✓ WorkedThe parser preserves document hierarchy across most sections of the long report, keeping headings and supporting text structurally aligned through the majority of the output.Reading Order & Structure◐ MixedHeading hierarchy is reconstructed inconsistently, so some parts of the document are flattened even though the underlying content is still recovered.Table Preservation✓ WorkedThe financial table is reconstructed into a usable structured layout without losing the relationships between headers, rows, and corresponding values.Visual Content Retention✓ WorkedVisual assets are extracted into page-specific folders, keeping charts, signatures, and other document elements tied to their source pages instead of collapsing them into one consolidated output.
Provenance
- Observation
- ec898578-9641-4f6a-b228-bebde59abc51
- Evidence run
- 6e3160de-fe46-4b45-b071-72560b5c5d0e
- Study
- Convert a Complex PDF into Clean Markdown with an API
- Research task
- 86b9h7t37
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "mistral-ai",
scenario: "hybrid-earnings-report"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 4 other tools
measured on Markdown Quality
PDF.ai✗ FailedMarkdown export failed on an 84-page hybrid earnings report; the report records "Output MD: Output failed."PDFVector◐ MixedReturned markdown-like content inside JSON, but the document structure was largely flattened rather than preserved as faithful markdown.Reducto⚠ StruggledProduces usable but not clean markdown: the output contains 12 <signature> tags, 2 <empty> placeholders, and literal HTML tags such as <b>, <i>, and <u> instead of pure CommonMark.Tensorlake✓ WorkedRenders the extraction as structured, copyable markdown in the Document Markdown view rather than a flat text dump.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com