Renders the extraction as structured, copyable markdown in the Document Markdown view rather than a flat text dump.
What was measured
Markdown Quality
Produces clean, well-structured, usable markdown rather than a flat text dump.
decisive for this rankingtransformation
The ranking is specifically about producing clean Markdown, so the quality and usability of the Markdown output directly measures success. (3 of 3 judges)
What was given, what came back
Test input: Hybrid Earnings Report · pdf · group: hybrid-earnings-report
Input — what we sent
A long, real-world hybrid annual report used to benchmark end-to-end PDF-to-markdown conversion. It combines native text, financial tables, charts/graphics, and scanned signature/stamp regions, stressing preservation of document structure, reading order, and embedded visual content across a multi-section report.
Why this input is hard
- · Native digital text extraction
- · Complex financial table preservation
- · Chart and graphic retention
- · Image-based signatures and stamps
- · Reading order across a long multi-section report
- · Markdown quality and consistency
Output — unretouched

Loading file...
Also checked on this input — same tool, 6 other criteria
Advanced Features✓ WorkedProvides separate chart extraction for the SG&A waterfall, outputting chart metadata with a title, axis labels, 10 category labels, and a numeric series instead of only inline prose.Complex Document Handling✓ WorkedProcesses an 84-page mixed-content annual report without collapsing the hierarchy, keeping text, tables, charts, and scanned signatures usable within the extracted workflow.Reading Order & Structure✓ WorkedPreserves heading order and section relationships in an 84-page hybrid annual report, keeping the narrative, bullets, and figure placement aligned with the source flow.Table Preservation✓ WorkedKeeps a 5-year financial summary table structured, retaining the row/column relationships across sales, expenses, EBIT, and per-share rows instead of flattening the table.Text & OCR Completeness◐ MixedOn a blurry Ernst & Young signoff, the OCR preserves the firm reference but makes a symbol-level mistake by rendering the ampersand as '+', so the text is close but not exact.Visual Content Retention✓ WorkedRetains scanned signature-page visuals as figure blocks plus signer text, including named executives and dates, rather than dropping the signature regions from the output.
Provenance
- Observation
- 449b0ccc-3d5d-467a-b98b-79653db91504
- Evidence run
- 6e3160de-fe46-4b45-b071-72560b5c5d0e
- Study
- Convert a Complex PDF into Clean Markdown with an API
- Research task
- 86b9h7t37
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "tensorlake",
scenario: "hybrid-earnings-report"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 4 other tools
measured on Markdown Quality
Mistral AI✓ WorkedThe export is usable markdown rather than a flat text dump: it includes an overall markdown file plus page-wise markdown files delivered in a downloadable ZIP.PDF.ai✗ FailedMarkdown export failed on an 84-page hybrid earnings report; the report records "Output MD: Output failed."PDFVector◐ MixedReturned markdown-like content inside JSON, but the document structure was largely flattened rather than preserved as faithful markdown.Reducto⚠ StruggledProduces usable but not clean markdown: the output contains 12 <signature> tags, 2 <empty> placeholders, and literal HTML tags such as <b>, <i>, and <u> instead of pure CommonMark.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com