Separately extracts chart content into textual summaries and value tables, including a waterfall chart and bar charts with year/value pairs and growth statistics.
What was measured
Advanced Features
Provides separate table/chart extraction and flags low-confidence OCR or ambiguous regions.
context, not decisivetransformation
Separate extraction modes and OCR confidence flags are useful workflow features, but they do not by themselves determine whether the Markdown conversion is good. (3 of 3 judges)
What was given, what came back
Test input: Hybrid Earnings Report · pdf · group: hybrid-earnings-report
Input — what we sent

Hybrid Earnings Report
A long, real-world hybrid annual report used to benchmark end-to-end PDF-to-markdown conversion. It combines native text, financial tables, charts/graphics, and scanned signature/stamp regions, stressing preservation of document structure, reading order, and embedded visual content across a multi-section report.
Why this input is hard
- · Native digital text extraction
- · Complex financial table preservation
- · Chart and graphic retention
- · Image-based signatures and stamps
- · Reading order across a long multi-section report
- · Markdown quality and consistency
Output — unretouched



Also checked on this input — same tool, 4 other criteria
Complex Document Handling◐ MixedHandles an 84-page hybrid report end to end, but quality degrades on signatures and multicolumn pages rather than staying uniform across the document.Reading Order & Structure✗ FailedBreaks the two-column reading order on the annual-report page, so headings and body text no longer follow the source navigation cleanly.Table Preservation◐ MixedReconstructs the 2011–2015 financial summary table with rows and columns intact, but misses a small number of currency symbols in the extracted values.Visual Content Retention✗ FailedDoes not retain handwritten signature content as visual output on the signatures page, leaving only partial surrounding legal text and no clearly identifiable signatures.
Provenance
- Observation
- e260a099-6895-499f-ab5d-82e11a0cf513
- Evidence run
- 6e3160de-fe46-4b45-b071-72560b5c5d0e
- Study
- Convert a Complex PDF into Clean Markdown with an API
- Research task
- 86b9h7t37
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "upstage-ai",
scenario: "hybrid-earnings-report"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 3 other tools
measured on Advanced Features
Landing AI✓ WorkedEmits semantic attestation blocks for handwritten-signature regions and a blurry signature stamp, distinguishing the signer/company name and the signature legibility instead of leaving those regions unannotated.Reducto✓ WorkedThe separate schema-driven extract.run endpoint cleanly recovers both a financial table and a donut chart: all requested table rows and all five donut percentages come back exactly.Tensorlake✓ WorkedProvides separate chart extraction for the SG&A waterfall, outputting chart metadata with a title, axis labels, 10 category labels, and a numeric series instead of only inline prose.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com