Converting a complex PDF into clean Markdown with an open-source library
This benchmark is for self-hostable open-source libraries that convert real-world PDFs into Markdown.
Benchmark overview
What is included and excluded
How well a self-hostable open-source library turns complex PDFs into faithful Markdown that a reader can trust without opening the original.
A reader learns where a library preserves content and structure, and where it drops, reorders, invents, or weakens information. The point is not pretty output; it is whether the converted document still supports search, retrieval, QA, indexing, and human review.
In scope
- Converting complex, real-world PDFs into Markdown with an open-source library.
- Preserving text, reading order, headings, tables, figures, charts, scanned text, equations, and code.
- Judging completeness, accuracy, faithfulness, and honesty against the source PDF.
Out of scope
- Structured field extraction.
- Web-page conversion.
- Form filling.
- Handwriting.
- Document classification or routing.
- PDF creation or editing.
Participating tools
Tools in this benchmark’s public roster. Publication availability is not a performance ranking.
| Tool | Published Results | Explore |
|---|---|---|
| doc2mark | No published result | View tool in this benchmark → |
| Docling | No published result | View tool in this benchmark → |
| GPT-5.6 terra | No published result | View tool in this benchmark → |
| LiteParse | No published result | View tool in this benchmark → |
| Marker | No published result | View tool in this benchmark → |
| MarkItDown | No published result | View tool in this benchmark → |
| MinerU | No published result | View tool in this benchmark → |
| olmOCR | No published result | View tool in this benchmark → |
| PaddleOCR-VL | No published result | View tool in this benchmark → |
| PyMuPDF4LLM | No published result | View tool in this benchmark → |
Capabilities & scenarios
12 scenarios grouped by 8 capabilities. Open a group to explore its scenarios in this benchmark.
Text Fidelity1 scenario
Whether ordinary digital text comes through complete, unchanged, and not invented in the Markdown output.
Capability in this benchmark → · Global definition →
- An ordinary digital text documentNo published results
Reading Order & Layout1 scenario
Whether the source page is read in the natural sequence, so text is not interleaved or reordered by layout.
Capability in this benchmark → · Global definition →
- A page laid out in multiple columnsNo published results
Heading & Section Structure2 scenarios
Whether headings and subheadings become real Markdown structure, with section hierarchy intact and footnotes kept attached outside the body flow.
Capability in this benchmark → · Global definition →
- A document with styled headings and subheadingsNo published results
- Footnotes at the bottom of the pageNo published results
Table Extraction2 scenarios
Whether tables keep their rows, columns, headers, and cell values as a table.
Capability in this benchmark → · Global definition →
- A simple, clearly formatted tableNo published results
- A table that continues across a page breakNo published results
Figures & Charts2 scenarios
Whether figures and charts remain placed, referenced, and captioned or titled, with a text trace that still points back to the source.
Capability in this benchmark → · Global definition →
- A document with figures and captionsNo published results
- A document with a data chartNo published results
Scanned Document OCR2 scenarios
Whether image-only pages are turned into accurate text instead of being skipped or left unread.
Capability in this benchmark → · Global definition →
- A cleanly scanned documentNo published results
- A document mixing digital and scanned pagesNo published results
Equations & Mathematical Notation1 scenario
Whether equations survive as math markup or a clear fallback instead of garbled prose.
Capability in this benchmark → · Global definition →
- A document containing mathematical equationsNo published results
Code Extraction1 scenario
Whether code survives as a fenced block with line breaks and indentation intact.
Capability in this benchmark → · Global definition →
- A document containing a code blockNo published results
Results overview
Current published evidence in this benchmark.
No published results for this benchmark yet.
Publication availability is separate from test coverage and unpublished research progress.
How the benchmark works
A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.
One test case for one tool
Each verdict applies to one test case run against one tool.
Pass, fail, or not gradable
A result is pass, fail, or not gradable; not gradable is not a failure.
Versioned scenarios and registered resources
Use the registered scenario versions and the registered resource versions for every run.
Recorded stimulus, complete output, relevant references
Keep the recorded stimulus, the complete output, and the relevant registered references with the verdict.
Resources and fixtures
The registered material and systems that create a consistent test environment for this benchmark.
PDF→Markdown fixture corpus — round 1
A shared pinned PDF corpus used as input for both PDF-to-Markdown benchmark versions.