Benchmark · Version 1

Converting a complex PDF into clean Markdown with an open-source library

This benchmark is for self-hostable open-source libraries that convert real-world PDFs into Markdown.

Benchmark overview

8Capabilities
12Scenarios
10Participating tools
What is included and excluded

How well a self-hostable open-source library turns complex PDFs into faithful Markdown that a reader can trust without opening the original.

A reader learns where a library preserves content and structure, and where it drops, reorders, invents, or weakens information. The point is not pretty output; it is whether the converted document still supports search, retrieval, QA, indexing, and human review.

In scope

  • Converting complex, real-world PDFs into Markdown with an open-source library.
  • Preserving text, reading order, headings, tables, figures, charts, scanned text, equations, and code.
  • Judging completeness, accuracy, faithfulness, and honesty against the source PDF.

Out of scope

  • Structured field extraction.
  • Web-page conversion.
  • Form filling.
  • Handwriting.
  • Document classification or routing.
  • PDF creation or editing.

Participating tools

Tools in this benchmark’s public roster. Publication availability is not a performance ranking.

ToolPublished ResultsExplore
doc2markNo published resultView tool in this benchmark →
DoclingNo published resultView tool in this benchmark →
GPT-5.6 terraNo published resultView tool in this benchmark →
LiteParseNo published resultView tool in this benchmark →
MarkerNo published resultView tool in this benchmark →
MarkItDownNo published resultView tool in this benchmark →
MinerUNo published resultView tool in this benchmark →
olmOCRNo published resultView tool in this benchmark →
PaddleOCR-VLNo published resultView tool in this benchmark →
PyMuPDF4LLMNo published resultView tool in this benchmark →

Capabilities & scenarios

12 scenarios grouped by 8 capabilities. Open a group to explore its scenarios in this benchmark.

Text Fidelity1 scenario

Whether ordinary digital text comes through complete, unchanged, and not invented in the Markdown output.

Capability in this benchmark → · Global definition →

  1. An ordinary digital text documentNo published results
Reading Order & Layout1 scenario

Whether the source page is read in the natural sequence, so text is not interleaved or reordered by layout.

Capability in this benchmark → · Global definition →

  1. A page laid out in multiple columnsNo published results
Heading & Section Structure2 scenarios

Whether headings and subheadings become real Markdown structure, with section hierarchy intact and footnotes kept attached outside the body flow.

Capability in this benchmark → · Global definition →

  1. A document with styled headings and subheadingsNo published results
  2. Footnotes at the bottom of the pageNo published results
Table Extraction2 scenarios

Whether tables keep their rows, columns, headers, and cell values as a table.

Capability in this benchmark → · Global definition →

  1. A simple, clearly formatted tableNo published results
  2. A table that continues across a page breakNo published results
Figures & Charts2 scenarios

Whether figures and charts remain placed, referenced, and captioned or titled, with a text trace that still points back to the source.

Capability in this benchmark → · Global definition →

  1. A document with figures and captionsNo published results
  2. A document with a data chartNo published results
Scanned Document OCR2 scenarios

Whether image-only pages are turned into accurate text instead of being skipped or left unread.

Capability in this benchmark → · Global definition →

  1. A cleanly scanned documentNo published results
  2. A document mixing digital and scanned pagesNo published results
Equations & Mathematical Notation1 scenario

Whether equations survive as math markup or a clear fallback instead of garbled prose.

Capability in this benchmark → · Global definition →

  1. A document containing mathematical equationsNo published results
Code Extraction1 scenario

Whether code survives as a fenced block with line breaks and indentation intact.

Capability in this benchmark → · Global definition →

  1. A document containing a code blockNo published results

Results overview

Current published evidence in this benchmark.

0Current published Results
0Scenarios with published Results
0Tools with published Results

No published results for this benchmark yet.

Publication availability is separate from test coverage and unpublished research progress.

Explore scenarios →

How the benchmark works

A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.

Evaluation unit

One test case for one tool

Each verdict applies to one test case run against one tool.

Outcomes

Pass, fail, or not gradable

A result is pass, fail, or not gradable; not gradable is not a failure.

Consistency

Versioned scenarios and registered resources

Use the registered scenario versions and the registered resource versions for every run.

Evidence standard

Recorded stimulus, complete output, relevant references

Keep the recorded stimulus, the complete output, and the relevant registered references with the verdict.

Full benchmark methodology →

Resources and fixtures

The registered material and systems that create a consistent test environment for this benchmark.

Converting a complex PDF into clean Markdown with an open-source library — Benchmark definition | AI Demos