Benchmark methodology

Converting a complex PDF into clean Markdown with a hosted API

The public evaluation method approved for this benchmark definition.

How the benchmark works

A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.

Evaluation unit

One test case, one tool

One test case is scored for one tool at a time.

Outcomes

Pass, fail, or partial

Each run ends pass, fail, or partial. A missing capability scores 0 and stays in the arithmetic — a product must not rank higher by having less product. not_applicable, not_measured and blocked (with the reason, never a 0) are distinct states.

Consistency

Versioned scenarios and resources

Use versioned scenarios and registered resource versions so runs stay comparable.

Evidence standard

Recorded stimulus and complete output

Record the input stimulus, the complete output, and the relevant registered references.

Active test cases are partly withheld to reduce benchmark gaming; the public page still lists the capabilities, scenarios, method, and resource types.

Converting a complex PDF into clean Markdown with a hosted API methodology | AI Demos