Benchmark methodology

Converting a complex PDF into clean Markdown with an open-source library

The public evaluation method approved for this benchmark definition.

How the benchmark works

A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.

Evaluation unit

One test case for one tool

Each verdict applies to one test case run against one tool.

Outcomes

Pass, fail, or not gradable

A result is pass, fail, or not gradable; not gradable is not a failure.

Consistency

Versioned scenarios and registered resources

Use the registered scenario versions and the registered resource versions for every run.

Evidence standard

Recorded stimulus, complete output, relevant references

Keep the recorded stimulus, the complete output, and the relevant registered references with the verdict.

Active test cases are partly withheld to reduce benchmark gaming; capabilities, scenarios, method, and resource types stay public.

Converting a complex PDF into clean Markdown with an open-source library methodology | AI Demos