AI Demos Research — the structured-intelligence platform. Every verdict on these pages opens to the execution behind it.
Graded 28 September 2026

Can Adobe PDF Extract API preserve a code block's line structure and indentation intact?

Adobe PDF Extract API did not preserve the code block as a preformatted or fenced block. The nine-line Python block was collapsed into five ordinary paragraphs, the shell block into two lines, and no output line carried leading whitespace. The whitespace-insensitive character content still matched the source, but the block structure was lost.

1 of 1 test case failed

Every test case in this scenario has a result.

Pass rate0%0 of 1 with a result
Coverage1 of 1test cases with a result
0PassNo test case passed.
1FailMissed at least one thing it was expected to do; the reason and proof are on the row.
0Not gradableEvery result here could be graded.
0UntestedEvery test case in this scenario has a result.
The pass rate is a summary. The evidence is the test case below: what we sent, what we checked, what the tool returned, and the proof.

The test case

Each test case is judged on its own: Pass, Fail, Not gradable, or Untested. The scenario result above counts this row.

Code block — fenced, structure intactFailEvidence
What we sent
InputDocument (PDF)
Source PDF with a nine-line Python block and a shell block. From outside the graded run; shown for context, not graded.
Open the PDF ↗
Key momentsp. 2 ↗
What the tool returned
Markdown51 lines
Tarnbeck Institute of Hydrology · Data Services

# Reloading a station export after a gauge reset

Operations note DS/OP/07 · Data services · Reviewed 14 January 2026

# Symptoms

The station's daily total is short and the loader's own log says nothing is wrong. In the log the day looks like a clean run: the export was fetched, every block was read, and the summary line reports fewer readings than the station's counter did. The gap is always at the end of the day rather than the beginning, and it is always a whole number of blocks.

The other sign is a duplicate warning in the block log, with two entries carrying the same sequence number and different checksums. One warning is normal after a reset. A run of them means the reset happened part way through the day and everything after it is being rejected block by block.

# Why the loader rejects the second block

Block sequence numbers come from the station and not from the loader. The loader keeps the highest number it has committed for the day and treats anything at or below it as a repeat, which is the right thing to do on a poor link: a station that retransmits sends the same block twice, and the loader must not count it twice.

A reset returns the counter to zero, so the blocks after it carry sequence numbers the loader has already committed. Nothing about them looks new and the loader discards them for exactly the reason it should. The reset time is the only thing that distinguishes a continuation from a repeat, and the loader does not know it unless it is told.

# The fix

When a gauge is reset in the field the station's counter returns to zero, and the export for that day contains two blocks carrying the same sequence number. The loader accepts the first and rejects the second. That is the right behaviour for a duplicated transmission and the wrong behaviour here, and the day's record ends up short by however much arrived after the reset.

The fix is to requeue the affected blocks with the reset time supplied, so that the loader treats the second block as a continuation rather than a repeat. The function below is the one to call. It lives in tarnbeck.stations.maintenance and is safe to run more than once: a block that already carries a good checksum is skipped.

from tarnbeck.stations import Export, reload_window

def reload_after_reset(station, reset_at): export = Export.open(station) window = reload_window(reset_at, margin_minutes=45)

for block in export.blocks(window): if block.checksum_ok(): continue

export.requeue(block, reason="gauge reset")

return export.commit(dry_run=False)

Run it from the loader host and not from a workstation, and take the reset time from the field log rather than from the station clock, which is the thing that was wrong. The margin defaults to forty-five minutes either side of the reset; widen it only if the field log is vague about the time, because a wide window requeues blocks that were never in doubt.

If commit returns a count lower than the number of blocks in the window, stop there and raise it with data services before running anything again. A short count means blocks were rejected for a second reason, and requeuing them repeatedly will not find out what it was.

# Running it

Run it from the loader host, in the maintenance environment, with the station code and the reset time as they appear in the field log:

$ ssh loader-01.tarnbeck.example

$ tarnbeck-maint reload-after-reset --station VQ1 \ --reset-at '2026-01-29T09:15' --confirm

The command prints the window it will use, the number of blocks in it and the number it intends to requeue, and then waits. Read the three numbers before answering. If the window is wider than the field log justifies, stop and narrow it rather than letting it run.

Without --confirm the command does everything except commit, which is the safe way to see what it would do. The dry run takes a few seconds and there is no reason to skip it.

# Afterwards

Copied from Proof 1 · Output file (Markdown), lines 1–51 · an excerpt; the link below opens the whole file

Open the Markdown file ↗
Proof 1Output file (Markdown)
Markdown output built from PDF extraction structured data. From outside the graded run; shown for context, not graded.
Part of this file is printed above, under What the tool returned. Show the whole file
Markdown79 lines
Tarnbeck Institute of Hydrology · Data Services

# Reloading a station export after a gauge reset

Operations note DS/OP/07 · Data services · Reviewed 14 January 2026

# Symptoms

The station's daily total is short and the loader's own log says nothing is wrong. In the log the day looks like a clean run: the export was fetched, every block was read, and the summary line reports fewer readings than the station's counter did. The gap is always at the end of the day rather than the beginning, and it is always a whole number of blocks.

The other sign is a duplicate warning in the block log, with two entries carrying the same sequence number and different checksums. One warning is normal after a reset. A run of them means the reset happened part way through the day and everything after it is being rejected block by block.

# Why the loader rejects the second block

Block sequence numbers come from the station and not from the loader. The loader keeps the highest number it has committed for the day and treats anything at or below it as a repeat, which is the right thing to do on a poor link: a station that retransmits sends the same block twice, and the loader must not count it twice.

A reset returns the counter to zero, so the blocks after it carry sequence numbers the loader has already committed. Nothing about them looks new and the loader discards them for exactly the reason it should. The reset time is the only thing that distinguishes a continuation from a repeat, and the loader does not know it unless it is told.

# The fix

When a gauge is reset in the field the station's counter returns to zero, and the export for that day contains two blocks carrying the same sequence number. The loader accepts the first and rejects the second. That is the right behaviour for a duplicated transmission and the wrong behaviour here, and the day's record ends up short by however much arrived after the reset.

The fix is to requeue the affected blocks with the reset time supplied, so that the loader treats the second block as a continuation rather than a repeat. The function below is the one to call. It lives in tarnbeck.stations.maintenance and is safe to run more than once: a block that already carries a good checksum is skipped.

from tarnbeck.stations import Export, reload_window

def reload_after_reset(station, reset_at): export = Export.open(station) window = reload_window(reset_at, margin_minutes=45)

for block in export.blocks(window): if block.checksum_ok(): continue

export.requeue(block, reason="gauge reset")

return export.commit(dry_run=False)

Run it from the loader host and not from a workstation, and take the reset time from the field log rather than from the station clock, which is the thing that was wrong. The margin defaults to forty-five minutes either side of the reset; widen it only if the field log is vague about the time, because a wide window requeues blocks that were never in doubt.

If commit returns a count lower than the number of blocks in the window, stop there and raise it with data services before running anything again. A short count means blocks were rejected for a second reason, and requeuing them repeatedly will not find out what it was.

# Running it

Run it from the loader host, in the maintenance environment, with the station code and the reset time as they appear in the field log:

$ ssh loader-01.tarnbeck.example

$ tarnbeck-maint reload-after-reset --station VQ1 \ --reset-at '2026-01-29T09:15' --confirm

The command prints the window it will use, the number of blocks in it and the number it intends to requeue, and then waits. Read the three numbers before answering. If the window is wider than the field log justifies, stop and narrow it rather than letting it run.

Without --confirm the command does everything except commit, which is the safe way to see what it would do. The dry run takes a few seconds and there is no reason to skip it.

# Afterwards

Check the day's totals against the gauge's own counter before closing the ticket. The loader reports the number of blocks committed and not the number of readings, and a day can come back with the right number of blocks and the wrong number of readings if a block was truncated in the field.

Note the reset in the station log, with the time taken from the field log and not the time the reload ran. The next person to look at a gap in the series will look there first, and a reload with nothing written beside it is indistinguishable from a gap nobody has dealt with.

# If the count is short

A short count means blocks were rejected for a second reason and the reload has not found out what it was. Do not run it again: a second pass requeues the same blocks, reports the same count, and fills the log with attempts that bury the original failure.

Take the sequence numbers the command reports as rejected and look them up in the block log. A checksum failure on a block truncated in the field cannot be fixed by reloading and needs the station visited. A rejection with no reason recorded is a loader fault, and goes to data services with the run's log attached.

# The block header, for reference

Each block in an export begins with a fixed header, which is what the loader reads to decide whether it has seen the block before:

SEQ=00184 STATION=VQ1 START=2026-01-29T08:45Z

COUNT=0180 INTERVAL=15s CRC=8f2a41d6

SEQ is the station's counter and is the field a reset returns to zero. START is the station clock, which is not to be trusted after a reset - which is why the reset time comes from the field log and not from the export.

CRC covers the readings and not the header, so a block with a good checksum and a wrong sequence number is possible and is exactly what a reset produces. Nothing in the header says a reset happened.

# Rolling back

A committed reload can be undone within the same day by reverting the station's day to the snapshot the loader takes before it commits. The snapshot is kept for seven days and is named for the run.

After seven days the only route back is a fresh fetch of the export from the station, which is possible while the station's own buffer still holds the day - about three weeks - and impossible afterwards. A reload that has committed and cannot be rolled back is an incident and is raised as one, whatever the size of the gap.
Open the Markdown file ↗
Expected vs. Found
What this scenario evaluates

These are scenario-level criteria. Each test case's Expected and Found are listed separately.

  • Whether the code remains a distinct preformatted or fenced block.
  • Whether line breaks and indentation are preserved.
  • Whether the code is not reflowed into ordinary prose.
  • Whether the code block is still recognisable as code rather than flattened text.
How this scenario is judged →
✗Found: Code blocks were flattened into paragraphs.
Why this result

The conversion flattened code blocks into ordinary paragraphs, so the required preformatted structure was not preserved.

Also observed on this row

Characters preserved. The output preserved all source characters, with no additions or deletions in the text itself.

Page furniture stripped. Page numbers did not appear in the output, and the running header stayed as body text.

Tested by Chandresh Bisht · evidence dated 10 September 2026

Configuration and setup

How this tool was set up for the run and what the test needed in place. Each row is a fact from the run's records; a fact the records do not hold is left out, not guessed.

Software that produced the output
Adobe PDF Extract API
Build or version
PDF Services REST (Extract) json_export=233, page_segmentation=55, schema=1.1.0, structure=1.1218.0, table_structure=5
Plan or tier
free tier
Model
PDF Services REST (Extract) json_export=233, page_segmentation=55, schema=1.1.0, structure=1.1218.0, table_structure=5
Surface
REST API
Set up before the run
A PDF containing a code block was provided, and the complete PDF was converted to Markdown with the product's default pipeline in one pass, never as a cropped page.
Tested
10 September 2026 · Chandresh Bisht

How this scenario is graded

How we decide Pass, Fail and Not gradable. The same rules apply to every tool tested on this scenario.

How results are decided

Each test case gets one result per tool: Pass, Fail or Not gradable. A test case we haven't run yet shows Untested. There are no partial results.

The rules
  • Pass: the tool did everything the test expected, and nothing it said contradicts the correct answer.
  • Fail: at least one expected behaviour clearly didn't happen; the row says which and quotes the tool.
  • Not gradable: our evidence couldn't settle the outcome (for example a record we needed is missing). It is never counted as a fail, and the row says what's missing.

Where this sits in the benchmark

This page is one cell of a larger study: one tool, one scenario. Only this benchmark's frame appears here.

History of this result

What has happened to this result since it was first published. Runs and grades are never overwritten: a retest or a re-grade publishes a new result and keeps the earlier one readable.

from the publication record
7 October 2026First publishedConverting a complex PDF into clean Markdown with a hosted API v1

Act on this result

Nothing filed here edits the run or the grade. A challenge opens a review, and a review can produce a new run or a re-grade — which becomes the current result and leaves this one in the history.

This matches what I see

You run the same kind of test against your own setup and get the same behaviour.

Agree →
This does not match

Yours behaves differently. Tell us what you got, with a screenshot if you have one.

Disagree →
Point out an issue

Something here is wrong — a reference value, a transcription, a grade.

Report an issue →
Request a retest

On a newer build, a larger dataset, or your own setup.

Request a retest →
We have fixed this

Tell us what changed and we schedule a rerun of the failing test case. The old result stays as history.

Vendor notice →
Filed against this evidenceNothing yet. Challenges, counter-evidence and fix notices appear here with their outcome, and stay on the page after they are resolved.
Cite this result
aidemos.com/benchmarks/pdf-to-markdown-apis/results/adobe-pdf-extract-api/a-document-containing-a-code-block · 0 pass · 1 fail · coverage 1/1 · graded 2026-09-28

The same record is available as structured data through the AI Demos MCP server, with the counts, the coverage and every per-test-case reason carried as fields.

Verify the proof files

These files support this result. Open a file to inspect the original evidence.

Input · Document (PDF) ↗application/pdf · 78 KB
File fingerprints (SHA-256)

A fingerprint identifies the exact file used for this result.

Proof 1 · Output file (Markdown)c0819b381d032ff9a035b9382213817eda6911bedf0d7a5d4023a16f22e14157
Input · Document (PDF)cd27d1e77ec9818ff769bde1b49d95e711373888507af4b007054c27b7a52db8