Extend AI in Structured Document Extraction
Scenario-level performance from current published Results.
3 scenarios with published Results · 19 scenarios in the benchmark
How Extend AI performed
Open a capability to explore its scenarios. Each row reports the test set in its published Result; counts are not combined into an overall score.
Nested & Repeated Fields4 scenarios · 1 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| The set is legitimately empty | 1 Pass0 Fail0 Not gradable 1 of 1 test case passedWhat happenedExtend AI returned an empty set for the legitimately empty document. The extracted set of damage or exception entries is empty, and the source references are empty too. Page 3 says no damage or exceptions were reported. | 1/1 assessed1/1 gradablePublished test set | View Result → |
| A repeated set of records | No published result | ||
| The set contains a non-record | No published result | ||
| The set continues across a page break | No published result | ||
OCR2 scenarios · 1 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| A document mixing digital and scanned pages | 1 Pass0 Fail0 Not gradable 1 of 1 test case passedWhat happenedExtend AI returned the requested values from a document mixing digital and scanned pages. The decision-relevant finding is that the image-only page was processed rather than skipped: the scanned-page values and the digital-page values both came back correct in one extraction call. | 1/1 assessed1/1 gradablePublished test set | View Result → |
| A cleanly scanned document | No published result | ||
Source Grounding2 scenarios · 1 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| Trace a value to its location | 1 Pass0 Fail0 Not gradable 1 of 1 test case passedWhat happenedExtend AI traced the returned value to the correct document location. It returned 6103 with a citation to page 2 on "Subtotal" beside £6,103.00. The PDF shows "Subtotal" once on page 2 and not on page 1. | 1/1 assessed1/1 gradablePublished test set | View Result → |
| The field is absent | No published result | ||
Field Extraction5 scenarios · 0 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| The field is absent | No published result | ||
| The same field across layouts | No published result | ||
| The value is directly available | No published result | ||
| The value must be derived | No published result | ||
| The value needs a supplied definition | No published result | ||
Querying3 scenarios · 0 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| Aggregate across documents | No published result | ||
| Filter records across documents | No published result | ||
| The answer was never in the schema | No published result | ||
Review and Correction3 scenarios · 0 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| Correct a field value | No published result | ||
| Corrected data goes downstream | No published result | ||
| Repair a record | No published result | ||
Training1 scenario · 0 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| A field the tool can only get right from examples | No published result | ||
Reading these Results
Published evidence and test coverage answer different questions.
Which scenarios have a Result?
A published Result is public evidence for this tool on one scenario. “No published result” does not say whether testing has taken place.
What does each Result cover?
Assessed includes Pass, Fail and Not gradable. Gradable includes Pass and Fail. Both use the pinned test count in that published Result.
Inventory is not testing progress
The 19 scenarios describe this benchmark’s scope. They are not an assumed applicability or test-coverage denominator for Extend AI.