Returns directly consumable JSON for the invoice workflow, with the demo reaching extracted results and the report stating the JSON is immediately usable for integration.
What was measured
Structural Clean Output
Is the JSON directly consumable by a downstream AI pipeline or system without requiring a structural transformation layer?
context, not decisivetransformation
Being directly consumable by a downstream pipeline is valuable, but it is more about integration convenience than whether the tool actually extracts the data correctly. (3 of 3 judges)
What was given, what came back
Test input: Invoice PDF · pdf · group: business-document-extraction
Input — what we sent
A two-page broadcast advertising invoice PDF with nested header metadata, billing and remit addresses, and eight line items spanning a page break. It was used to stress hierarchical line-item extraction, amount precision, time/day parsing, code extraction, and summary validation.
Why this input is hard
- · Nested line-item hierarchy extraction
- · Multi-page continuity across a page break
- · Precision on large dollar amounts and totals
- · Parsing complex time slots and day patterns
- · Extraction of Ad IDs and reconciliation codes
- · Mapping structured metadata sections correctly
- · Financial summary validation
- · Handling political advertising compliance text
Output — unretouched
Loading file...
Also checked on this input — same tool, 6 other criteria
Extraction Accuracy✗ FailedMisreads at least one alphanumeric identifier: line 1's source Ad-ID NRCCWI071005 is extracted as NRCCW1071005.Extraction Accuracy✓ WorkedExtracts invoice metadata and totals consistently, including invoice_number "4064621-1", invoice_date "10/28/12", aired_spots 8, gross_total 29750, and net_amount_due 25287.5.Extraction Accuracy✗ FailedLeaves several station-level fields null—address, city, state, postal_code, main_phone, and billing_phone—even though the station block is present in the schema and related info exists elsewhere in the invoice.Schema Adherence✓ WorkedReconstructs the invoice hierarchy into invoice_metadata, advertiser, station, billing_address, remit_address, flight_dates, line_items, and summary objects rather than returning flat OCR text.Semantic Field Enrichment✓ WorkedPopulates derived scheduling fields such as day_of_week "M" and days_pattern "MTWT" on invoice line items.Table & Record Completeness✓ WorkedExtracts all 8 advertising line items as separate records, preserving the page-break continuation instead of merging or omitting rows.
Provenance
- Observation
- dbf84eab-a8e4-49e1-b7a5-ddf3319de6bc
- Evidence run
- a061b9e7-a9c5-443d-a171-b296aaf51b8c
- Study
- Extract and query structured data from documents using natural language
- Research task
- 86b9y25e5
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "landing-ai",
scenario: "business-document-extraction"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 3 other tools
measured on Structural Clean Output
Datalab◐ MixedThe invoice JSON is structurally usable but not order-stable: the top-level keys are emitted in a fully shuffled order instead of the requested schema order.Extend AI✓ WorkedThe invoice output includes per-section OCR confidence metadata alongside the extracted fields while remaining machine-readable JSON that can be consumed without structural transformation.Retab⚠ StruggledPreserves valid JSON but does not keep the schema's original key order, so downstream consumers that expect schema-consistent ordering need an extra formatting step.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com