Returns accurate product data for the hydrated page, including the correct product name Nike Air Force 1 '07 and a long size list; the visible output shows 14 fully readable size pairs from M 7 / W 8.5 through M 14 / W 15.5.

✓ Worked🧾 artifact-verifiedinput + output shownTested Jun 23, 2026Skyvern
What was measured
Output Quality

Conceptually matches the prompt and is visually usable.

decisive for this rankingtransformation

Whether the result matches the page and is actually usable is the main outcome the ranking is trying to measure. (3 of 3 judges)

What was given, what came back

Test input: Nike Air Force 1 '07 size options extraction · mixed
Input — what we sent
Input, verbatim
https://www.nike.com/t/air-force-1-07-mens-shoes-jBrhbr/CW2288-111 — Wait for the size selection options to fully render. Extract the product name, price, and a list of all available shoe sizes.

A Nike product page with client-side JavaScript hydration used to test whether a headless scraper waits for dynamic DOM content before extracting product details and all available shoe sizes.

Output — unretouched
Completed extraction of Nike Air Force 1 product data
Completed extraction of Nike Air Force 1 product data
Also checked on this input — same tool, 9 other criteria
Interaction Stability⚠ StruggledHandles the page extraction itself, but the run’s live recording pipeline can fall out of sync during hydration: the report states the screen capture froze on the initial page view, making the recording unwatchable and hard to debug.Interaction Stability◐ MixedThe extraction completed, but the screen-capture recorder fell out of sync and froze on an early page state, so runtime observability degraded.Interaction Stability✗ FailedThe live recording pipeline went out of sync during hydration and froze on the initial page view even though backend extraction still completed, showing unreliable handling of dynamic UI state changes.JS DOM Hydration✓ WorkedWaits for the client-rendered product page to hydrate and extracts the size grid instead of stopping at the initial shell.JS DOM Hydration◐ MixedThe extraction pipeline recovered the hydrated Nike size data, but the screen-recording/visual trace subsystem was out of sync and froze on an initial page view, so the captured recording did not reflect the final dynamic state.JS DOM Hydration✓ WorkedThe tool can wait for client-side hydration and capture the rendered product state, including a populated size-selection grid; in this run it extracted a structured schema of the Nike Air Force 1 '07 size options and showed multiple size variants rather than an empty shell.Schema Extraction Integrity✓ WorkedAccurately extracts hydrated product-size data into structured output; the visible result enumerates multiple size pairs, including M 7 / W 8.5, M 7.5 / W 9, M 8 / W 9.5, and through M 15 / W 16.5 in the shown excerpt.Schema Extraction Integrity✓ WorkedOutputs a structured size list with many men/women pairs, with the visible payload spanning at least M 6 / W 7.5 through M 14 / W 15.5.Schema Extraction Integrity✓ WorkedCan accurately preserve a hydrated page’s structured output, extracting a complete schema of all 22 shoe sizes without corrupting the requested JSON structure.
Provenance
Observation
bb5b1f48-5766-4e41-b40f-989875cac6cc
Evidence run
06e1dbd6-5518-4af8-aa1a-735259a75b4f
Study
Scrape Web Pages Into Clean Markdown or Structured Data Using AI
Research task
86b9jm3a3
Tested at
Jun 23, 2026
Source
first-party
Evidence state
verified
Proof shown
input + output shown
Cost / latency
not captured
Repeat run
not captured
Tester
not captured

The last three rows are honest blanks, not placeholders — our capture has no field for them yet.

Query this
get_evidence({
  tool: "skyvern"
})
MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 3 other tools
measured on Output Quality
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com