Keeps the conversation going across multiple turns, but on the final comparison it reuses the delivered-versus-cancelled thread instead of the pending-paid edge-case context, so context selection breaks on one follow-up.

⚠ Struggledinput onlyTest date not recordedQuerio
What was measured
Follow-Up Context

Does the tool remember previous answers correctly?

transformation

What was given, what came back

Test input: Order pipeline breakdown with paid-pending edge case and last-month comparison · text · group: ecommerce-nl2sql-benchmark
Input — what we sent
The exact prompt
How many orders do we have at each stage right now?

Follow-up 1: What percentage of our orders were successfully delivered vs cancelled?

Follow-up 2: Are there any orders that are pending but already paid?

Follow-up 3: Compare that to last month — same breakdown, I want to see if things have improved or got worse.

A deeper operational analysis of current order stages, delivery-versus-cancellation rates, pending-but-paid edge cases, and a month-over-month comparison of the same breakdown.

Why this input is hard
  • · Order pipeline analysis
  • · Percentage calculation
  • · Edge-case detection
  • · Payment/order status joins
  • · Multi-turn context retention
  • · Month-over-month comparison
  • · Ambiguity handling for 'same breakdown'
Output — unretouched
No output artifact
The verdict rests on the tester's written observation alone — no file was captured for this cell.
Provenance
Observation
36b9f3cc-d31a-4e7d-a424-b7ec650d518e
Evidence run
db2bb5d5-0e0e-4cb3-8d76-3555c45c23cd
Study
Query Live Databases Using Plain English with AI
Research task
86b9y6c99
Tested at
not recorded
Source
first-party
Evidence state
observed
Proof shown
input only
Cost / latency
not captured
Repeat run
not captured
Tester
not captured

The last three rows are honest blanks, not placeholders — our capture has no field for them yet.

Query this
get_evidence({
  tool: "querio",
  scenario: "ecommerce-nl2sql-benchmark"
})
MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 2 other tools
measured on Follow-Up Context
From the same study (page rebuilt from a later run)
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com