Carries the top-3 customer identities forward across follow-ups and regenerates fresh SQL on each turn, preserving context correctly across at least two drill-down questions.
What was measured
Follow-Up Context
Does the tool remember previous answers correctly?
transformation
What was given, what came back
Test input: Best customers with unpaid-order and payment-method follow-ups · text · group: ecommerce-nl2sql-benchmark
Input — what we sent
The exact prompt
Who are my best customers — the ones who order the most and spend the most? Follow-up 1: For the top 3 from that list — do any of them have unpaid orders? Follow-up 2: What payment methods do these top 3 usually use?
A conversational multi-table customer analysis that identifies best customers by both order volume and spend, then drills into unpaid orders and payment methods for the top 3.
Why this input is hard
- · Ambiguous business-term interpretation
- · Multi-table joins across customers orders and payments
- · Aggregation and ranking
- · Follow-up context retention
- · Scoped drill-down to the top 3 customers
- · Readable customer-level output
Output — unretouched



Also checked on this input — same tool, 5 other criteria
Ambiguity Handling✓ WorkedWhen the phrase "best customers" was ambiguous, it did not guess a single meaning; it split the request into separate order-count and spend rankings.Chart / Visualization Support✓ WorkedSupports visualization on the customer-analysis thread by generating bar-chart views for both spend ranking and order-count ranking.Plain English Query Handling✓ WorkedHandles an informal conversational request about top customers without needing the user to translate it into SQL.Result Readability◐ MixedThe tables are structurally clear, but using customer UUIDs instead of customer names makes the best-customer results less immediately readable for nontechnical users.SQL Visibility✓ WorkedShows the generated SQL inline so users can inspect the aggregation and ordering logic, including the GROUP BY and ORDER BY clauses.
Provenance
- Observation
- f17a5005-134e-4912-8a04-b8fcd44f94dd
- Evidence run
- db2bb5d5-0e0e-4cb3-8d76-3555c45c23cd
- Study
- Query Live Databases Using Plain English with AI
- Research task
- 86b9y6c99
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "querio",
scenario: "ecommerce-nl2sql-benchmark"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 3 other tools
measured on Follow-Up Context
AskYourDatabase✓ WorkedKept the top-3 customer context across the follow-up chain by hardcoding the same three customer IDs into the unpaid-order check.Basedash◐ MixedIt retained the earlier ranking context only partially: the unpaid-orders follow-up checked the top 3 highest spenders first, then had to run a separate pass for the top 3 by order count instead of carrying one unambiguous 'top 3' thread forward.Definite✓ WorkedIt retained context across follow-ups, allowing the user to ask about unpaid orders and then payment methods for the top 3 from the earlier spend-ranked list.
This evidence is published in
From the same study (page rebuilt from a later run)
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com