The customer ranking output was easy to scan because it separated frequent shoppers from highest spenders into clearly named tables with totals and ranks.
What was measured
Result Readability
Is the answer easy for a non-technical user to understand?
transformation
What was given, what came back
Test input: Best customers with unpaid-order and payment-method follow-ups · text · group: ecommerce-nl2sql-benchmark
Input — what we sent
The exact prompt
Who are my best customers — the ones who order the most and spend the most? Follow-up 1: For the top 3 from that list — do any of them have unpaid orders? Follow-up 2: What payment methods do these top 3 usually use?
A conversational multi-table customer analysis that identifies best customers by both order volume and spend, then drills into unpaid orders and payment methods for the top 3.
Why this input is hard
- · Ambiguous business-term interpretation
- · Multi-table joins across customers orders and payments
- · Aggregation and ranking
- · Follow-up context retention
- · Scoped drill-down to the top 3 customers
- · Readable customer-level output
Output — unretouched

Also checked on this input — same tool, 9 other criteria
Ambiguity Handling✓ WorkedHandled the ambiguous phrase by not collapsing it into one ranking; it produced separate order-count and total-spend rankings instead of guessing silently.Business Insight✓ WorkedGenerated useful takeaways automatically, including Rahul Sharma as the standout all-rounder, Mohan Vishe as the most loyal frequent buyer, and Vikram Singh as the largest spender.Business Insight✓ WorkedThe follow-up reasoning went beyond listing payments by flagging Mohan Vishe's unpaid shipped orders as a process risk and noting that Rahul Sharma's pattern was largely UPI-driven.Chart / Visualization Support✗ FailedVisualization did not auto-generate on this customer analysis flow; the report says it required an additional prompt every time.Follow-Up Context✓ WorkedThe tool preserved conversational context across turns by reusing the previously identified top three customers in the unpaid-order follow-up.Follow-Up Context✓ WorkedKept the top-3 customer context across the follow-up chain by hardcoding the same three customer IDs into the unpaid-order check.Plain English Query Handling✓ WorkedThe tool understood follow-up phrasing naturally, including 'For the top 3 from that list' and 'What payment methods do these top 3 usually use?'.Plain English Query Handling✓ WorkedAccepted an informal customer-analytics question and split it into two dimensions without requiring SQL: who orders most and who spends most.SQL Generation✓ WorkedRan 2 SQL queries simultaneously for the main request, and then generated a filtered follow-up query for the exact top 3 customers.
Provenance
- Observation
- 05dbced4-de87-4e7a-a1bf-84da4f7a6be8
- Evidence run
- db2bb5d5-0e0e-4cb3-8d76-3555c45c23cd
- Study
- Query Live Databases Using Plain English with AI
- Research task
- 86b9y6c99
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "askyourdatabase",
scenario: "ecommerce-nl2sql-benchmark"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 2 other tools
measured on Result Readability
Definite✓ WorkedThe ranked customer table was readable and easy to scan, with rank, orders, total spend, and average order value clearly laid out for the top 20 customers.Querio◐ MixedThe tables are structurally clear, but using customer UUIDs instead of customer names makes the best-customer results less immediately readable for nontechnical users.
This evidence is published in
From the same study (page rebuilt from a later run)
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com