Can turn numeric results into plain-English takeaways; the delivered-vs-cancelled follow-up automatically summarized the split and framed delivered orders as more common than cancelled orders.
What was measured
Business Insight
Does it explain what the result means?
transformation
What was given, what came back
Test input: Order pipeline breakdown with paid-pending edge case and last-month comparison · text · group: ecommerce-nl2sql-benchmark
Input — what we sent
The exact prompt
How many orders do we have at each stage right now? Follow-up 1: What percentage of our orders were successfully delivered vs cancelled? Follow-up 2: Are there any orders that are pending but already paid? Follow-up 3: Compare that to last month — same breakdown, I want to see if things have improved or got worse.
A deeper operational analysis of current order stages, delivery-versus-cancellation rates, pending-but-paid edge cases, and a month-over-month comparison of the same breakdown.
Why this input is hard
- · Order pipeline analysis
- · Percentage calculation
- · Edge-case detection
- · Payment/order status joins
- · Multi-turn context retention
- · Month-over-month comparison
- · Ambiguity handling for 'same breakdown'
Output — unretouched

Also checked on this input — same tool, 2 other criteria
Ambiguity Handling⚠ StruggledWhen a follow-up phrase could refer either to the full order-stage breakdown or to the immediately preceding pending-paid subresult, it did not ask for clarification and instead chose the narrower pending-paid interpretation.Chart / Visualization Support✓ WorkedCan auto-generate a correct chart for a simple month-over-month comparison, showing Last Month at 11 pending-paid orders versus Current Month at 20.
Provenance
- Observation
- 82eb80c3-8afa-429f-b476-46edd043f78e
- Evidence run
- db2bb5d5-0e0e-4cb3-8d76-3555c45c23cd
- Study
- Query Live Databases Using Plain English with AI
- Research task
- 86b9y6c99
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "draxlr",
scenario: "ecommerce-nl2sql-benchmark"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 3 other tools
measured on Business Insight
AskYourDatabase✓ WorkedThe follow-up summaries explained what the raw counts meant operationally, including delivery/cancellation rates and the backlog reduction story.Definite✓ WorkedIt explained what the numbers meant, including the 14% cancellation rate, the completed-orders caveat, the 2 pending-but-paid orders that need attention, and the warning that May is still early because 17 of 21 orders remain pending.Querio✓ WorkedExplains the comparison rather than just reporting numbers, including that the month-to-date view is an early signal and should not be treated as a stable trend yet.
This evidence is published in
From the same study (page rebuilt from a later run)
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com