Tool in benchmark · Version 1

FutureSmart Database Agent in AI Database Agents

Scenario-level performance from current published Results.

4 scenarios with published Results · 28 scenarios in the benchmark

How FutureSmart Database Agent performed

Open a capability to explore its scenarios. Each row reports the test set in its published Result; counts are not combined into an overall score.

Conversational Interaction6 scenarios · 1 with published Results
ScenarioPublished outcomesTest coverageResult
The user changes direction mid-conversation
0 Pass1 Fail0 Not gradable
1 of 1 test case failed
What happened

When the user changed direction, FutureSmart Database Agent dropped the earlier filter but did not give the unfiltered result the scenario calls for. In the run, it returned a 16-row by-state breakdown instead of the single overall cancellation rate, so the answer did not match the unfiltered result.

1/1 assessed1/1 gradablePublished test setView Result →
Customer corrects information given earlierNo published result
The agent offers what to ask nextNo published result
The follow-up is ambiguousNo published result
The next question refers to the previous answerNo published result
The user narrows what they just askedNo published result
Question Answering4 scenarios · 1 with published Results
ScenarioPublished outcomesTest coverageResult
The user's wording doesn't match how the values are stored
0 Pass1 Fail0 Not gradable
1 of 1 test case failed
What happened

FutureSmart Database Agent failed to bridge the user's wording to the stored value, and it returned an empty answer as if it were a fact. The generated SQL searched only for gift voucher phrasing and never referenced gift_card, while the answer text said no matching orders existed and suggested broadening the search. The capture cannot show table contents, so the stored values themselves were not visible.

1/1 assessed1/1 gradablePublished test setView Result →
The answer is available in the connected dataNo published result
The connected data cannot answer the questionNo published result
The question uses a term the company defines itselfNo published result
Reporting3 scenarios · 1 with published Results
ScenarioPublished outcomesTest coverageResult
The report is run again after the data has changed
0 Pass1 Fail0 Not gradable
1 of 1 test case failed
What happened

FutureSmart Database Agent did not demonstrate that a report rerun after data changes preserved the same relative window. It returned 57 for the preceding 30 days as of September 2, 2026, and the visible answer surface showed no save/report control or refreshed report to compare after new rows were added.

1/1 assessed1/1 gradablePublished test setView Result →
The user wants several findings in one placeNo published result
The user wants to keep an answer they just gotNo published result
Training4 scenarios · 1 with published Results
ScenarioPublished outcomesTest coverageResult
The user sets a standing rule for all future answers
1 Pass0 Fail0 Not gradable
1 of 1 test case passed
What happened

FutureSmart Database Agent kept the standing exclusion rule across both later answers to unrelated questions. Its generated SQL filtered out cancelled orders in both cases, and the returned results were $51,848.02 for revenue and 32 orders for the count.

1/1 assessed1/1 gradablePublished test setView Result →
The agent used the wrong definition and the user corrects itNo published result
The database's names are cryptic and the user documents themNo published result
The user supplies example questions and the queries they trustNo published result
Analytics & Observability4 scenarios · 0 with published Results
ScenarioPublished outcomesTest coverageResult
A user reports a bad answer and the operator has to find itNo published result
Reported analytics match what actually happenedNo published result
Someone needs to see what people have been askingNo published result
The operator needs to find the questions the agent is failing onNo published result
Answer Presentation4 scenarios · 0 with published Results
ScenarioPublished outcomesTest coverageResult
The answer is a plain fact or a short listNo published result
The answer is a trend or a comparisonNo published result
The answer needs a summary and its detail togetherNo published result
The user asks to see the same answer a different wayNo published result
Data Access Control3 scenarios · 0 with published Results
ScenarioPublished outcomesTest coverageResult
An excluded column is asked for directlyNo published result
An excluded table is needed to answer the questionNo published result
Two users are given different data scopeNo published result

Reading these Results

Published evidence and test coverage answer different questions.

Publication availability

Which scenarios have a Result?

A published Result is public evidence for this tool on one scenario. “No published result” does not say whether testing has taken place.

Test coverage

What does each Result cover?

Assessed includes Pass, Fail and Not gradable. Gradable includes Pass and Fail. Both use the pinned test count in that published Result.

Scenario scope

Inventory is not testing progress

The 28 scenarios describe this benchmark’s scope. They are not an assumed applicability or test-coverage denominator for FutureSmart Database Agent.

FutureSmart Database Agent in AI Database Agents | AI Demos