FutureSmart Database Agent in AI Database Agents
Scenario-level performance from current published Results.
4 scenarios with published Results · 28 scenarios in the benchmark
How FutureSmart Database Agent performed
Open a capability to explore its scenarios. Each row reports the test set in its published Result; counts are not combined into an overall score.
Conversational Interaction6 scenarios · 1 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| The user changes direction mid-conversation | 0 Pass1 Fail0 Not gradable 1 of 1 test case failedWhat happenedWhen the user changed direction, FutureSmart Database Agent dropped the earlier filter but did not give the unfiltered result the scenario calls for. In the run, it returned a 16-row by-state breakdown instead of the single overall cancellation rate, so the answer did not match the unfiltered result. | 1/1 assessed1/1 gradablePublished test set | View Result → |
| Customer corrects information given earlier | No published result | ||
| The agent offers what to ask next | No published result | ||
| The follow-up is ambiguous | No published result | ||
| The next question refers to the previous answer | No published result | ||
| The user narrows what they just asked | No published result | ||
Question Answering4 scenarios · 1 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| The user's wording doesn't match how the values are stored | 0 Pass1 Fail0 Not gradable 1 of 1 test case failedWhat happenedFutureSmart Database Agent failed to bridge the user's wording to the stored value, and it returned an empty answer as if it were a fact. The generated SQL searched only for gift voucher phrasing and never referenced gift_card, while the answer text said no matching orders existed and suggested broadening the search. The capture cannot show table contents, so the stored values themselves were not visible. | 1/1 assessed1/1 gradablePublished test set | View Result → |
| The answer is available in the connected data | No published result | ||
| The connected data cannot answer the question | No published result | ||
| The question uses a term the company defines itself | No published result | ||
Reporting3 scenarios · 1 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| The report is run again after the data has changed | 0 Pass1 Fail0 Not gradable 1 of 1 test case failedWhat happenedFutureSmart Database Agent did not demonstrate that a report rerun after data changes preserved the same relative window. It returned 57 for the preceding 30 days as of September 2, 2026, and the visible answer surface showed no save/report control or refreshed report to compare after new rows were added. | 1/1 assessed1/1 gradablePublished test set | View Result → |
| The user wants several findings in one place | No published result | ||
| The user wants to keep an answer they just got | No published result | ||
Training4 scenarios · 1 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| The user sets a standing rule for all future answers | 1 Pass0 Fail0 Not gradable 1 of 1 test case passedWhat happenedFutureSmart Database Agent kept the standing exclusion rule across both later answers to unrelated questions. Its generated SQL filtered out cancelled orders in both cases, and the returned results were $51,848.02 for revenue and 32 orders for the count. | 1/1 assessed1/1 gradablePublished test set | View Result → |
| The agent used the wrong definition and the user corrects it | No published result | ||
| The database's names are cryptic and the user documents them | No published result | ||
| The user supplies example questions and the queries they trust | No published result | ||
Analytics & Observability4 scenarios · 0 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| A user reports a bad answer and the operator has to find it | No published result | ||
| Reported analytics match what actually happened | No published result | ||
| Someone needs to see what people have been asking | No published result | ||
| The operator needs to find the questions the agent is failing on | No published result | ||
Answer Presentation4 scenarios · 0 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| The answer is a plain fact or a short list | No published result | ||
| The answer is a trend or a comparison | No published result | ||
| The answer needs a summary and its detail together | No published result | ||
| The user asks to see the same answer a different way | No published result | ||
Data Access Control3 scenarios · 0 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| An excluded column is asked for directly | No published result | ||
| An excluded table is needed to answer the question | No published result | ||
| Two users are given different data scope | No published result | ||
Reading these Results
Published evidence and test coverage answer different questions.
Which scenarios have a Result?
A published Result is public evidence for this tool on one scenario. “No published result” does not say whether testing has taken place.
What does each Result cover?
Assessed includes Pass, Fail and Not gradable. Gradable includes Pass and Fail. Both use the pinned test count in that published Result.
Inventory is not testing progress
The 28 scenarios describe this benchmark’s scope. They are not an assumed applicability or test-coverage denominator for FutureSmart Database Agent.