Tool in benchmark · Version 1

BlazeSQL in AI Database Agents

Scenario-level performance from current published Results.

11 scenarios with published Results · 28 scenarios in the benchmark

How BlazeSQL performed

Open a capability to explore its scenarios. Each row reports the test set in its published Result; counts are not combined into an overall score.

Analytics & Observability4 scenarios · 1 with published Results
ScenarioPublished outcomesTest coverageResult
The operator needs to find the questions the agent is failing on
0 Pass1 Fail0 Not gradable
1 of 1 test case failed
What happened

BlazeSQL failed to surface unanswered questions as failures. In the recording, Query review showed one ordinary row marked 'Not verified' and no failures or unanswered view appeared anywhere in the sections shown. The social-media complaints question appeared in Chat with a normal answer, and the other unanswerable question is not shown in the admin view.

1/1 assessed1/1 gradablePublished test setView Result →
A user reports a bad answer and the operator has to find itNo published result
Reported analytics match what actually happenedNo published result
Someone needs to see what people have been askingNo published result
Answer Presentation4 scenarios · 2 with published Results
ScenarioPublished outcomesTest coverageResult
The answer is a plain fact or a short list
1 Pass0 Fail0 Not gradable
1 of 1 test case passed
What happened

BlazeSQL answered the plain fact in prose, so the evidence does not show the result being confined to a chart or a one-cell table. In the captured run, the reply said there were 100 orders ever, and the same frame also showed a one-row result table and no chart.

1/1 assessed1/1 gradablePublished test setView Result →
The user asks to see the same answer a different way
1 Pass0 Fail0 Not gradable
1 of 1 test case passed
What happened

BlazeSQL passed this scenario: it showed the same revenue answer in a different form. In the recorded conversation, the first result appeared as a table with seven categories, then the follow-up request switched that same result to a bar chart. The bar lengths matched the table values within the chart's scale.

1/1 assessed1/1 gradablePublished test setView Result →
The answer is a trend or a comparisonNo published result
The answer needs a summary and its detail togetherNo published result
Conversational Interaction6 scenarios · 3 with published Results
ScenarioPublished outcomesTest coverageResult
Customer corrects information given earlier
1 Pass0 Fail0 Not gradable
1 of 1 test case passed
What happened

BlazeSQL replaced the earlier July answer with the corrected June answer. In the visible exchange, the first turn returned 21 orders for July 2026, and the follow-up turn rebuilt the query for June 2026 and returned 20 orders, matching the corrected period only.

1/1 assessed1/1 gradablePublished test setView Result →
The agent offers what to ask next
1 Pass0 Fail0 Not gradable
1 of 1 test case passed
What happened

BlazeSQL offered what to ask next after answering the revenue question. It ended the August 2026 revenue reply with “Would you like a breakdown by category, or any other analysis?”, so the capture shows one specific follow-up rather than only a generic closing.

1/1 assessed1/1 gradablePublished test setView Result →
The next question refers to the previous answer
1 Pass0 Fail0 Not gradable
1 of 1 test case passed
What happened

BlazeSQL understood the follow-up as referring to the prior answer. In the first turn, it returned the top five customers by spend. In the follow-up, it counted orders for those same five customers, kept them in the same order, and did not ask who they were again. The shown SQL carries that customer list forward in a named CTE.

1/1 assessed1/1 gradablePublished test setView Result →
The follow-up is ambiguousNo published result
The user changes direction mid-conversationNo published result
The user narrows what they just askedNo published result
Data Access Control3 scenarios · 1 with published Results
ScenarioPublished outcomesTest coverageResult
An excluded table is needed to answer the question
0 Pass1 Fail0 Not gradable
1 of 1 test case failed
What happened

BlazeSQL failed this scenario. In the shown run, it gave a confident answer instead of saying it could not answer, even though the needed table was excluded and the SQL used only permitted tables.

1/1 assessed1/1 gradablePublished test setView Result →
An excluded column is asked for directlyNo published result
Two users are given different data scopeNo published result
Question Answering4 scenarios · 2 with published Results
ScenarioPublished outcomesTest coverageResult
The connected data cannot answer the question
1 Pass0 Fail0 Not gradable
1 of 1 test case passed
What happened

BlazeSQL said it could not answer from the available data when asked how many 5-star reviews were received this month. It listed ten available tables, gave no number, and pointed to adding review data or using a related source instead. The key finding is that it declined rather than inventing a count.

1/1 assessed1/1 gradablePublished test setView Result →
The question uses a term the company defines itself
1 Pass0 Fail0 Not gradable
1 of 1 test case passed
What happened

BlazeSQL did not answer revenue as a bare number; it stated the definition it applied and showed the breakdown it used. In the captured Sep 2, 2026 run, it reported order revenue, subscription revenue, refunds, and an adjusted net revenue total, with the order line explicitly marked as excluding cancelled orders.

1/1 assessed1/1 gradablePublished test setView Result →
The answer is available in the connected dataNo published result
The user's wording doesn't match how the values are storedNo published result
Reporting3 scenarios · 1 with published Results
ScenarioPublished outcomesTest coverageResult
The user wants to keep an answer they just got
1 Pass0 Fail0 Not gradable
1 of 1 test case passed
What happened

BlazeSQL saves an answer and its query, then reopens the saved item with the same three rows under the saved name. The saved item appears in a separate Saved area, and the reopened table matches the earlier chat answer.

1/1 assessed1/1 gradablePublished test setView Result →
The report is run again after the data has changedNo published result
The user wants several findings in one placeNo published result
Training4 scenarios · 1 with published Results
ScenarioPublished outcomesTest coverageResult
The database's names are cryptic and the user documents them
1 Pass0 Fail0 Not gradable
1 of 1 test case passed
What happened

BlazeSQL used the user’s documentation to answer from the cryptic table and column. It reported 6 returns currently awaiting manual fraud review, split across 5 approved and 1 received. The stored note was applied to returns.flg_x2, and the follow-up answer came from that documented column.

1/1 assessed1/1 gradablePublished test setView Result →
The agent used the wrong definition and the user corrects itNo published result
The user sets a standing rule for all future answersNo published result
The user supplies example questions and the queries they trustNo published result

Reading these Results

Published evidence and test coverage answer different questions.

Publication availability

Which scenarios have a Result?

A published Result is public evidence for this tool on one scenario. “No published result” does not say whether testing has taken place.

Test coverage

What does each Result cover?

Assessed includes Pass, Fail and Not gradable. Gradable includes Pass and Fail. Both use the pinned test count in that published Result.

Scenario scope

Inventory is not testing progress

The 28 scenarios describe this benchmark’s scope. They are not an assumed applicability or test-coverage denominator for BlazeSQL.

BlazeSQL in AI Database Agents | AI Demos