Benchmark · Version 1

AI Database Agents

This benchmark looks at database agents for business users who want answers from connected relational data without writing SQL.

Benchmark overview

7Capabilities
28Scenarios
11Participating tools
What is included and excluded

AI Database Agents evaluates whether a person who cannot or does not want to write SQL can ask a live relational database for business answers correctly, in the company’s own terms, and honestly when the data cannot answer.

Readers learn how these products behave across question answering, conversation, training, answer presentation, reporting, analytics and observability, and data access control. It is a buyer-facing comparison of behaviour in live database use, not a check on SQL generation alone.

In scope

  • asking questions of connected database data conversationally
  • how the answer is presented
  • whether it can be turned into a report
  • whether the agent can be trained
  • whether configured data restrictions hold

Out of scope

  • database administration
  • migrations
  • index tuning
  • arbitrary writes
  • DBA automation
  • debugging SQL a user already wrote
  • data-engineering pipelines
  • authoring BI dashboards from scratch
  • question-answering over documents or spreadsheets

Participating tools

Tools in this benchmark’s public roster. Publication availability is not a performance ranking.

ToolPublished ResultsExplore
AskYourDatabase1 scenario with a published ResultView tool in this benchmark →
BlazeSQL11 scenarios with published ResultsView tool in this benchmark →
FutureSmart Database Agent4 scenarios with published ResultsView tool in this benchmark →
AI for DatabaseNo published resultView tool in this benchmark →
Anomaly AINo published resultView tool in this benchmark →
BasedashNo published resultView tool in this benchmark →
camelAINo published resultView tool in this benchmark →
DefiniteNo published resultView tool in this benchmark →
DotNo published resultView tool in this benchmark →
DraxlrNo published resultView tool in this benchmark →
QuerioNo published resultView tool in this benchmark →

Capabilities & scenarios

28 scenarios grouped by 7 capabilities. Open a group to explore its scenarios in this benchmark.

Question Answering4 scenarios

Answers a question about connected data correctly, using the company’s meaning and the values stored in the database, and says when the data cannot answer.

Capability in this benchmark → · Global definition →

  1. The answer is available in the connected dataNo published results
  2. The question uses a term the company defines itself1 tool with a published Result
  3. The user's wording doesn't match how the values are stored1 tool with a published Result
  4. The connected data cannot answer the question1 tool with a published Result
Conversational Interaction6 scenarios

Keeps a working session coherent across turns, so context carries forward, follow-ups can narrow or change direction, ambiguity can be raised, and earlier information can be corrected.

Capability in this benchmark → · Global definition →

  1. Customer corrects information given earlier1 tool with a published Result
  2. The next question refers to the previous answer1 tool with a published Result
  3. The user narrows what they just asked1 tool with a published Result
  4. The user changes direction mid-conversation1 tool with a published Result
  5. The follow-up is ambiguousNo published results
  6. The agent offers what to ask next1 tool with a published Result
Training4 scenarios

Uses company-specific information to improve later behaviour and to generalise beyond the exact example given.

Capability in this benchmark → · Global definition →

  1. The agent used the wrong definition and the user corrects itNo published results
  2. The database's names are cryptic and the user documents them1 tool with a published Result
  3. The user supplies example questions and the queries they trustNo published results
  4. The user sets a standing rule for all future answers1 tool with a published Result
Answer Presentation4 scenarios

Chooses and changes the form of the answer, such as prose, a table, a chart, or a combination, when that better fits the question.

Capability in this benchmark → · Global definition →

  1. The answer is a plain fact or a short list1 tool with a published Result
  2. The answer is a trend or a comparisonNo published results
  3. The answer needs a summary and its detail togetherNo published results
  4. The user asks to see the same answer a different way1 tool with a published Result
Reporting3 scenarios

Turns analysis into a durable artifact that can be reopened or run again later without changing its meaning.

Capability in this benchmark → · Global definition →

  1. The user wants to keep an answer they just got1 tool with a published Result
  2. The report is run again after the data has changed1 tool with a published Result
  3. The user wants several findings in one placeNo published results
Analytics & Observability4 scenarios

Lets the buyer inspect what people asked and what happened, and checks that reported numbers match reality.

Capability in this benchmark → · Global definition →

  1. Reported analytics match what actually happenedNo published results
  2. Someone needs to see what people have been askingNo published results
  3. A user reports a bad answer and the operator has to find itNo published results
  4. The operator needs to find the questions the agent is failing on1 tool with a published Result
Data Access Control3 scenarios

Respects configured restrictions on which tables, columns, and rows the agent may use when answering.

Capability in this benchmark → · Global definition →

  1. An excluded table is needed to answer the question1 tool with a published Result
  2. An excluded column is asked for directlyNo published results
  3. Two users are given different data scopeNo published results

Results overview

Current published evidence in this benchmark.

16Current published Results
16Scenarios with published Results
3Tools with published Results

Publication availability is separate from test coverage and unpublished research progress.

Explore scenarios →

How the benchmark works

A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.

Graded artefact

Answer, not SQL

The benchmark grades the answer, never the SQL; SQL is captured only as instrumentation.

Evaluation unit

Scenario is a condition

A scenario is a condition, not a stimulus; changing the stimulus creates a test case, not a new scenario.

Test form

Specifications first

Test cases are specifications, not runnable tests, and concrete literals are bound later at fixture implementation.

Test design

Simplest direct verification

Each scenario starts with the simplest, most obvious, natural test that directly verifies it.

On-trial grouping

Analytics split under test

Analytics & Observability is tested through aggregate and individual groups so execution can show whether they are really separate capabilities.

Full benchmark methodology →

Resources and fixtures

The registered material and systems that create a consistent test environment for this benchmark.

Database

Cedarline commerce database

A shared PostgreSQL 16 commerce database used for benchmark evaluations.

AI Database Agents — Benchmark definition | AI Demos