A high score is not a release test for a data agent
[Ad Space — Insert ad script here]
Before a data agent answers a business question, I want to know what it gets wrong.
Most demos stop too early. The agent writes SQL, the query runs, and a chart appears. That proves the workflow works. It does not prove the number is right.
A better test starts backwards.
Take questions the data team already understands. Ask the agent. Compare its results with answers someone is willing to own. Keep those questions out of the examples the agent can retrieve at runtime.
That gives you a useful release control. But only if you are careful about what the score actually means.
Easy benchmarks can flatter the agent
Public text-to-SQL numbers still shape a lot of pilot conversations. They should not shape a launch decision on their own.
Spider 2.0 was built to look more like enterprise data work. Its 632 tasks sit in environments such as BigQuery and Snowflake, often with more than 1,000 columns. Solving a task can require several queries and more than 100 lines of SQL.
In the ICLR 2025 paper, a code-agent framework using o1-preview solved 21.3 percent of those tasks. The same paper reports 91.2 percent on the original Spider benchmark and 73.0 percent on BIRD. The project page also shows GPT-4o at 10.1 percent on Spider 2.0 versus 86.6 percent on Spider 1.0. Later leaderboard rows have moved, but the gap between clean academic tasks and enterprise-style workflows has not disappeared.
That does not mean your warehouse agent will be wrong four times out of five. It means a score from a cleaner distribution tells you very little about a messy one.
BEAVER reaches a similar conclusion from enterprise query logs. It is a large enterprise Text-to-SQL benchmark: thousands of question-SQL pairs, hundreds of tables, and many domains. It also breaks failures into subtasks such as table retrieval, join keys, column mapping, domain knowledge, and query decomposition.
Even after researchers gave the model the correct tables, joins, columns, domain facts, and decomposition as oracle hints, execution accuracy reached 30.1 percent. Without those hints, agentic methods on a strong recent model sit much lower. Better context helped. It did not make evaluation unnecessary. And execution accuracy itself only means the generated SQL returned the same rows as the gold SQL on that database state. It does not prove the gold SQL used the approved business definition.
The expected answer can also be wrong
A backwards test needs an answer key. That answer key deserves as much scrutiny as the agent.
Jin, Choi, Zhu, and Kang audited two public text-to-SQL benchmarks. Expert reviewers found annotation errors in 52.8 percent of BIRD Mini-Dev and 62.8 percent of Spider 2.0-Snow.
When the researchers corrected a subset and rescored 16 agents, relative execution accuracy moved from -7 percent to +31 percent. Rankings moved by as many as nine places. An earlier CIDR 2026 slice of the same line of work told the same story in a narrower cut: gold labels were frequently wrong enough to move measured performance.
Those figures describe benchmark labels, not the quality of metrics inside a typical company. But the lesson transfers cleanly.
If finance changes the definition of revenue and your test still uses the old query, a regression test rewards the wrong answer. If “last quarter” moves with the calendar, the expected result changes even when the agent does not. If two teams disagree about active customers, the score will quietly pick a side.
The gold answer is part of the system. Version it, date it, snapshot the data or tolerances it depends on, and give it an owner.
Do not let an LLM be the final judge
It is tempting to ask another model whether the generated SQL looks correct. That is useful for triage. I would not use it as the final check on a business number.
JudgeBench , published at ICLR 2025, tests model judges on difficult pairs where one answer is objectively wrong. Many strong judges, including GPT-4o, performed only slightly better than random guessing. Separate EMNLP 2025 research found that model judges can also prefer their own generations, so a naive score gap can mix bias with genuine quality.
Neither study is specifically about analytics. They still expose the problem: plausible output is not the same as verified output.
Platform tooling is getting better about the evaluation workflow, and that is useful. It does not remove the need for judgment about what “correct” means.
Snowflake’s Cortex Analyst evaluation design is a practical example of the right direction. A verified query pairs a business question with expected SQL. When that query is selected for evaluation, Snowflake removes it from the semantic view used by the agent. The example cannot be both the hint and the test. Unselected verified queries can still guide generation. That distinction matters: a high score with the test examples available measures retrieval and paraphrase matching as well as reasoning.
But Snowflake’s sql_correctness metric is still produced by a versioned model judge. Its documentation says generated SQL can vary between runs, recommends running the same evaluation several times before setting a threshold, and warns that time-relative questions can make the ground truth stale. Evaluation sets are also not auto-built from query history. You still have to curate them.
Databricks documents a related pattern: evaluation sets from curated requests or production traces, plus expert review of agent outputs and production traffic that gets negative feedback. That is a workable operating loop. It is not published evidence that a particular accuracy number predicts safe open-ended use.
The tools help you run the test. They do not decide what your organization should trust.
What I would require before release
I would not ask for one impressive accuracy number. I would ask for a small operating system around the misses.
1. Separate release questions from development examples
Keep three sets if you can: cases for iteration, cases for validation, and a locked release set that is not used as runtime guidance or a tuning target. Snowflake’s temporary removal of selected verified queries is one concrete implementation of that principle. Without it, teams can overfit the scorecard.
2. Compare executed results, not only SQL that compiles
A query can be valid SQL and still be wrong for the business. Prefer a deterministic match against a certified result table, metric service, or approved calculation. Use SQL-string similarity only as a secondary diagnostic.
3. Treat every expected answer as a versioned record
Store more than a prompt and a gold query. Keep the intended interpretation, metric-definition version, owner, effective date, permitted sources, risk tier, and the data snapshot or tolerances used when the case was certified. Re-certify after definition or model changes.
4. Stratify the set by consequence and shape
A suite of familiar dashboard questions measures familiar dashboard questions. Include certified recurring KPIs, multi-table joins, time logic, cohort and attribution cases, edge conditions, access-boundary requests, ambiguous wording, and questions that should get a clarification or refusal rather than a number.
5. Route by evidence, not by the model’s prose confidence
Auto-answer only when deterministic preconditions pass: approved definition, allowed sources, successful execution, result checks, and no unresolved ambiguity. Send disagreements, unseen high-risk patterns, access-sensitive asks, and failed checks to a named owner.
6. Report two rates separately
Publish how often the agent is allowed to answer without review, and how often those automatic answers are wrong. An agent that answers a narrow set of questions reliably may be more useful than one that answers everything with a better-looking average.
What this test cannot fix
A golden set can become a false certificate.
If it contains only easy, stable asks, it will not estimate performance on new definitions, unusual joins, permission boundaries, or questions nobody has asked before. Spider 2.0 and BEAVER both show how sharply results change when the workload gets closer to real enterprise analysis.
It also cannot repair missing context. If the agent cannot see the approved metric, does not understand the grain, or has access it should not have, the test will reveal some failures but it will not remove their cause. In those cases the binding work is definition ownership, permission-aware retrieval, clarification behavior, and certified calculation paths, not a larger scorecard alone.
There is also no primary study I found that shows replaying the last 30 days of human decisions at some match rate predicts a safe launch. Historical traces are valuable inputs to an evaluation set. They are not a safety certificate by themselves.
The point of the backwards test is not to prove that an agent is safe for open-ended use.
It is to make its known limits visible before a business user discovers them in a board deck.
Closing
A high score can be useful. It can also be comforting for the wrong reasons.
The release control that survives skepticism is narrower than a leaderboard win: held-out questions, current definitions, result-level checks, and a named person for the misses.
If your team already tests a data agent, what are you comparing: SQL that ran, a judge’s opinion, or a result against a definition someone still owns?
[Ad Space — Insert ad script here]