The bigger opportunity in AI data agents is the boring layer
[Ad Space — Insert ad script here]
Everyone is trying to build AI data agents right now.
I get it. The demos look great. Ask a question in plain English, get a SQL answer a few seconds later, and put “self-service analytics, finally” on the next slide.
But I think the bigger opportunity for data teams is somewhere else .
It is the boring layer underneath the agent: shared meaning, scoped access, deterministic checks, lineage, and a human who owns the decision when it matters.
Agents are becoming widely available. Trustworthy context is still scarce.
What keeps going wrong
A lot of the public evidence points in the same direction, even if the exact percentages differ.
MIT’s NANDA work on GenAI in business (reported mid-2025) found that most enterprise pilots still fail to show real P&L impact. Their framing was not “the model is dumb.” It was closer to a learning and workflow gap.
Gartner has been blunt about agentic projects too: more than 40 percent are at risk of cancellation by the end of 2027 because of cost, unclear value, and weak risk controls. Model performance is not the main reason they give.
Fivetran’s 2026 readiness work is more data-specific. Most enterprises say they are investing in agentic AI while admitting the data foundation is not ready. Lineage and quality keep showing up near the top of the blockers.
You can challenge any one of these surveys. The overall pattern is harder to dismiss.
When agents fail on warehouse data, it rarely looks like a broken model. It looks like this:
- two definitions of revenue
- a join that is “correct” in SQL and wrong in the business
- an agent with a service account that can see too much
- an answer nobody can reproduce next Wednesday
- a pilot with no owner once the build team moves on
That is not an intelligence problem. That is an operating problem.
The clearest number I found
One data point stood out to me, but it needs a caveat.
dbt Labs published a 2026 benchmark comparing text-to-SQL against a semantic layer on the same modeled project. Same models. Same questions. Different interface to the business logic.
Accuracy improved when the agent used the semantic layer. On the subset of questions the semantic layer could answer, one setup showed a 50 percentage-point difference.
I would not oversell that result:
- It is a vendor benchmark.
- The method is open and inspectable, which is better than most marketing charts.
- Even dbt notes that putting a full enterprise schema into context is not realistic at scale. The raw-schema baseline in the lab may be easier than what many teams face in production.
Still, it is a useful result: same model, better plumbing , better answers.
That fits how I think about this as a Head of Data working on analytics engineering and semantic layers. The hard question is usually not whether the LLM can generate SQL. It is whether the system understands what the business means, and whether we can verify how it reached the answer.
The boring layer, named
By “boring layer,” I do not mean buying a governance platform and hoping it solves the problem.
I mean five practical things that sit between the model and production data.
1. Semantic context
Metrics, entities, joins, synonyms, and approved definitions. This cannot live only in a glossary slide. It has to be available in a form the agent can actually use. Snowflake, ThoughtSpot, dbt, and others are converging on this idea for a reason: schemas do not carry business meaning. For the broader shift from BI convenience to shared AI infrastructure, I wrote about that earlier in semantic layers as AI infrastructure .
2. Governed access
The agent should inherit the user’s permissions or use a policy-enforced data product. A shared warehouse credential with broad access is a poor substitute. Databricks has been explicit about this. OpenAI’s enterprise data agent follows the same pattern by reusing existing table, row, and column rules. Control what the agent can touch instead of pretending every possible plan can be reviewed in advance.
3. Deterministic validation
SQL execution is deterministic for the same query and warehouse state. LLM SQL generation is not, even when temperature is zero. That makes the boundary between generation and execution the right place for tests, contracts, allowlists, and consistency checks. Research on text-to-SQL hallucinations is useful here because it identifies failure modes that look correct but are not. A query can be valid SQL and still be wrong .
4. Lineage, freshness, evidence
If you cannot explain where a number came from, how fresh it is, and what the agent did, you do not have a reliable analytics product. Monte Carlo and others are treating agent observability as a ship-or-no-ship concern for a reason.
5. Ownership and approval
Not every action needs a human. High-impact ones do. Risk-based routing beats blanket review. And someone has to own the agent in production after the pilot team leaves. If that sentence is hard to finish, the pilot is already in trouble.
None of this is glamorous, but each part makes the next one more valuable.
You might not need a new platform
This is the part people skip.
Two of the more useful production write-ups I found did not start with a large new AI governance suite.
One example, involving Vercel and Cube, gave the agent a well-documented semantic layer as plain files and removed most of the tool sprawl. The result depended on more structure, not more magic.
Spotify’s migration-agent work took a similar approach: tight constraints and explicit mapping rules instead of open-ended access.
My takeaway is not that teams should never buy tools. It is simpler:
If your semantic layer is messy, undocumented, and contradictory, an agent will not heal it. It will just fail faster, with more confidence.
Finish the thin layer you already have. Then wire the agent to it.
That also matches how I think about product work in this space: integrate with the existing stack, add discipline around execution, and avoid creating a second platform unless the problem genuinely requires one.
The honest disagreement
I do not think this is fully settled.
dbt has said that a semantic layer is not required for every AI workflow. Debugging a failed job is a different problem from answering “What was Q4 revenue?”
Some agent builders argue that model intelligence is already good enough for most use cases, while architecture, monitoring, and trust remain the real blockers. That seems reasonable. Some go further and expect today’s scaffolding to shrink as models improve.
Maybe some of it will.
I would still fund the boring layer today.
Permissions, validation, lineage, and ownership stay useful even if the model gets smarter. Waiting for a future model to invent your metric definitions is not a strategy. It is a delay.
What I would do next
If a Head of Data asked me where to focus next quarter, I would not start by selecting an agent framework.
I would ask:
- Which five metrics would be unsafe to get wrong in front of the board?
- Are those definitions encoded somewhere an agent can use, or only in a dashboard and somebody’s head?
- What can the agent access that a human in the same role cannot?
- What check runs before a number leaves the system?
- Who owns this when the pilot is over?
If those answers are fuzzy, another agent demo will not help much.
If they are clear, almost any capable model becomes more useful.
Closing
I am not against agents. I am against skipping the work that makes them trustworthy.
Everyone is racing to build the visible part. The quieter work is still the scarce part: meaning, access, checks, evidence, and ownership.
That is the opportunity I would bet on.
If you are already shipping agents on a warehouse, I would like to know which of these five pieces you built first, and which one is still held together by a Slack thread.
[Ad Space — Insert ad script here]