Companies are increasingly discussing how AI agents fit into data teams. Some of the work that used to be done by an analyst, BI specialist or data engineer can be handled by an agent. It receives a question, finds the data, writes a query, returns an answer and ideally explains what the result means.
That’s an appealing prospect for company leaders: faster reporting, less waiting for the data team and more decisions backed by data. But there is a lot of work between connecting a model to a database and getting a trustworthy result. You cannot just give an agent access to the warehouse and expect it to figure out everything else.
An agent does not replace ETL pipelines. Those pipelines still bring data in and prepare it for reporting. An agent can change what happens next: it takes a question about prepared data, turns it into a query, and returns an answer that needs to be checked. That is why data teams also need to define metrics, relationships between data, and the context an agent needs to understand the question.
If the agent does not know company definitions, the data model, source limitations, the domain and the purpose of the question, it will sound confident even when it is wrong. That is not only a technical inconvenience. It is a management risk, because a bad answer can look just as convincing as a good one.
Agentic BI makes the work we used to put off more valuable. Good data modeling, metadata, column descriptions, clear table names, metric definitions, and warehouse maintenance are essential infrastructure that helps the agent write fewer queries, consume fewer tokens, put less load on the database and, most importantly, increase the chance that it answers correctly.
What I mean by agentic BI
In this article, an agent is an AI tool that works toward a goal, beyond following precise instructions such as “click this button.” Based on rules, permissions and available tools, it decides what to do next: find context, write SQL, run a query, check the result, suggest a dashboard change or send a change for review.
Agentic BI is the use of such an agent within a company’s data environment. The goal is not merely a nicer chat interface on top of dashboards. The ambition is bigger. A salesperson, marketer or manager does not have to wait for the data team for every question. They can ask an agent that knows the data warehouse, ETL processes, metrics, dashboards and company rules. The agent answers, or suggests what is missing in a dashboard: a filter, dimension, metric or another view of the data.
But that raises the bar for the environment. An agent can handle more than many companies currently assume. The larger the part of the work we hand over to it, the more important its foundations become. It needs preparation, care, human oversight and a system that improves over time based on where the agent failed.
Agentic BI is not just chat over dashboards. It is a test of whether a company can share its knowledge about data.
What context actually means
Context is the set of information, knowledge and experience we use to make decisions. A person builds it over years: from meetings, production incidents, conversations with the business, historical compromises in the data and unwritten rules that move between people inside the company.
For an AI agent, it is similar. It is not enough to know where a table lives. It needs to understand what the company considers an order. When someone asks about results, it needs to know whether to use revenue, margin, order count or marketing efficiency. And when someone asks about Slovakia, it needs to know how to select the right customers or orders in the warehouse. It also needs to know the time zone, reporting exceptions and historical artifacts in the data warehouse.
There is no single universal context. In data teams, at least three layers meet: general business context, domain context and data context. These layers overlap. They complement each other and together form the picture of the company that an analyst uses when answering a question. The agent does not have that picture without explicit help.
Business context is mainly an understanding of what the company does, where it is going, what goals it has, how its processes work and why some decisions are being made right now.
Domain context describes what individual teams deal with, which tools they use, which projects they have open and what they expect from data. Data context is the specific meaning of metrics, tables, columns, times, currencies and rules in the warehouse. It builds on business and domain context: if we do not know how the company and the team work, it is hard to describe precisely which table or metric should answer the question.
Marketing makes this easy to see. If the agent is supposed to answer a marketing team, it is not enough for it to know impressions, traffic and conversion. It needs to know what projects the team is working on, where the team is struggling and what the team actually expects from the answer.
The agent is a new colleague, not a mind reader
Think of introducing an agent as onboarding a new colleague. When an analyst joins a company, it is not enough to say: "Here is the database, figure it out." They need to know where the sources of truth are, which tables are historical, what column names mean, how the main KPIs are calculated and when they should ask rather than guess.
The agent needs the same information. It has one disadvantage: it was not in your meetings, it did not hear the historical debates and it cannot read the minds of senior people on the team. If knowledge stays in people's heads, the agent has nowhere to get it from.
A team with a data warehouse, dashboards and a few people who "know where things are" is not necessarily a bad team. It just may not be ready for the agentic era. That difference matters. Agentic BI does not punish weak teams. It quickly shows which know-how was never written down, tested or made transferable.
A simple question is often not simple
Take an everyday question: how do you calculate the number of orders? A model without context may answer with syntactically correct SQL:
SELECT COUNT(order_id)
FROM orders But the right answer in your company may look very different:
SELECT COUNT(DISTINCT order_id)
FROM orders
WHERE order_status <> 'canceled'
AND order_stream = 'online'
That difference is the whole point. Should the agent count all rows or unique orders? Do canceled orders belong in the number? Are we looking only at the online channel or all channels? Do we count a cart creation, a payment, warehouse confirmation or something else as an order?
If those answers live only in the head of a senior analyst, the agent has nowhere to get them from. And if it starts guessing, the company gets a fast answer, but not necessarily an answer it can use for decisions.
A business question hides technical questions inside it
It is even clearer with a question such as: "How did Slovakia perform in Q1 this year compared with last year?" It sounds like a routine question from leadership. For a data agent, it is a bundle of definitions and decisions it has to find somewhere.
First, it needs to know how we measure "performed": revenue, margin, orders, number of customers, retention, marketing efficiency or a combination of metrics. It must identify who is asking, because a CEO will often need a different answer than a marketing team. And it needs to know how to filter Slovakia in the warehouse: by customer country, delivery address, market, currency or organizational unit. On top of that come Q1, time zone, currency and the choice of the right tables.
Imagine an anonymous e-commerce company that expanded to Slovakia. Sales tracks revenue by market, marketing by campaign currency, logistics by delivery address and finance by accounting entity. When the CEO asks how Slovakia is doing, it is not only about SQL. First, it must be clear which definition of the Slovak market is relevant for the decision.
The environment hasn’t given the model what it needs. An agent without context does not know which meanings are correct inside the company. It can write a query, but it cannot guarantee that the query matches what the person actually needed to find out.
Agentic BI does not start with model selection. It starts with pulling rules, exceptions and agreements out of people's heads.
What data context needs to include
When I say context is infrastructure, I do not mean one long Notion document. The agent does not need a novel about the company. It needs a set of machine-usable and human-usable sources that it can find, cite, verify and use while working.
In a data environment, I would start with the following blocks. These are not theoretical categories. They are the places where an unprepared agent most often starts guessing.
Metric definitions
Every important metric needs a clear definition: a business-readable name, description, entity, calculation, primary key, exceptions, default filters and information about where it is used. Ideally, this is not only text for people, but a structured source that can feed the agent, tests and documentation.
{
"entity": "orders",
"business_name": "Orders",
"description": "An order represents a confirmed customer purchase.",
"category": "core_kpi",
"kpi_tier": "main",
"is_main_kpi": true,
"primary_key": "order_id",
"default_metric": {
"name": "Order count",
"calculation": "count(distinct order_id)"
}
}
The important point is that a metric is not just a formula. It is an agreement between the business and the data team. Without it, the agent will keep revisiting decisions that should already have been decided.
With simple metrics such as order count or total revenue, the agent will often be at least roughly right. With more complex metrics, such as margin, problems show up quickly. What belongs in costs? How do we handle shipping, returns, discounts, refunds, marketplace fees or the difference between an accounting and management view? If the company does not have the answer described in a structured way, the agent will fill it in from patterns it knows from training data. That may not be your definition.
The slide above is a illustrative example, not a production measurement or a universal benchmark. It illustrates the principle: when a metric exists in a structured form, the agent does not have to keep guessing the logic from scratch.
Table and column metadata
The second big area is metadata. Not the data itself, but information about what tables and columns mean. Tables and columns need descriptions. Columns with the same name across tables should have the same meaning and data type.
Seemingly small details can cause real problems here. A postal code is not an integer just because it consists of digits. Currencies, time zones and units do not belong in assumptions. They belong in names, descriptions and rules. The more the agent has to infer, the more room it has for bad interpretation.
A practical detail: column descriptions can also act as instructions for the agent. Tell it when to use a column, when to avoid it and what to watch out for. Input for AI is often cheaper than bad output.
Companies often neglect metadata at the column level. For an agent, it is signage. The better you describe what a column means, when it is created, which units the value uses and which exceptions it has, the less the agent has to try and guess.
The more the agent has to guess, the more expensive and less trustworthy its answer becomes.
Technical infrastructure
Besides the meaning of data, the agent needs to understand the environment it works in. It must know which tools the company uses and what they are for: where the warehouse runs, where transformations live, where tests run, where deployment happens, where logs are read and which actions it can perform through CLI, API or MCP.
If the system can only be controlled by clicking through an interface, agents will eventually get stuck. Not because they cannot click, but because click-based workflows are slow, harder to audit and much harder to describe as repeatable procedures.
The data model
The data model needs its own map. It is not enough that it exists in the heads of the people who built the warehouse. The agent needs to know whether you use a star schema, snowflake, Data Vault or your own layers. It needs to understand what staging, marts and reporting tables are for, where the sources of truth are and where there are intermediate results.
Testing requirements belong here, too: what a new change must satisfy before it can be trusted. This is exactly the place for `.skills`, repository instructions and architecture rules. Not as AI decoration, but as a way to keep work consistent across people and agents.
Data types, currencies and time zones
Then there are rules that look boring but decide whether an answer is correct. These include data types, currencies, time zones, units and exception rules. A person who has worked in the company for a long time often treats them as obvious. The agent has no reason to take them for granted.
That is why they need to be written down. For example: all timestamps are converted to UTC. Financial values are reported in EUR and exchange rates come from a specific table. Refunds belong to the same period as the original order, or to the processing period. Without such rules, the agent will not answer according to the company’s agreed definitions. It will create its own variant.
Tests, evaluations and data contracts
Once outputs affect users, hoping the agent gets it right is not enough. You need tests over data, evaluations over typical questions and data contracts between sources, transformations and consumers. A new version that may change answers should pass a set of examples just like an application passes tests.
What a warehouse ready for agents looks like
All of this can sound abstract, so I would use a simple test. Take a new person who knows the technology but does not know your company. Give them access to the warehouse, documentation and descriptions. If they can find their way around without constantly looking for the one person who "is the only one who knows how it works", you have a good foundation for an agent too.
A clean warehouse does not mean a perfect warehouse. It means a warehouse where tables have understandable names, columns have descriptions, model layers make sense, main metrics have definitions and there is a map that helps both people and agents decide where to start. A perfect state is not necessary. Knowledge that can be shared is essential.
The biggest risk is not token cost, but loss of trust
When an agent lacks good context, the company pays several times. Part of the cost is visible immediately: the agent consumes more tokens because it has to reason more and ask for more explanation. It generates more SQL queries because it tries paths it would not need to try with a good model and metadata. It puts load on the database, which can become a real operating cost in larger companies.
The worst cost is loss of trust. If the agent starts returning unreliable answers, or even breaks dashboards that the team uses, people stop trusting not only the agent, but also the data team and the numbers in general. Financial costs can be optimized. Lost trust in data is much harder to repair.
With queries, it is important to distinguish volume from efficiency. If the agent generates many queries because it answers many user questions, that can be an understandable operating cost. But if it needs dozens of attempts for one common question, that is a signal that it lacks context, metadata or the right data model.
Repeated agent work needs to be controllable
When agents repeatedly work with company data or change business logic, a prompt alone is not enough. The company needs to control which context the agent reads, which actions it can take, what it changed, and how the result is checked. CLI, MCP, APIs, GitHub, and CI/CD pipelines offer practical ways to expose, version, and review that work. That is why I treat them as important parts of designing production data and agent systems.
Why? Because the agent needs to do work repeatably. It needs to read context, run commands, change files, open pull requests, run tests, write results and let a human inspect the diff. It must be clear which action it performed, what it was based on, what it changed and how we know the output passed.
That is why the technology stack must also be selected based on whether it can be controlled programmatically. If a tool has no API, CLI, MCP interface, auditable export, GitHub integration or path into CI, it will be a weak point for agent workflows. It may be good for a person in a browser, but weak for a system that should work repeatedly, testably and under control.
Manual clicking and unclear permissions are technical debt in an agentic environment. Not on the first day, but once an agent is supposed to work over production logic, data or decisions, the lack of programmatic interfaces becomes a serious obstacle to safe operation.
A click-based tool can be good for a person. For an agent, a better interface is one that can be described, run, versioned and checked.
In practice, this means that agents working with data are not only a data team project. They also concern how the company versions logic, describes systems, checks changes, stores decisions and quickly turns unwritten know-how into usable context.
Humans must stay in the loop
In agentic BI, human-in-the-loop is a necessity, not a weakness of the system. In plain terms: a human must stay in the loop. At the beginning, someone from the data team should check the questions the agent receives, the SQL queries it generates and the answers it returns. Not to keep the agent manually controlled forever, but to improve the system.
In practical terms, that means capturing questions the agent could not answer, ambiguous metric interpretations, expensive or unnecessary queries and answers where the agent was unsure. Some of that oversight can move into evaluations over time. Some will stay with people, because someone has to decide whether an answer matches business reality, not just SQL syntax.
The valuable result of oversight is not just a corrected answer. The important output is added context: a better column description, a clearer metric definition, a new evaluation example, an exception rule or documentation of a process that used to live only in one person's head.
Company checklist: what to prepare before giving an agent access to company data
If I had to recommend first steps to a company, I would not start with a large AI rollout. I would use this checklist and verify whether the company has a contextual foundation to build on.
- Clarify why you want agentic BI and what kind of decision or work it should improve.
- Choose the ten most common business questions and write down how a good analyst answers them today.
- Describe the company's main KPIs: those used in team goals, leadership and regular reporting.
- Create structured definitions for the main metrics, including exceptions, filters, owner and typical uses.
- Add descriptions for key tables and columns, especially where the name is not enough or can mislead.
- Write down rules for currencies, time zones, data types, refunds, cancellations and other company exceptions.
- Describe the layers of the data model: what is a source of truth, what is an intermediate output and what is a reporting mart.
- Create an agent map: where the data is, how queries are run, where tests live and what the agent must not do.
- Build an evaluation set from real questions and expected answers.
- Connect changes to GitHub, CI and review so context and logic change under control.
- When choosing tools, prefer those with API, CLI, MCP or another programmatically controllable entry point.
I would not wait for a perfect description. Waiting can be more expensive than an incomplete start. The first version of metric definitions and documentation does not have to be ideal. It can be migrated, sharpened and extended. But without a first version, the agent has nothing to build on and the team has nothing to improve.
This does not sound as impressive as a demo where an agent answers a question in chat on the first try. But this work decides whether the agent ends up as a toy after a month, or becomes a real part of how the team works.
When the company is not ready yet
The biggest warning sign is simple: key knowledge exists only in people's heads. If the team repeatedly says, "we need to wait for Pepa; he’s the only one who knows how it works", agentic BI will hit a wall very quickly. The agent cannot use knowledge that is not available anywhere.
That does not mean the company must not start. It means the first work is not model selection or buying a new tool. The first work is moving critical knowledge from people's heads into metrics, metadata, documentation, tests and rules that both humans and agents can use.
Make the knowledge usable
AI agents will not remove the need for data teams. They will more likely bring into the open everything good data teams have carried in their heads until now: definitions, compromises, rules, domain knowledge and the feel for what a question really means.
Companies that want to use agents seriously need to make that context available, structured, testable and controllable through tools. Then the agent has a chance to help. Without that, it is only a faster way to produce convincing-sounding uncertainty.
If you are deciding how to connect data, metrics, and AI, see my data and AI architecture design. If your team first needs to align around a specific question, the From ETL to AI agents workshop may help. If you do not yet know the state of your data environment, start with a data audit.