Can You Trace Where an AI Answer Came From? Data Lineage for AI


A figure appears in a board paper. It is specific, it is confident, and it was produced by an AI assistant that someone in finance asked a plain-English question. A director asks the obvious question: where did that number come from? The room goes quiet, and the honest answer is that nobody can reconstruct it.
In the reporting world that preceded this, the chain was traceable. A number on a dashboard came from a measure, which came from a table, which came from a nightly load out of a named source system, and a competent analyst could walk that path in an afternoon. Generative AI collapses the chain. Retrieval selects documents, the model paraphrases them into prose, and unless somebody kept a record, the trail ends at an interface that no longer remembers what it read.
This article covers what data lineage means in practice, why AI turns it from a technical housekeeping concern into a governance one, how lineage for AI differs from the lineage your BI team already understands, what an adequate record actually contains, and where a mid-market organisation should start.
Key Takeaways
Lineage answers "where did this come from". Without it, an AI output is an assertion rather than something you can stand behind.
AI breaks the traceable chain. Retrieval varies, sources are unstructured, and the output is prose rather than a value read from a cell.
Record retrieval, not just results. What the model read, under whose permissions, and at what version is the part almost nobody captures.
What Is Data Lineage in Plain Terms?
Data lineage is the documented path a piece of information travels from where it originated to where it appeared in front of someone, including every transformation applied along the way. It answers three questions: where did this come from, what was done to it, and who touched it. Nothing more exotic than that.
Most organisations already have lineage of a sort for their structured reporting, even if nobody uses the word. The warehouse documentation lists source systems. Transformation code sits in version control. Someone can point at a report field and name the table behind it. It may be patchy and live in a person's head rather than a tool, but it exists.
The distinction worth holding onto is between technical lineage, which maps tables and columns, and business lineage, which explains in words how a business concept such as active customer was defined and by whom. Boards care about the second. Auditors and regulators end up asking for both.
Why Does AI Turn Lineage Into a Board Issue?
Because AI removes the human checkpoint that used to sit between raw data and an assertion. A report was built by someone, reviewed by someone, and carried an implicit signature. An AI answer carries none of that, yet it arrives in the same fluent, authoritative register, and it is increasingly used to inform decisions that a board is accountable for.
Directors are expected to satisfy themselves that the information they rely on is reliable. That obligation does not soften because the information was generated rather than compiled. If an AI-derived figure informs a capital decision, a disclosure or a customer commitment, the organisation needs to be able to show its provenance on request, and "the assistant said so" is not a provenance record.
There is a privacy dimension as well. If an answer surfaced personal information, the organisation may need to establish which records were involved and whether the person who received the answer was entitled to them. That is a lineage question wearing a compliance hat, and it is very difficult to answer retrospectively if nothing was logged at the time.

What Actually Breaks When Nobody Can Trace an Answer?
Three things, in escalating order of cost. You cannot confirm whether the answer was wrong. You cannot establish who else received the same wrong answer. And you cannot fix the underlying cause, because you never learn which source produced it.
The first is the quiet one. When a figure is disputed and nobody can trace it, the argument becomes a contest of confidence rather than evidence, and people stop trusting the tool for anything consequential. That is an expensive way to discover a governance gap, particularly after a programme sold on productivity.
The second and third are what turn a mistake into an incident. Establishing scope is the first real task in any response, as we set out in our guide to building an AI incident response plan, and scope is impossible to determine without a record of what was retrieved and shown. Root cause is equally blocked: a wrong answer might come from an outdated policy document, a duplicate file, a mislabelled record or a genuine model failure, and these have entirely different remedies. Most of them are accumulated data quality debt already sitting in the estate rather than anything the model did.
How Is AI Lineage Different From Traditional BI Lineage?
In three specific ways. Retrieval is variable rather than fixed, the sources are unstructured documents rather than tables, and the output is generated language rather than a value read from a cell. Each of those breaks an assumption that conventional lineage tooling was built on.
Variability is the least intuitive. A dashboard measure resolves the same way every time it runs. A retrieval step selects whichever content it judges most relevant to the phrasing of the question, so two similar questions can draw on different documents and produce different answers, which is the mechanism behind Copilot giving different answers to the same question. Lineage therefore has to be captured per interaction, not documented once per pipeline.
The unstructured problem is equally awkward. The relevant unit is not a table and column but a passage inside a document, which has a version, an owner, a location and a set of permissions, all of which may have changed since. And because the model paraphrases, the wording a user reads may not appear anywhere in the source, which is exactly the gap where hallucinations take hold.
What Does Adequate Lineage Look Like in Practice?
Adequate means you can reconstruct, weeks later, what a specific answer was built from. That requires a record at five layers, and most organisations today capture only the last one.
Layer | What to record | Question it answers |
Source | System, document, version and date | Where did this originate? |
Access | Identity and permissions at retrieval time | Was this person entitled to see it? |
Retrieval | Which passages were actually returned | What did the model read? |
Model | Model version, system instructions, settings | What produced this wording? |
Output | Timestamp, citations shown, user response | What did the person actually see? |
The retrieval layer is the one that gets skipped, and it is the one that matters most. Without it you know a question was asked and an answer was given, but not what sat in between, which is precisely the part under dispute when an answer turns out to be wrong.
Citations are the user-facing half of this. An assistant that shows which documents an answer drew on gives every reader the ability to check before acting, which is worth more day to day than any log. Designing that path properly, so retrieval respects entitlements and records what it did, is architectural work rather than a configuration setting, and it belongs in the same conversation as trusted data architecture and governance.
Where Should a Mid-Market Organisation Start?
Not with a lineage platform. Start by narrowing the scope to the content your AI tools can already reach, because that set is usually far smaller than the estate as a whole and far larger than anyone assumes.
Three steps carry most of the value. First, inventory the repositories in scope and record for each one an owner, a currency date and whether its permissions are trusted. Second, require citations in any assistant used for decisions, so that provenance is visible at the point of use rather than reconstructed afterwards. Third, agree an interaction log with a defined retention period, so that questions asked in March can still be answered in June.
That sequence also surfaces the uncomfortable findings early: the policy library with three versions of the same document, the folder inherited from an acquisition nobody has permission-reviewed since. Establishing that picture is the substance of preparing your data for AI, and it is considerably cheaper to do before a deployment than during an investigation.
Traceability Is What Turns an AI Answer Into Evidence
An untraceable answer is an opinion delivered with unusual confidence. A traceable one is evidence, because somebody can follow it back, test it, and either confirm it or correct it at the source. That distinction determines whether AI output can safely be used in decisions the organisation must later defend.
The organisations handling this well are not the ones with the most sophisticated tooling. They decided early which sources their assistants could read, insisted answers show their working, and kept enough of a record to reconstruct a decision months later. None of that requires a large programme, and all of it is harder to retrofit once the tools are in daily use.
If you cannot currently answer where a given AI-generated figure came from, that gap is worth closing before the next one reaches a board paper.




