Your AI Has Sources. Can It Prove What They Support?
In a 2023 human evaluation of four generative search engines, researchers found that only 51.5% of generated sentences were fully supported by their citations, and only 74.5% of citations supported the sentence they were attached to. The answers looked referenced. About half of the generated sentences were not fully supported by the citations provided.
That gap sits underneath a common assumption in enterprise AI: once a retrieval-augmented system returns an answer with citations, the answer has become trustworthy. The presence of the citation becomes a proxy for a verification step that may never have occurred.
It skips several questions.
- Does the cited passage support this particular claim, or only the topic?
- Is the quotation still present in the current version of the source?
- Did the underlying document change after it was ingested?
- Can the generated output be connected back to the exact evidence used, rather than to a document title?
- What happens when the evidence is insufficient?
As AI systems move from answering questions to informing consequential decisions and taking actions, the provenance question moves with them. "This came from document X" is a pointer. It is not yet a proof.
A reference can be present while the support is absent
In a typical retrieval-augmented generation workflow, a retriever selects passages relevant to a query and a language model generates an answer conditioned on the retrieved material. In systems that expose citations, those passages or their source documents may then be referenced alongside the answer.
The references can be real documents. The relationship between a reference and the sentence it appears to support is a separate matter.
Three things can fail independently. The passage can be relevant to the subject and still not support the specific claim made. The claim can be supported by a passage that has since been edited, withdrawn or superseded. And the output can be impossible to connect to the exact evidence used because the system recorded a document, not a location in it.
The ALCE benchmark, published at EMNLP 2023, measured citation quality in end-to-end systems that retrieve evidence and generate answers with citations. On one of its datasets, even the best models lacked complete citation support 50% of the time.
These figures relate to the systems, datasets and evaluation designs used in the studies. They should not be read as a universal failure rate for retrieval-augmented systems. The finding that carries is narrower and more durable: a citation can be present, look correct and still not support the sentence it is attached to, and a reader has no way to tell from the surface.

The NIST Generative AI Profile names the mechanism. Confabulation, it says, covers "confidently stated but erroneous or false content", and GAI outputs "may also include confabulated logic or citations that purport to justify or explain the system's answer, which may further mislead humans into inappropriately trusting the system's output". A citation that looks like evidence can increase trust while adding none.
Provenance belongs to the claim, not the document
Document-level provenance answers one question: which source was consulted. For a system whose output will be acted on, that is the weakest useful answer.
Consider what a careful auditor or lawyer means by a citation. A pinpoint reference names the document, the edition or version, the page or paragraph, and the words relied on. A reader can go to that place and judge for themselves whether the words carry the weight placed on them. "See the contract" is not a citation. Neither is "see the policy manual" when the manual has been revised twice since the answer was written.
Claim-level provenance restores that discipline. For each statement a system makes, four things need to be connected:
- The claim, as generated.
- The evidence: the exact quotation relied on, verbatim, with its location in the source.
- The source state: which revision of the source the quotation was taken from, so a later change is detectable.
- The support check: whether the quotation actually supports the claim, assessed separately rather than inferred from the presence of a citation.
The same NIST profile describes high-integrity information as content that "can be linked to the original source(s) with appropriate evidence", that "has a clear chain of custody", and that "creates reasonable expectations about when its validity may expire".
The last clause is particularly important for retrieval systems. Evidence has a shelf life. A source changes after ingestion, and a claim that depended on the earlier text may no longer have the same evidentiary basis.
What an evidence-bound record looks like in practice
A company knowledge base or internal repository offers a useful example of how claim-level provenance can work in practice. When an AI system generates content from internal policies, research, decisions or other organizational material, the objective should be to preserve the connection between each material claim and the evidence supporting it.
An evidence-bound approach can retain the relevant source, the specific passage relied on and the version of that source used at the time. The support relationship can then be evaluated separately from the generation of the answer. If sufficient evidence cannot be established, the system can flag, qualify or withhold the claim rather than presenting a citation that merely appears relevant.
Source changes also need to be accounted for. When an underlying document is revised, claims that relied on the earlier version may need to be reassessed. Maintaining that relationship makes it possible to identify which generated outputs are affected instead of treating the knowledge base as a static collection of documents.
This changes the role of a citation. It becomes part of a traceable evidence record: what was claimed, what supported it, which version of the source was relied on, and whether that support remains valid as the underlying information changes.
The broader principle extends beyond company knowledge bases. Wherever AI-generated information informs a consequential decision, provenance should make the evidentiary basis of that information inspectable and capable of being reassessed.
What this does not solve
Claim-level provenance can establish that a statement is supported by evidence from a particular source and version. It cannot establish that the source itself is correct, current or appropriate. A claim may be faithfully grounded in a document that contains inaccurate or outdated information. Provenance makes the evidentiary basis inspectable; the quality of that evidence remains a separate question.
The same applies to the assessment of whether evidence supports a claim. Where that judgment is automated, it also needs to be evaluated rather than treated as inherently reliable. Claim-level provenance therefore strengthens traceability, but it does not eliminate the need to assess source quality and the reliability of the systems interpreting that evidence.
From what the AI knew to what it did
For a governed agent, knowledge provenance is the first link in a longer chain. If an agent offers a regulatory interpretation, recommends an action or provides information used for a consequential decision, the questions run in sequence.
What did the AI know? What supported it? What did the AI decide? What did it then do?
Regulation increasingly treats parts of that chain as operational requirements rather than transparency principles alone. For high-risk systems, Article 12 of the EU AI Act requires that the system technically allow automatic recording of events over its lifetime, with logging capabilities that ensure a level of traceability appropriate to the system's intended purpose. Article 11 requires technical documentation to be drawn up before the system is placed on the market and kept up to date.
Both obligations apply specifically to high-risk systems, but the design lesson generalizes: traceability is something the system must produce continuously, and documentation that was true at deployment is not sufficient later.
NIST frames the same point as lifecycle risk management. Its Generative AI Profile is organized around managing risk "across various stages of the AI lifecycle", and among its suggested actions concerning third-party content is to maintain records of changes "to promote content provenance, including sources, timestamps, metadata". Provenance, in that framing, is not a property of an answer. It is a record the system keeps as it operates.
This is the broader chain Blade Labs is building toward with ZeroH.
For governed AI agents, traceability should extend beyond the information that informed an output. It should also make consequential runtime decisions and actions reconstructable: what the agent proposed to do, which policy decision applied, whether an approval was required, and what happened when the action reached execution.
On governed runtime paths evaluated so far, ZeroH records evidence associated with policy decisions, approvals and execution outcomes within the relevant runtime context. The aim is to connect these operational records with a broader evidence model in which organizations can inspect both the basis for an AI-generated claim and the subsequent decisions or actions that followed from it.
The distinction matters. A citation can help answer what supported the AI's output. Runtime evidence can help answer what the system decided and did next. Together, these forms of traceability move toward a more complete record of AI activity—from evidence, to decision, to action.
Learn how ZeroH approaches runtime governance and verifiable evidence at zeroh.io.
Frequently asked questions
What does GAI mean?
GAI stands for Generative Artificial Intelligence. It refers to AI systems that generate content such as text, images, audio or code. In this article, the term appears in connection with the NIST Generative Artificial Intelligence Profile.
What does RAG mean?
RAG stands for Retrieval-Augmented Generation. It is an approach in which an AI system retrieves information from external sources and uses that information as context when generating an answer. RAG can help ground an answer in source material, but retrieval alone does not establish that every generated claim is supported by the cited evidence.
What is ALCE?
ALCE stands for Automatic LLMs' Citation Evaluation. It is a benchmark introduced by researchers Gao, Yen, Yu and Chen to evaluate systems that retrieve supporting evidence and generate answers with citations. Its evaluation includes citation quality as well as correctness and fluency.
What is claim-level provenance in AI?
Claim-level provenance connects an AI-generated statement to the specific evidence used to support it. Rather than pointing only to a document, it can identify the exact quotation or passage, its location, the relevant source revision and whether the evidence passed a support check.
Is a citation enough to make an AI-generated answer trustworthy?
No. A citation identifies a source, but its presence does not establish that the cited passage actually supports the claim beside it. Research on generative search systems has found that answers can contain citations without every generated statement being fully supported by them.
What is the difference between document-level and claim-level provenance?
Document-level provenance identifies the source consulted by the system. Claim-level provenance creates a more precise connection between an individual statement and the evidence supporting it. This makes it possible to inspect what evidence was relied on and which version of the source was used.
Why does source versioning matter for AI provenance?
Sources can be amended, replaced or superseded after an AI system has used them. Recording the source revision makes it possible to identify which version supported a claim and to determine which outputs may need to be reassessed when the underlying evidence changes.
Does claim-level provenance prove that an AI answer is correct?
No. Provenance can establish that a claim is supported by particular evidence, but it does not establish that the underlying source itself is correct, current or appropriate. Source quality and the accuracy of the support check remain separate assurance questions.
How is provenance different from AI logging?
Provenance addresses the evidence behind a claim: what information supported it and from which source state. Runtime logging can record what happened during execution, including policy decisions, approvals and actions. For governed AI agents, these records can form different parts of a broader traceability chain from evidence to decision to action.
What does NIST mean in this context?
NIST is the U.S. National Institute of Standards and Technology. The article refers to the NIST Generative Artificial Intelligence Profile (NIST AI 600-1), which addresses risks associated with generative AI, including confabulation and information integrity.
What happens when an AI system cannot find sufficient evidence for a claim?
That depends on the system's governance design. Possible responses include refusing to make the claim, explicitly qualifying the answer or escalating it for human review. In an evidence-bound system, insufficient evidence should be surfaced rather than hidden behind a plausible-looking citation.
How does ZeroH approach verifiable AI evidence?
ZeroH is being developed around the principle that consequential AI activity should be reconstructable from retained evidence. On governed runtime paths evaluated so far, this includes evidence associated with policy decisions, approvals and execution outcomes. Blade Labs' evidence-bound library work separately connects generated claims to supporting quotations and source revisions. Together, these approaches explore traceability from the evidence supporting an AI output through to subsequent decisions and actions.
References
1. Liu, Zhang and Liang, Evaluating Verifiability in Generative Search Engines