TL;DR
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A Hugging Face report describes ProvenanceGuard, a post-generation verifier for MCP agents that checks both whether a claim is supported and whether the answer attributes it to the right source. In a test using 40 answers, it caught 138 of 139 claims experts said should not pass, while also flagging 67 claims experts considered supported.
A paper on ProvenanceGuard describes a way to check whether an AI agent using the Model Context Protocol (MCP) has matched each claim to the source it names or implies. In tests on medical-agent answers, the system caught 138 of 139 claims that human experts said should not pass, while also flagging some claims experts considered supported.
The method targets cross-source conflation: a claim may be supported somewhere in an agent’s collected evidence but attributed to the wrong tool output. For example, a refund window might appear in a policy document, while an answer incorrectly says it came from an account record. A verifier that pools both sources could mark the claim as supported without detecting the attribution error.
ProvenanceGuard runs after an agent generates an answer and reads its captured MCP trace, retaining tool outputs and source IDs. It breaks the answer into claims, identifies a relevant source for each, checks support, compares that source with the one the answer names or implies, and returns claim-level verdicts plus an overall allow-or-block decision.
The reported experiment used local models for claim decomposition, source matching and support checks. The authors say those choices describe the evaluated setup, rather than a requirement of the method. They also describe a repair step that can revise a blocked answer or offer a safe fallback, followed by another verification check.
Why Source Attribution Matters
For an MCP agent, where a fact came from can change how a reader should interpret it. A detail from a patient record is not interchangeable with a finding from medical research, and a customer’s account data is different from a general policy. An answer that names the wrong source can mislead even when the underlying fact appears somewhere in the evidence.
The authors argue that familiar factuality checks, including RAGAS faithfulness and systems such as MiniCheck, AlignScore and SummaC, generally assess support against pooled evidence. In their usual form, they do not establish which tool output supports each claim or whether that matches the answer’s attribution. ProvenanceGuard is intended to preserve that distinction for data-sensitive review.
The reported results show a cautious trade-off. The system caught nearly all claims experts judged should not pass, but it also held 67 claims experts considered supported, sending them for review or repair. That may suit settings where teams prefer extra scrutiny, though the report does not establish how the approach performs in other workflows or at larger scale.
From Pooled Evidence to Source IDs
MCP lets an agent draw on multiple tools, such as search, structured records, databases and metadata. As the report describes it, the agent can combine those outputs in one answer. Once evidence is combined, a checker that sees only the pooled material may find a matching fact but lose the connection between that fact and its original source.
The paper frames source identity as part of verification. ProvenanceGuard does not retrain the agent; it examines the recorded tool trace after the response has been generated. The system’s claim and source checks can also be adapted to hosted models, according to the authors, but a hosted setup would need its own testing and calibration. The results in the report come from the local configuration.
The medical evaluation used 281 real traces involving patient records, research articles and other tools. For the main test, experts reviewed 361 claims from 40 answers that had been set aside from the data used to develop the system. The report presents medicine as a useful setting because patient-specific information and general research findings must remain distinct.
Limits of the Reported Test
The reported evaluation covers 40 answers in a medical-agent test; the material provided does not establish performance across other fields, agent designs or larger deployments. It also does not give enough detail here to assess how results might change with different models, thresholds or source-matching methods.
The system’s cautious behavior is visible in the result: it flagged 67 claims experts considered supported, in addition to catching 138 claims experts said should not pass. The report describes these as held for review or repair, but the supplied account does not quantify review time, repair success or the effect on users. For claims with an identifiable source, it selected the correct source about 86% of the time in this test; how that rate transfers to other settings is unknown.
Testing Beyond the Medical Traces
The next step is further testing on agents and source types beyond the medical traces described in the report. Teams considering a hosted-model version would need to test and calibrate that setup separately, as the authors state; the local-model results do not establish its performance.
Further evaluations could clarify how often the repair step produces an acceptable, source-grounded answer and how the system’s cautious decisions affect review workloads. The paper is available on Hugging Face and, according to the report, on arXiv. The material provided does not specify a release schedule or a broader deployment plan.
Key Questions
What does ProvenanceGuard check?
It checks whether claims in an MCP agent’s answer are supported by relevant tool outputs and whether those sources match the answer’s stated or implied attribution.
What is cross-source conflation?
It is when a claim is supported by evidence from one source but the answer attributes it to another, such as describing a policy detail as if it came from an account record.
How did it perform in the reported test?
Experts said 139 claims should not pass; ProvenanceGuard caught 138 and let one through. It also held 67 claims that experts considered supported. The test covered 361 claims from 40 medical-agent answers.
Does the system require the local models used in the experiment?
No. The report says those models were used for the evaluated setup, not required by the method. A hosted-model setup is possible but would need separate testing and calibration.
Source: rss
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
