AI News Feed
Market watch
Large Language Models

ProvenanceGuard Verifies Source Attribution in MCP-Based LLM Agent Answers

A paper published on the Hugging Face blog describes ProvenanceGuard, a post-generation verification layer for MCP-based LLM agents that checks not only whether a claim is supported by evidence but whether it is attributed to the source the answer names or implies.

A source-blind verifier may pass such a claim because the fact exists somewhere in the pooled evidence, while a source-aware verifier should not. The blog post gives the example of a customer support agent that answers, "According to the account record, this plan includes a 30-day refund window." The refund window may be real, but it is stated in a policy document rather than the account record the answer points to. Pool the two sources together and the claim looks supported; keep them separate and the attribution is wrong. In a data-sensitive setting, the paper argues, a wrong attribution can be as damaging as a wrong fact. The same pattern appears in a clinical agent, where a patient-specific medication detail taken from a patient-history tool becomes misleading once the answer presents it as a finding from the medical literature.

The blog post states that faithfulness scores, useful as they are, are not enough for MCP agents, because an answer carries provenance that is sometimes explicit, as in "according to the account record," and sometimes implicit. ProvenanceGuard keeps that connection between claim and source available for inspection.

ProvenanceGuard runs after a black-box MCP agent produces an answer and never collapses the evidence into one anonymous context. According to the blog post, it carries source identity through the pipeline by reading the captured MCP trace, including tool outputs and their source IDs, without retraining the agent. It then performs five steps in sequence: breaking the answer into specific claims, finding the source most relevant to each claim, checking whether that source actually supports it, comparing that source with the one the answer names or implies, and emitting both a per-claim source verdict and a global allow-or-block decision for the answer.

For the experiments in the paper, the team used local models so the captured traces could be processed in a controlled, offline setup: MiniLM helps find the relevant source, a DeBERTa NLI verifier checks whether that source supports the claim, and a local language model helps break answers into claims. The verifier also checks literal values closely, so a number, date or identifier absent from the source cannot pass merely because the sentence sounds plausible. A calibrated decision step combines these signals, and blocked answers can go through RARR-style repair, a source-grounded revision or safe fallback, which the verifier then checks again.

The blog post says those named models are the setup that was evaluated rather than a requirement of ProvenanceGuard, and that the same steps could be adapted to hosted models, though a new setup would need its own testing and calibration. The reported results come from the local configuration, whose conservative decision policy the authors describe as suited to data-sensitive review, where getting the source right matters more than producing the fastest possible answer.

The system was tested on answers from a medical agent that had used patient records, research articles and other tools, giving 281 real traces to study. In the main test, human experts checked 361 claims from 40 answers set aside from the data used to develop the system. The paper reports that experts judged 139 claims should not pass and that ProvenanceGuard caught 138 of them, letting one through. It also held 67 claims the experts considered supported, sending them for review or repair. For claims with an identifiable source, it picked the right source about 86 percent of the time. Four other support checkers were run on the same claims, and the paper says ProvenanceGuard scored highest on its measure of how well a system catches claims that should be blocked.

The blog post says the method can also be used in other fields when an agent keeps a record of its tool outputs and source IDs.