groundy
Agents & Frameworks

MCP Agents Cite the Wrong Source: Pooled vs Source-Aware Fact Checks

Pooled faithfulness checkers miss cross-source conflation in MCP agents. ProvenanceGuard adds per-claim source verification, though author-reported results remain unreplicated

Published 12 references
A scuffed translucent green dinosaur holds a yellow three-lobed piece above two resin slabs: one with a matching cavity and one with a square cavity, on an ivory background.
On this page14 sections

If your MCP agent answers from several tool outputs at once, a pooled faithfulness checker can pass a claim that credits the wrong source: the fact exists somewhere in the evidence, so the check succeeds even though the attribution is wrong. The fix is a per-claim, per-source verification layer keyed to your tool and source IDs. What follows is when pooled checking is enough, when it is not, and what the only published numbers are actually worth.

The failure: a true fact credited to the wrong tool output

The ProvenanceGuard paper from Ander Alvarez, Santhiya Rajan, Samuel Mugel, Román Orús, and colleagues at Multiverse Computing names this failure mode “cross-source conflation”: a claim that is supported somewhere in the pooled evidence while being attributed to the wrong source. The paper’s motivating example is a customer support agent that says: “According to the account record, this plan includes a 30-day refund window.” The refund window is real, but it lives in a policy document, not the account record.

Pool the account record and the policy document into one evidence blob and the claim checks out. Nothing about the pooled check can see that the sentence’s “according to” clause points at the wrong place. The customer hears a fabricated provenance for a true fact, and in a dispute, that provenance is what gets quoted back.

This is not a hypothetical stitched together for one paper. The Aegis agent-failure taxonomy found that exploitation failures, where the agent already holds the right information but uses it incorrectly, account for the majority of failures in nearly every workload studied: 75% in airline, 88% in retail, 60% in file system, and 33% in CRM under GPT 4.1. The authors trace the tool-output-processing subtype, its most common form in airline, retail, and CRM, to context distraction rather than an intrinsic limitation of the LLM; exploitation failures also include domain-rule and user-instruction errors. Misattribution is one flavor of that broader pattern, and it gets more likely as agents weave more tool outputs into a single answer.

Why pooled checkers pass it

RAGAS faithfulness, MiniCheck, AlignScore, and SummaC were designed to answer one question: is this claim supported by the available evidence? As the authors put it in their Hugging Face guest post, these systems “ask whether a claim is supported by the available evidence once that evidence has been pooled together,” and in their usual form they do not indicate which tool output supports each claim or whether it matches the source the answer names.

Worth being precise about the frame here: this is a scope gap, not a benchmark defeat. These checkers were never designed for attribution. Judging them on source identification is like judging a spellchecker on grammar. The gap only matters when your agent’s answers assert or imply a specific source, which is exactly what agents grounded in MCP tool outputs tend to do.

The failure mode also predates MCP. The fact-checking community already scores citation correctness (whether text attributed to a citation is entailed by the corresponding evidence) and citation completeness in the CLEF-2026 CheckThat! lab. And in regulatory compliance QA, a citation-closure study documents systemic attribution failures from the common pattern of appending citations as free-form footers after generation, arguing instead for schema-level binding between individual claims and their sources. Different fields keep rediscovering the same problem: support is not attribution.

What source-aware verification actually changes

ProvenanceGuard’s design, per the paper, treats attribution as a separate check with its own routing step:

  1. Consume captured MCP traces with stable tool IDs, source IDs, and raw tool outputs.
  2. Decompose the answer into atomic claims.
  3. Route each claim to source-specific evidence rather than the pool.
  4. Check support with natural-language inference plus an attention-derived token-alignment proxy.
  5. Separately compare the claim’s stated attribution against the routed source.

The output is different in kind, not just in score: per-claim claim-to-source verdicts plus an answer-level allow/block decision, with optional RARR-style repair that revises a blocked answer before re-verification. That per-claim record is the part pooled checkers structurally cannot produce, and it is the part an auditor actually wants.

The gate is deliberately conservative. The paper reports it removes nearly all held-out claims that should not pass but “also sends some supported claims to review or repair.” That tradeoff has a cost in reviewer time, and you should budget for it; a gate that blocks nothing is not a gate. The same conservatism shows up in citecheck, an MCP server for bibliographic verification that separates analysis from modification because a plausible-but-incorrect match can silently damage a manuscript.

The numbers, and who reported them

Every quantitative result here comes from the authors themselves, in the guest post, the paper, or the artifact repository. No independent replication exists in the available evidence, and the artifact repo states plainly that “reproducible code and data are not currently released.” The baselines were re-run by the paper’s team, not the original tool authors. Treat the table below as author-reported [unverified] until someone replicates it.

CheckerReject/block F1 (held-out)Emits claim-to-source IDs?
ProvenanceGuard0.802Yes
MiniCheck0.783No
RAGAS Faithfulness0.758No
AlignScore0.662No
SummaC-ZS0.436No

Setup details, all author-reported: evaluation on 281 medical-domain MCP-agent traces, with human experts checking 361 claims from 40 answers held out from development data. In controlled probes with 50 deliberately injected source-attribution swaps over frozen MCP evidence, the verifier caught all 50 with no retained wrong attribution.

Two more results discipline the headline. On a harder benchmark with several similar sources, block F1 rose to 0.846 but exact-source identification fell to 50.3% of claims. The authors themselves call telling similar sources apart an important area for improvement. Read that as: the gate can reliably say “this attribution is wrong,” but when two sources look alike, it coin-flips on which one is right. And the reported results come from a local model configuration; the authors note that adapting to hosted models “would need its own testing and calibration.”

The guest post also claims NVIDIA NVFlow merged an optional grounding-verification stage for its finance agent using this approach. That claim appears only in the authors’ own post [unverified], so do not build a business case on it.

When pooled checking is enough

The decision turns on one question: does the answer name or imply a specific source?

If your agent synthesizes across outputs without asserting provenance (“plans in this tier often include refund windows”), pooled faithfulness checking is doing the job it was built for. There is no attribution to verify. Keep RAGAS or MiniCheck as a coarse hallucination screen and spend your effort elsewhere.

If the answer asserts provenance (“according to the account record”), or if a downstream consumer will act differently depending on which record a fact came from, pooled checking is blind to the error class that matters. That includes most regulated workflows, but also ordinary support tooling, because “the account record says X” and “the policy doc says X” have different dispute paths even when X is identical. This mirrors the trust problem Groundy covered in description-code inconsistency: wherever the runtime has no independent way to check a claim about a tool, every unverified assertion runs on trust.

A practical middle position: run the pooled checker on everything, and route only source-asserting answers through the attribution layer. The conservative gate’s tendency to send some good claims to review is easier to absorb when it fires on a subset.

Integration points inside the MCP agent loop

Source-aware verification is not a library you import; it needs provenance plumbing first. Three integration points, in order of invasiveness:

Trace-time capture. The verifier consumes MCP traces with stable tool IDs and source IDs, so you need per-tool-execution provenance before anything else. PROV-AGENT shows one way: a @flowcept_agent_tool decorator creates an activity record for each tool execution, linked to its inputs and outputs via W3C PROV relationships, enabling queries about where erroneous data originated and how it propagated. If your tracing does not already bind outputs to tool call IDs, start here.

Post-generation gate. Existing MCP observability already records call sequences and can run evaluators per output; the telemetry-aware development paper describes a hallucination-watchdog pattern where an LLM subscribes to each new answer. A source-aware gate slots into the same surface: subscribe to completed answers, decompose, route, check, allow or block. This is ProvenanceGuard’s own operating point.

Generation-time binding. The most invasive option is to make attribution structural rather than checked after the fact, which is the CAMS argument covered next.

The by-construction alternative and its bill

CAMS (arXiv:2606.23989) argues attribution should be a structural property of generation, not a downstream prediction: claims are anchored to sources as the summary is built. The authors report staying ahead of an end-to-end baseline on AlignScore (.88 vs .74) and citation precision (79 vs 52). These are also author-reported numbers [unverified].

The cost accounting is unusually honest. The claim-anchored pipeline runs roughly 8x an end-to-end call and adds about 65 seconds of machine time per summary, while cutting human verification from 31 seconds to 9 seconds per claim. That trade breaks even at under three checked claims per summary, and holds only where the output is actually read for verification. CAMS explicitly targets regulated, legal, medical, and journalistic review, not high-throughput summarization.

So the spectrum is: pooled check (cheap, attribution-blind), post-hoc source-aware gate (moderate cost, conservative blocking, per-claim verdicts), and by-construction attribution (highest machine cost, lowest human verification cost, no retrofit path for existing generators). The right point depends on how expensive a wrong attribution is in your workflow, and how many claims a human would otherwise check.

Regulated workflows: attribution as an audit artifact

The second-order consequence of pooled verification is an audit gap. When an agent’s answer passes a faithfulness check, the passing score erases any record of attribution errors inside it. For a compliance reviewer, “the answer was supported” and “the answer cited the right record” are different assertions, and only the second survives a dispute.

The stakes are rising with the tool population. A formal security analysis of the MCP ecosystem cites, second-hand, over 10,000 active servers, 177,000 registered tools, and 97 million monthly SDK downloads as of early 2026, with action-capable tools growing from 27% to 65% of all tools between November 2024 and February 2026. Those figures are cited within that paper rather than measured by it, but the direction is consistent with what Groundy found in enterprise MCP filtering: the governance burden is shifting to whoever owns the allowlist, and now to whoever owns the verification layer. Per-claim claim-to-source verdicts, emitted as structured records, are the difference between “we checked” and “here is what we checked.”

What still breaks

Three limitations should shape how much you trust any of this today. First, exact-source identification at 50.3% on similar sources means the verifier’s strongest current skill is detecting that an attribution is wrong, not telling you which source was right. Second, the entire quantitative case rests on 281 medical-domain traces with no released code or data, a single testbed domain, and baselines re-run by the proposing team. Third, failure-mode taxonomies in this area are generally easier to publish than to operationalize; the Model or Harness? taxonomy notes that judge-based validation is hard to deploy in production because judge accuracy remains limited, a caveat that applies to any verifier built on model judgments, including this one.

Bottom line for teams shipping Monday

Keep your pooled faithfulness checker for coarse hallucination screening; it was built for that, and the author-reported F1 gap (0.802 vs 0.783 for MiniCheck) is not the reason to switch, since the paper’s own paired trace bootstrap reports the 0.019 difference as not statistically significant (two-sided p≈0.26). The reason to add a source-aware layer is the output type: per-claim claim-to-source verdicts keyed to your MCP tool and source IDs, which pooled checkers cannot emit at any score. Add it wherever answers name or imply a specific record, policy, or document, and archive the verdicts as audit artifacts in regulated workflows. Pilot on your own traces before trusting the reported F1, because every ProvenanceGuard number in this article is author-reported, unreplicated, and drawn from one medical-domain testbed, and the technique’s hardest case, telling similar sources apart, is a coin flip even by its own authors’ account.

Frequently Asked Questions

When is pooled faithfulness checking sufficient for MCP agents?

If your agent synthesizes across outputs without asserting provenance (“plans in this tier often include refund windows”), pooled faithfulness checking is doing the job it was built for. There is no attribution to verify. Keep RAGAS or MiniCheck as a coarse hallucination screen and spend your effort elsewhere.

What is the cost difference between CAMS and end-to-end generation?

The claim-anchored pipeline runs roughly 8x an end-to-end call and adds about 65 seconds of machine time per summary, while cutting human verification from 31 seconds to 9 seconds per claim. That trade breaks even at under three checked claims per summary, and holds only where the output is actually read for verification.

How accurate is ProvenanceGuard at identifying the correct source when sources are similar?

On a harder benchmark with several similar sources, block F1 rose to 0.846 but exact-source identification fell to 50.3% of claims. The authors themselves call telling similar sources apart an important area for improvement. Read that as: the gate can reliably say “this attribution is wrong,” but when two sources look alike, it coin-flips on which one is right.

References

Follow the links in the article for context. The supporting material is collected here for further reading.

  1. ProvenanceGuard paperarxiv.orgAccessed
  2. Aegis agent-failure taxonomyarxiv.orgAccessed
  3. Hugging Face guest posthuggingface.coAccessed
  4. CLEF-2026 CheckThat! labarxiv.orgAccessed
  5. citation-closure studyarxiv.orgAccessed
  6. citecheckarxiv.orgAccessed
  7. artifact repogithub.comAccessed
  8. PROV-AGENTarxiv.orgAccessed
  9. telemetry-aware development paperarxiv.orgAccessed
  10. CAMSarxiv.orgAccessed
  11. formal security analysis of the MCP ecosystemarxiv.orgAccessed
  12. Model or Harness? taxonomyarxiv.orgAccessed

Join the discussion

Share a useful perspective or ask a question about this article.

Discussion guidelinesComments privacy