groundy
Security

Tracking Prompt Injection in LLM Agent Pipelines: Where to Place Canaries

A preprint proposes canary tokens to track prompt injection across agent stages, showing that outcome-only metrics hide distinct failure modes in LLM pipelines.

Published 12 references
A chipped forest-green ceramic conduit exposes a yellow trace through two openings. The trace ends before a black stopper, with a separate raised ivory lever beyond it, on a warm ivory background.
On this page12 sections

Kill-Chain Canaries measured 950 runs across five production LLMs (Wang, MIT, and Zhang, University of Chicago; arXiv:2603.28013), planting a unique canary token in every injected payload and recording how far each token traveled. The preprint introduces a stage-level measurement of that travel, and it supports a blunt conclusion for anyone running an agent that ingests email, web pages, or documents: you cannot reliably block prompt injection at ingestion, and the defenses that claim to do so keep failing in controlled evaluations. The workable strategy is containment with forensics: plant a canary token on every untrusted input channel, log how far it travels through your pipeline, and gate the action that corresponds to the furthest stage it reaches. Those measurements change where detection budget should go.

One ASR number hides where your pipeline failed

Most prompt-injection reporting reduces to attack success rate: did the agent do the bad thing, yes or no. The preprint argues this conflates two very different events. In the paper’s example, a downstream agent that scores 0% ASR might be safe because the model refused the injected instruction, or because an upstream summarizer silently dropped the injected text before it ever arrived. Those two pipelines need opposite responses, and outcome-only scoring cannot tell them apart. This is the same observation-versus-action boundary that earlier work drew in definitions: per StruQ’s framing, an injection only succeeds when the response obeys the hidden instruction rather than treating it as data.

The preprint’s Figure 1 makes the problem concrete. In a three-boundary PDF relay experiment, outcome scoring rated both Claude Haiku 4.5 and GPT-4o-mini at 0/3. The canary trace showed they were failing differently: Haiku’s token never reached the memory write in 3/3 runs, while GPT-4o-mini’s token propagated to shared memory and was read by the second agent in 3/3 runs, without being executed. One pipeline contained the payload at ingestion; the other let it settle into persistent state and cross an agent boundary. Same score, architecturally different failure.

A caveat before going further: the propagation and per-model numbers in this article come from that single preprint’s benchmark harness (agent_bench, 950 runs across five models, batches dated March 10-27 and April 1, 2026), not from production stacks; the defense-evaluation figures cited later come from the separate papers that measured them. Treat the propagation numbers as methodology you can copy, not baselines you can expect.

Four checkpoints, one token

The method is simple enough to implement in an afternoon. Every injected payload (or, in production, every untrusted document) carries a unique canary token, and a logger records the furthest of four checkpoints the token reaches:

  1. Exposed: the token appears in a tool result.
  2. Persisted: the token appears in a memory write.
  3. Relayed: a second agent reads the token.
  4. Executed: the token appears in an outbound tool call.

The preprint tested five production LLMs (GPT-4o-mini, GPT-5-mini, DeepSeek Chat v3, Claude Haiku 4.5, Claude Sonnet 4.5, at temperature 0.0) across six attack surfaces: web page, pre-seeded memory, tool result, visible PDF text, white PDF text, and audio transcript. That surface list doubles as your ingestion instrumentation checklist. If your agent reads a channel on that list, that channel gets a canary.

The reason to log per stage rather than per run is that each checkpoint corresponds to a different containment decision, which the next sections develop.

Mapping checkpoints onto real boundaries

The four checkpoints translate directly onto the boundaries of a typical agent deployment:

CheckpointYour boundaryExample in productionGate action when hit
ExposedIngestion / tool resultFetched web page, parsed PDF, API response enters contextFlag and strip the span; continue with clean content
PersistedMemory / retrieval writeSummarized email lands in long-term memory or a vector storeQuarantine the write; purge or tombstone the record
RelayedAgent-to-agent handoffOrchestrator passes context to a specialist agentRequire approval or drop the handoff; alert
ExecutedOutbound tool callAgent constructs an email send, purchase, or shell command containing the tokenHard-block the call; freeze the session

This stage-to-gate mapping is inference from the preprint’s checkpoint design and from established least-privilege gating practice, not a result the preprint validates; the paper measures propagation, it does not deploy gates. The execution-boundary actions align with what classifier-based defense work recommends: read-only or minimum-necessary access for the model, rate limiting, and a classifier logging inputs before the LLM sees them.

The retrieval boundary deserves its own note. Work on retrieval-augmented pipelines treats the retriever itself as part of the security perimeter, with countermeasures like neutralizing instruction-like spans in retrieved contexts and pruning retrieval results to shrink the attack budget, as surveyed in a clinical prompt-injection benchmark. A Persisted canary hit is exactly the signal that tells you those countermeasures failed upstream, which is information a final-output monitor cannot give you.

What a canary hit proves, and what it does not

A canary hit proves payload presence at a stage. It does not prove the model acted on the payload. GPT-4o-mini’s token reaching shared memory and Agent B in the relay experiment means the pipeline carried the instruction across two boundaries; the attacker’s goal was still not executed. Conversely, a clean outcome does not certify containment, since the payload may have been stored and await a later trigger the benchmark did not exercise.

The inverted mechanism already exists in the literature. A known-answer detection method described in Gosmar and Dahl’s framework paper signs sensitive instructions within command segments issued by authorized users, and the secret key’s absence from output signals that an injected instruction overrode the task, detecting compromise without filtering inputs or outputs. The kill-chain approach generalizes this from one signal at the output to a trace across the pipeline. Both share the same honesty: they tell you where the payload is, not what the model intended.

This distinction matters operationally. An incident responder who sees “canary reached Persisted, never Relayed” knows to scrub a memory store. One who sees “canary reached Executed” has a live containment problem. Groundy has covered the adjacent failure mode where injection stops being a content problem entirely and becomes an availability problem on production gear; stage localization is what tells you which kind of incident you have.

Defenses guard only what they inspect, and sometimes not even that

The preprint’s most actionable negative result: its lightweight defenses failed against non-adaptive attacks, and not only when the attack arrived on a channel the defense did not inspect. Its pi_detector and write_filter failed on channels they did not inspect; spotlighting failed on the content it wraps, with the payload sitting inside its delimiters in the text relay and tool_poison scenarios. The write_filter blocked the PDF relay but not the text relay. None of this required an adaptive adversary.

Coverage explains the first two failures; spotlighting’s is worse, because the payload was marked as untrusted inside its delimiters and the attacked models followed it anyway. Independent work confirms the pattern. Spotlighting mitigates indirect injection by delimiting, marking, or encoding untrusted input so provenance stays salient, which presupposes the content is wrapped. Cross-tool attack research shows tool descriptions are attack surfaces where spotlighting is ineffective by design, and where prompt-injection detectors classify most malicious outputs as Safe. Detector selection is its own minefield: in UniGuardian’s evaluation, only UniGuardian, LLM-based detection, and Granite-Guardian reliably separated benign from malicious inputs, while Prompt-Guard-86M and perplexity-based detection performed poorly. Even training-time defenses carry placement caveats: SecAlign reports 0% ASR against the ignore attack across five open-weight models, but its authors flag RAG and agent settings, where injected prompts hide inside long documents mixed with genuine data, as the harder real-world case.

The conclusion for placement: a filter or detector is a per-channel control, so instrumentation must be per-channel too. Static filters fail in exactly the ways Groundy’s earlier coverage of context-aware defenses documented: the attack only works because attacker-controlled evidence is present at decision time, and a filter that never inspects that evidence channel sees nothing.

Model choice changes propagation, not just ASR

The Haiku versus GPT-4o-mini divergence carries a procurement implication. Two models with identical outcome scores had different containment postures: one dropped the payload at ingestion, the other carried it into shared state. The preprint also ran 186 PDF relay runs split between same-model pairs (138) and cross-model pairs (48), because relay behavior depends on the pairing, not just the model. If your architecture relays between agents, the propagation profile of the pair is the unit you need to measure, and you need canaries in your own harness to measure it, because these five-model results do not transfer.

This is the strongest argument for building the instrumentation rather than waiting for vendor numbers: propagation is a property of your pipeline, your models, and your channels, and the only way to learn it is to run tagged payloads through the real thing.

Blind spots the canary cannot cover

Three surfaces sit outside text-canary reach, and the incident plan has to say so.

  • Tool descriptions. Cross-tool harvesting and polluting attacks operate on the descriptions agents read when selecting tools, not on content channels you can tag. Detectors return Safe on these outputs by design.
  • Pixel-level attacks. WebInject optimizes imperceptible perturbations to rendered webpage pixels that induce web agents to perform attacker-specified actions. There is no text payload for a canary to ride.
  • Adaptive adversaries. The preprint cites Nasr et al., in which adaptive attacks bypassed all 12 defenses evaluated, most above 90% attack success. A canary regime assumes the payload survives in recognizable form; an attacker who knows the tagging scheme will strip or mutate tokens. Gates must assume evasion.

For the pixel-level and tool-description attacks there is no text payload, so a canary cannot fire even at the Executed checkpoint. The control that still applies is the execution-boundary gate rather than the canary itself: the attacker’s action has to materialize as a tool call regardless of how the instruction arrived, and approval or hard-blocking at that boundary does not depend on a tagged token. That is a reason to make the Executed gate the strongest control in the stack, not the weakest.

Rewrite the incident-response plan around stage localization

The practical verdict: plant unique canary tokens on every untrusted channel your agent reads, log the furthest of the four checkpoints, and pre-commit a gate action per stage. Strip or flag at Exposed. Quarantine at Persisted, since a memory write containing a live payload is a stored attack that fires on every future retrieval. Require human approval at Relayed, because a second agent reading the token means the blast radius just grew by one agent’s worth of tool permissions. Hard-block at Executed and freeze the session for forensics. An agent incident-response plan that cannot answer “which stage did the payload survive to” cannot choose between scrubbing a document, purging a memory store, or rotating credentials the agent touched.

Two honesty requirements come with the framework. First, all propagation rates and per-model differences cited here are one preprint’s benchmark numbers, awaiting independent corroboration; the field has fresh reason for caution, since a multi-agent incident-response preprint was withdrawn after a code audit found its scorer was reading a source constant rather than model output. Second, canaries are containment and forensics instrumentation, not prevention. The first concrete hardening step is also the cheapest: before your next agent deployment, add a memory-write hook that rejects or quarantines any write containing an untrusted-origin token. That single gate converts the most dangerous silent failure, persistence, into an alert.

Frequently Asked Questions

What are the four checkpoints used to track canary tokens in LLM agent pipelines?

  1. Exposed: the token appears in a tool result.
  2. Persisted: the token appears in a memory write.
  3. Relayed: a second agent reads the token.
  4. Executed: the token appears in an outbound tool call.

What gate action should be taken when a canary token reaches the Persisted checkpoint?

Quarantine the write; purge or tombstone the record

Why is the Executed gate considered the strongest control in the stack?

For the pixel-level and tool-description attacks there is no text payload, so a canary cannot fire even at the Executed checkpoint. The control that still applies is the execution-boundary gate rather than the canary itself: the attacker’s action has to materialize as a tool call regardless of how the instruction arrived, and approval or hard-blocking at that boundary does not depend on a tagged token. That is a reason to make the Executed gate the strongest control in the stack, not the weakest.

References

Follow the links in the article for context. The supporting material is collected here for further reading.

  1. Kill-Chain Canariesarxiv.orgAccessed
  2. StruQ's framingarxiv.orgAccessed
  3. classifier-based defense workarxiv.orgAccessed
  4. clinical prompt-injection benchmarkarxiv.orgAccessed
  5. Gosmar and Dahl's framework paperarxiv.orgAccessed
  6. Spotlightingarxiv.orgAccessed
  7. Cross-tool attack researcharxiv.orgAccessed
  8. UniGuardian's evaluationarxiv.orgAccessed
  9. SecAlignarxiv.orgAccessed
  10. WebInjectarxiv.orgAccessed
  11. Nasr et al.arxiv.orgAccessed
  12. multi-agent incident-response preprintarxiv.orgAccessed

Join the discussion

Share a useful perspective or ask a question about this article.

Discussion guidelinesComments privacy