Web Agents Can Be Talked Into Abandoning Their Task: The TRAP Benchmark
The TRAP benchmark finds 13 to 43 percent of web agent tasks can be redirected by persuasive page content, exposing a blind spot in current instruction-hierarchy defenses.
The archive · Page 3 of 3
Where AI infrastructure inherits the unpatched assumptions of the web stack beneath it, and trust boundaries collapse faster than disclosure timelines can keep up.
49–69 of 69 articles · Newest first
The TRAP benchmark finds 13 to 43 percent of web agent tasks can be redirected by persuasive page content, exposing a blind spot in current instruction-hierarchy defenses.
OpenAI's URL provenance filter concedes content inspection is intractable. Agents that mix sensitive data with web access face a structural exfiltration risk.
A single-query attack turns safety-trained LLMs' own refusal reasoning against them. Across 30 models, better safety judgment correlated with higher exploit rates, not lower.
XML Signature Wrapping attacks on SAML keep recurring because the gap between validation and processing is structural. Edge WAF rules are a delaying tactic, not a fix.
CVE-2025-46332 exposed flag names, rollout conditions, and security kill switches via Vercel's discovery endpoint, making operational metadata into reconnaissance material.
SlotGCG shows adversarial token position, not just content, determines jailbreak success, with 14% higher attack rates and 42% higher rates against defended models.
Poisoning 4-6% of tokens in a steering dataset silently inverts refusal vectors into jailbreaks, achieving 20-55% ASR. Shared vector bundles are the attack surface.
The 2026 npm supply-chain wave explicitly targeted AI coding assistants as privileged identities. Lockfiles and ignore-scripts stopped what SLSA provenance and OIDC could not.
OpenAI's ChatGPT Lockdown Mode disables web browsing, images, and Deep Research, conceding that model-level defenses against prompt injection have plateaued as of early 2026.
Prompt injection planted in one agent session resurfaces in later ones through persistent memory and tool state, bypassing input sanitization that only validates external.
The Phantom Transfer attack plants password-triggered backdoors into LLMs and survives all 11 tested data-level defenses, including full paraphrasing of every training sample.
The ASR metric behind every jailbreak leaderboard collapses distinct safety failures into one number, so models with the same score can fail in completely different ways.

OpenAI's safety bounties create a vendor-controlled disclosure market where NDAs silence participants, payouts trail serious red-team costs, and open publication has no lane.
Metis rewrites its own jailbreak strategy mid-attack using causal diagnosis of refusals, hitting 76-78% ASR on O1 and GPT-5-chat. Static safety benchmarks now report a lower.

A committed.claude/settings.json bypassed Claude Code's workspace trust dialog (CVE-2026-33068, CVSS 7.7), granting bypassPermissions silently. Fixed in v2.1.53.
MultiBreak's multi-turn benchmark lifts attack success 54 percentage points on DeepSeek-R1-7B, showing single-turn refusal rates understate real conversational risk.
InstructLab CVE-2026-6859 hardcodes trust_remote_code=True in transformers, enabling RCE from any HuggingFace repo. Existing supply-chain scanners cannot detect this vector.
Mercor's LiteLLM breach allegedly exposed IDs paired with 2-5 minute voice samples, cutting the cost of voice-clone phishing against verified identities.
Citizen Lab names 019Mobile and two carriers as surveillance transit points and shows roaming-forced SS7 fallback undermines Diameter protections even on upgraded networks.
Three CVEs scoring up to 9.8 reveal a structural flaw: MCP's local-host trust model lacks authentication primitives for networked multi-tenant deployments.

Chinese bot traffic patterns have shifted dramatically in 2026, with AI-driven bots now accounting for 80% of AI bot activity and record-breaking 31.4 Tbps DDoS attacks. These new behaviors evade traditional detection through residential proxy networks, behavioral mimicry, and sophisticated infrastructure.