security
security
more in this beat
- jun 05securityOpenAI Adds Lockdown Mode to ChatGPT, Shifting Prompt-Injection Risk to Users
- jun 05securityActivation Steering Was Sold as LLM Control. New Work Makes It an Attack Surface
- jun 05securityCatching LLM Agents Leaking Credentials From Their Own Activations
- jun 05securityThe 2026 npm Attacks Proved AI Coding Assistants Are a Supply-Chain Target
- jun 04securityChatGPT's New Lockdown Mode Borrows Apple's Name for a Prompt-Injection Kill Switch
- jun 04securityStudents Are Prompt-Injecting AI Graders to Score Full Marks
- jun 04securityRemoving an LLM Backdoor Post-Training Without the Poisoned Data
- jun 04securityStored Prompt Injection Now Persists Across AI Agent Sessions
- jun 04securityLLM Data Poisoning Survives the Data-Cleaning Defenses Built to Stop It
- jun 03securityWhy OpenAI Bets on Instruction Hierarchy to Stop Prompt Injection
- jun 03securityStopping Multi-Turn LLM Jailbreaks Without Retraining the Model
- jun 03securityAfrican Languages Are a Jailbreak Blind Spot for English-Tuned LLM Safety
- jun 03securityPoisoning Open-Source LLM Merges: One Bad Checkpoint Hijacks the Result
- jun 03securityAn Autonomous Research Agent Now Discovers SOTA LLM Jailbreak Attacks
- jun 03securityMalware Can Prompt-Inject the AI Agent Reverse-Engineering It
- jun 03securityCVE-Factory Turns Published CVEs Into Security Agent Training Data. A 32B Model Beats Claude 4.5 Sonnet.
- jun 02securityLLM Reasoning Traces Leak the Private Data They're Told to Hide
- jun 02securityVideo Jailbreaks Hit Multimodal LLMs by Splitting Payloads Across Clips
- jun 01securityVercel AI SDK CVE-2025-48985: Input Validation Bypass Hits LLM App Builders
- jun 01securityHijacking AI Agent Memory: One Conversation Can Plant a Persistent Trojan
- jun 01securityWhy Attack Success Rate Misleads LLM Jailbreak Benchmarks
- may 31securityJob Seekers Are Prompt-Injecting AI Resume Screeners. New Study Measures the Hit Rate
- may 31securityWhy Audio Jailbreaks Slip Past the Safety Training Built for Text LLMs
- may 31securityLoRA Adapter Backdoors Generalize Beyond Their Trigger Tokens
- may 29securityThree Labs Concede Browser Agents Cannot Stop Prompt Injection
- may 29securityVercel Firewall Now Blocks SAMLStorm. Can an Edge WAF Fix a SAML Signature Flaw?
- may 27securityVercel Could Block React2Shell at the Edge. Its Next 13 CVEs Had No Shortcut.
- may 27securityOpenAI Adds a GPT-5 System Card Addendum on Sensitive Conversations
- may 27securityMCP Tool Description Poisoning: New Benchmark Shows Agents Trust Manuals That Lie
- may 27securityOpenAI's New Safety Bug Bounty Pays Researchers for Jailbreaks and Policy Bypasses
- may 27securityAxios npm Compromise Forces Vercel Into Platform-Level Remediation
- may 27securityNext.js Dev Server CVE-2025-48068: Any Web Page Could Read Your Source Files
- may 26securityApple Names Claude in CVE Credit Line, Setting Vendor Attribution Precedent
- may 25securityCISA's Internal Data Leak Tests the Disclosure Standards It Sets for Others
- may 25securityTanStack npm Attack: When OIDC Trusted Publishing Becomes the Attack Vector
- may 25securityNx s1ngularity Attackers Used Local Claude Code and Gemini CLI to Steal Developer Tokens
- may 24securityOpenAI Ships Lockdown Mode and Elevated Risk Labels for ChatGPT Sessions
- may 23securityAI Jailbreaks Are Now a Reasoning Problem, Not a Prompt Problem
- may 23securityJailbreak Defense Now Lives in Model Weights, Not in Prompt Filters
- may 23securityVercel Blocks Deploys With Vulnerable next-mdx-remote by Default: Platform Mitigation Outpaces the CVE Cycle
- may 23securityVercel's Next.js Middleware Bypass Postmortem: What the Fix Reveals About Edge Runtime Auth
- may 23securityOpenAI's New Agent Defense Post Concedes Prompt Injection Is Architectural, Not Patchable
- may 23securityWhen Stronger Backdoor Triggers Backfire: An arXiv Theory Paper Inverts a Core Defense Assumption
- may 18securityDPrivBench: LLMs Score 99.5% on Textbook DP but Collapse on Advanced Reasoning
- may 18securityCatching Graph Neural Net Backdoors by Influence, Not Pattern
- may 18securityTrustFall: One Keypress in Claude Code, Gemini CLI, Cursor, and Copilot CLI Triggers Unsandboxed RCE
- may 18securityMini Shai-Hulud Ships the First Malicious npm With Valid SLSA Provenance
- may 18securityMultiBreak Benchmark: 10,389 Multi-Turn Jailbreak Prompts Raise ASR 54pp on DeepSeek-R1-7B
- may 18securityNext.js CVE-2026-44578: WebSocket Upgrade SSRF Hits 79,000 Self-Hosted Instances From 13.4.13 Onward
- may 18securityPraisonAI CVE-2026-44338: Legacy Flask API Ships With AUTH_ENABLED=False, First Scan in 3h44m