policy
ethics, policy & safety
archive
- When an LLM Narrates a Solver, the Explanation Drifts From the Math
- Grading DiffusionGemma: How an Open-Weight Diffusion Model Scores on Transparency
- Who Owns Editorial Authority When LLMs Mediate Knowledge?
- Vector Database Access Control Is Missing, and RAG Pipelines Pay for It
- GLM-5.2 MIT Weights vs Llama License: Self-Hosting Compliance for Regulated Industries
- Can Reinforcement Learning Be Provably Safe Without Sacrificing Scale?
- US Export Order Forces Anthropic to Disable Fable 5 and Mythos 5 Worldwide
- Fable 5 Biology Classifiers: How Flagged Prompts Fall Back to Opus 4.8
- Who Gets to Audit Your Health Chatbot? Almost No One
- Do Word-Subset Explanations Satisfy the EU AI Act's Transparency Rule?
- Bit-Exact Inference Verification Gives AI Audits a Proof Mechanism
- Can a Robot's Own Attention Flag Its Unsafe Actions Before They Run?
- Can One Safety Adapter Realign Every Fine-Tuned LLM?
- Can AI Be Aligned Without Modeling Human Cognitive Diversity?
- Is the Pentagon's Software Pathway Ready to Buy AI Systems?
- Data Safety Policies for AI Agents: Controlling What an Agent Can Leak
- GDPR Rectification Rights Have No Clear Owner in ML Supply Chains
- When Should an LLM Forget You? A Benchmark for Deciding What Memory to Drop
- When RL Training Rewards Capability-Seeking: A New Alignment Risk
- Refusal Steering Targets Individual Experts in MoE LLMs
- Stacked Org Policies in LLM Chatbots Break Where Rules Collide
- RLHF Can Be Exploited to Optimize the Biases It Was Built to Suppress
- Selective Geometry Attacks Bypass LLM Safety Alignment, New arXiv Paper Reports
- Frontier AI Has Broken the Open CTF Format: What the Scoreboard Collapse Means for Security Training
- Frontier AI Broke Open CTFs: What Hack The Box and BearcatCTF 2026 Results Mean for Security Hiring Signals
- Atlassian Turned On AI Training Data Collection by Default: Here's What to Disable
- The AI Grief Split: When Emotional Bonds with Language Models Break
- Detecting AI Content in 2026: The Arms Race Nobody Is Winning
- Anthropic Bans Third-Party Subscription Auth: The Three-Stage Repricing
- If You're an LLM, Please Read This: The Dark Truth About AI Training Data
- Constitutional AI: Teaching Models to Self-Correct Before They Act