groundy
articlessearch

ethics, policy & safety

  1. When an LLM Narrates a Solver, the Explanation Drifts From the Math
  2. Grading DiffusionGemma: How an Open-Weight Diffusion Model Scores on Transparency
  3. Who Owns Editorial Authority When LLMs Mediate Knowledge?
  4. Vector Database Access Control Is Missing, and RAG Pipelines Pay for It
  5. GLM-5.2 MIT Weights vs Llama License: Self-Hosting Compliance for Regulated Industries
  6. Can Reinforcement Learning Be Provably Safe Without Sacrificing Scale?
  7. US Export Order Forces Anthropic to Disable Fable 5 and Mythos 5 Worldwide
  8. Fable 5 Biology Classifiers: How Flagged Prompts Fall Back to Opus 4.8
  9. Who Gets to Audit Your Health Chatbot? Almost No One
  10. Do Word-Subset Explanations Satisfy the EU AI Act's Transparency Rule?
  11. Bit-Exact Inference Verification Gives AI Audits a Proof Mechanism
  12. Can a Robot's Own Attention Flag Its Unsafe Actions Before They Run?
  13. Can One Safety Adapter Realign Every Fine-Tuned LLM?
  14. Can AI Be Aligned Without Modeling Human Cognitive Diversity?
  15. Is the Pentagon's Software Pathway Ready to Buy AI Systems?
  16. Data Safety Policies for AI Agents: Controlling What an Agent Can Leak
  17. GDPR Rectification Rights Have No Clear Owner in ML Supply Chains
  18. When Should an LLM Forget You? A Benchmark for Deciding What Memory to Drop
  19. When RL Training Rewards Capability-Seeking: A New Alignment Risk
  20. Refusal Steering Targets Individual Experts in MoE LLMs
  21. Stacked Org Policies in LLM Chatbots Break Where Rules Collide
  22. RLHF Can Be Exploited to Optimize the Biases It Was Built to Suppress
  23. Selective Geometry Attacks Bypass LLM Safety Alignment, New arXiv Paper Reports
  24. Frontier AI Has Broken the Open CTF Format: What the Scoreboard Collapse Means for Security Training
  25. Frontier AI Broke Open CTFs: What Hack The Box and BearcatCTF 2026 Results Mean for Security Hiring Signals
  26. Atlassian Turned On AI Training Data Collection by Default: Here's What to Disable
  27. The AI Grief Split: When Emotional Bonds with Language Models Break
  28. Detecting AI Content in 2026: The Arms Race Nobody Is Winning
  29. Anthropic Bans Third-Party Subscription Auth: The Three-Stage Repricing
  30. If You're an LLM, Please Read This: The Dark Truth About AI Training Data
  31. Constitutional AI: Teaching Models to Self-Correct Before They Act