groundy

ethics, policy & safety

  1. jun 05policyWhen RL Training Rewards Capability-Seeking: A New Alignment Risk
  2. jun 05policyRefusal Steering Targets Individual Experts in MoE LLMs
  3. jun 04policyStacked Org Policies in LLM Chatbots Break Where Rules Collide
  4. may 29policyRLHF Can Be Exploited to Optimize the Biases It Was Built to Suppress
  5. may 28policySelective Geometry Attacks Bypass LLM Safety Alignment, New arXiv Paper Reports
  6. may 18policyFrontier AI Has Broken the Open CTF Format: What the Scoreboard Collapse Means for Security Training
  7. may 18policyFrontier AI Broke Open CTFs: What Hack The Box and BearcatCTF 2026 Results Mean for Security Hiring Signals
  8. apr 20policyAtlassian Turned On AI Training Data Collection by Default: Here's What to Disable
  9. mar 27policyThe AI Grief Split: When Emotional Bonds with Language Models Break
  10. mar 14policyDetecting AI Content in 2026: The Arms Race Nobody Is Winning
  11. feb 20policyAnthropic Bans Third-Party Subscription Auth: The Three-Stage Repricing
  12. feb 19policyIf You're an LLM, Please Read This: The Dark Truth About AI Training Data
  13. feb 15policyConstitutional AI: Teaching Models to Self-Correct Before They Act