ethics, policy & safety
Top in ethics, policy & safety
Public Sector AI Procurement Must Shift from Model Certification to Task Authorization
Public sector AI governance breaks under general-purpose models. One-time procurement certification fails because model behavior is not fixed. Teams must shift to task-level.
policyWhy Written AI Policies Fail to Control Agent Behavior
HANDBOOK.md benchmark shows frontier agents follow 124-page policies on only 36.2% of trials. Governance requires runtime enforcement, not documentation.
Why Vendor Model Cards Fail Clinical Ethics Procurement
PrinciplismQA exposes a gap generic leaderboards miss. Hospital compliance teams must build domain-specific ethics evals to test autonomy versus beneficence conflicts before.
policyTRIDENT Benchmark: LLM Safety Gaps in Finance, Medicine, and Law
TRIDENT benchmarks LLM safety in finance, medicine, and law using professional ethics codes. Results show specialized models often fail subtle ethical tests that generalists.
policyEU AI Act Traceability: Why ML Pipelines Fail Conformity Assessment
ML teams shipping high-risk systems to the EU face rejection if they cannot produce lifecycle-wide traceability artifacts. A May 2026 survey shows these mechanisms are often.
policyImplicit Bias in LLMs Passes NYC and EU Audits
ImplicitBBQ finds open-weight LLMs carry 6x more implicit than explicit bias. NYC Local Law 144 audits measure explicit outcomes only, leaving a gap compliance teams must.
policyEU Driver Monitoring: GDPR Compliance Without Consent
EU safety mandates require always-on cabin cameras, but GDPR prohibits consent for non-disableable systems. OEMs must use legal obligation and on-device processing to avoid.
policyEU AI Act bans emotion AI in schools, but permits it where models fail
EU AI Act bans emotion recognition in schools, yet permits it elsewhere without requiring accuracy. Research shows current models fail to read emotions reliably, shifting the.
- jul 21policyHuggingFace vs GitHub Models vs Replicate: Policy Compliance for Uploaders
- jul 20policySAMark Text Watermarking: Paraphrase Robustness and the Policy Gap
- jul 17policyA Digital Twin Can Validate AV Safety, but No Regulator Accepts the Evidence
- jul 17policyWhy EU AI Act Monitoring Will Miss Discontinuous LLM Alignment Failures
- jul 08policyCan US Export Controls Contain AI Built Without American Chips?
- jul 08policyEntropy Regularization Buys RL Robustness That Certification Can't Credit
- jul 08policyFusion's ML Disruption Predictors Have No Shared Validation Standard
- jul 08policyAI-Generated CSAM Risks Expose Filter-First Safety Gaps
- jul 08policyTreating AI Governance as Code Moves Compliance Into the Build Pipeline
- jul 08policyGitHub Issues Are Now Where GDPR and CCPA Compliance Gets Decided
- jul 08policyDeepSeek V4 Peak-Load Pricing Breaks Continuous Access for API Users
- jul 08policyLARA Shifts Model Safety from Training to Decode-Time Constraints
- jul 07policyStochastic Dominance Reveals Where RLHF Safety Filters Hide Tail Risk
- jun 29policyUncertainty-Aware Reward Discounting Cuts Reward Hacking 93.6% in a Preprint
- jun 28policyMedical AI Liability Needs a Clinical Harness
- jun 28policyDoes More AI Regulation Actually Reduce Corporate Control?
- jun 27policyWhen an LLM Sets Your Price, Whose Long-Term Value Wins?
- jun 27policyCombining LLMs Doesn't Escape Shared Failures: A 67-Model Test
- jun 27policyTask-Focused VLMs Suppress Hazards They Detect in Isolation, June 2026 Preprint Finds
- jun 25policy50 Years of Aviation Certification Expose a Structural Gap in AI Governance
- jun 25policyDo Reasoning Tokens Actually Make LLMs Safer? A New Paper Tests It
- jun 24policyMachine-Readable AI Usage Terms: Does ODRL's Permission Model Hold Up?
- jun 24policyWho Audits the Safety Rules an LLM Agent Evolves for Itself?
- jun 24policyWhen Vibe-Coded Software Is Safety-Critical, Who Verifies It?
- jun 24policyCan You Trust an AI Robustness Certificate? A Paper Says Verify It
- jun 23policyCan a Benchmark Catch When AI Discharge Summaries Drop Care Steps?
- jun 23policyDo LLM Personality Tests Measure Anything? A New Paper Says No
- jun 23policyCommunity LoRA Mining Raises a Consent Gap for Style Generation
- jun 21policyWhen an LLM Narrates a Solver, the Explanation Drifts From the Math
- jun 21policyGrading DiffusionGemma: How an Open-Weight Diffusion Model Scores on Transparency
- jun 21policyWho Owns Editorial Authority When LLMs Mediate Knowledge?
- jun 21policyVector Database Access Control Is Missing, and RAG Pipelines Pay for It
- jun 19policyGLM-5.2 MIT Weights vs Llama License: Self-Hosting Compliance for Regulated Industries
- jun 15policyCan Reinforcement Learning Be Provably Safe Without Sacrificing Scale?
- jun 13policyUS Export Order Forces Anthropic to Disable Fable 5 and Mythos 5 Worldwide
- jun 10policyFable 5 Biology Classifiers: How Flagged Prompts Fall Back to Opus 4.8
- jun 09policyWho Gets to Audit Your Health Chatbot? Almost No One
- jun 09policyDo Word-Subset Explanations Satisfy the EU AI Act's Transparency Rule?
- jun 09policyBit-Exact Inference Verification Gives AI Audits a Proof Mechanism
- jun 09policyCan a Robot's Own Attention Flag Its Unsafe Actions Before They Run?
- jun 09policyCan One Safety Adapter Realign Every Fine-Tuned LLM?
- jun 08policyCan AI Be Aligned Without Modeling Human Cognitive Diversity?
- jun 08policyIs the Pentagon's Software Pathway Ready to Buy AI Systems?
- jun 07policyData Safety Policies for AI Agents: Controlling What an Agent Can Leak
- jun 07policyGDPR Rectification Rights Have No Clear Owner in ML Supply Chains
- jun 05policyWhen Should an LLM Forget You? A Benchmark for Deciding What Memory to Drop
- jun 05policyWhen RL Training Rewards Capability-Seeking: A New Alignment Risk
- jun 05policyRefusal Steering Targets Individual Experts in MoE LLMs
- jun 04policyStacked Org Policies in LLM Chatbots Break Where Rules Collide
- may 29policyRLHF Can Be Exploited to Optimize the Biases It Was Built to Suppress
AI safety is a moving target dressed up as a settled science. Vendors publish leaderboard scores from single-turn evals; independent researchers show that configuration choices flip those rankings, that multi-step agents drift past guardrails their one-shot tests never probe, and that “aligned” often means filtered rather than principled. This beat sits in that gap, treating alignment as an empirical claim that has to survive replication, not a marketing posture.
The same pattern repeats outside the model. Training-data pipelines depend on consent regimes that were never granted; default-on data collection settings turn enterprise tools into harvesters; shadow libraries underwrite frontier capability while their authors go uncompensated. Regulators respond unevenly: state laws fragment faster than federal frameworks consolidate, transparency rules hinge on tests like “average consumer” that courts will spend years defining, and disclosure obligations land on platforms with no safe harbor before the technical standards exist.
Coverage tracks the second-order effects too. Junior-developer pipelines hollow out when seniors lean on AI pair-programmers. Companion chatbots accrue real psychological weight, and model deprecations produce real grief. Content homogenization, detector arms races, and the steady automation of online discourse all sit downstream of decisions made in places that resist scrutiny. The throughline is principled skepticism, not panic. When a safety claim, a consent assumption, or a policy fix doesn’t survive contact with how systems actually behave, that gap is the story.