ethics, policy & safety
Where AI safety claims collide with reproducible measurement, where training-data harvesting collides with consent, and where deployment outruns the laws and norms meant to constrain it.
- Google Play vs GitHub: How Each Handles Baseless AI Copyright Claims
- LLM Judge Scores Change Run to Run: Why AI Audits Need Variance Reporting
- Hugging Face's AI Action Plan Reply: Open Weights vs the Frontier Lab Lobby
- Why Temperature 0 Won't Save Your Financial AI Audit
- PayPal Blocks GrapheneOS: A Payment Contingency Guide for Open Source
- Does the EU AI Act Exempt Open-Weight Models Like Qwen and Llama?
- AI Agents Outgrow OAuth: What Task-Scoped Authorization Requires
- AI Bias Audits vs Ethics Audits: What Each Actually Catches
- Cloudflare's 1-Click Fix for Vibe-Coded Apps: What It Doesn't Solve
- DiverValue-Bench: Measuring LLM Value Divergence Across 74 Markets
- Android Privacy Policies vs. Runtime Logs: A 0.4% Alignment Study
- Why Machine Unlearning Can't Certify GDPR Erasure
- RA-Bench: Why Deepfake Detectors Fail on Re-Shared Crisis Video
- Dario Amodei on AI Regulation: Frontier Labs as Their Own Lobbyists
- Public Sector AI Procurement Must Shift from Model Certification to Task Authorization
- Why Written AI Policies Fail to Control Agent Behavior
- Why Vendor Model Cards Fail Clinical Ethics Procurement
- TRIDENT Benchmark: LLM Safety Gaps in Finance, Medicine, and Law
- EU AI Act Traceability: Why ML Pipelines Fail Conformity Assessment
- Implicit Bias in LLMs Passes NYC and EU Audits
- EU Driver Monitoring: GDPR Compliance Without Consent
- EU AI Act bans emotion AI in schools, but permits it where models fail
- HuggingFace vs GitHub Models vs Replicate: Policy Compliance for Uploaders
- SAMark Text Watermarking: Paraphrase Robustness and the Policy Gap
- A Digital Twin Can Validate AV Safety, but No Regulator Accepts the Evidence
- Why EU AI Act Monitoring Will Miss Discontinuous LLM Alignment Failures
- Can US Export Controls Contain AI Built Without American Chips?
- Entropy Regularization Buys RL Robustness That Certification Can't Credit
- Fusion's ML Disruption Predictors Have No Shared Validation Standard
- AI-Generated CSAM Risks Expose Filter-First Safety Gaps
- Treating AI Governance as Code Moves Compliance Into the Build Pipeline
- GitHub Issues Are Now Where GDPR and CCPA Compliance Gets Decided
- DeepSeek V4 Peak-Load Pricing Breaks Continuous Access for API Users
- LARA Shifts Model Safety from Training to Decode-Time Constraints
- Stochastic Dominance Reveals Where RLHF Safety Filters Hide Tail Risk
- Uncertainty-Aware Reward Discounting Cuts Reward Hacking 93.6% in a Preprint
- Medical AI Liability Needs a Clinical Harness
- Does More AI Regulation Actually Reduce Corporate Control?
- When an LLM Sets Your Price, Whose Long-Term Value Wins?
- Combining LLMs Doesn't Escape Shared Failures: A 67-Model Test
- Task-Focused VLMs Suppress Hazards They Detect in Isolation, June 2026 Preprint Finds
- 50 Years of Aviation Certification Expose a Structural Gap in AI Governance
- Do Reasoning Tokens Actually Make LLMs Safer? A New Paper Tests It
- Machine-Readable AI Usage Terms: Does ODRL's Permission Model Hold Up?
- Who Audits the Safety Rules an LLM Agent Evolves for Itself?
- When Vibe-Coded Software Is Safety-Critical, Who Verifies It?
- Can You Trust an AI Robustness Certificate? A Paper Says Verify It
- Can a Benchmark Catch When AI Discharge Summaries Drop Care Steps?
- Do LLM Personality Tests Measure Anything? A New Paper Says No
- Community LoRA Mining Raises a Consent Gap for Style Generation
AI safety is a moving target dressed up as a settled science. Vendors publish leaderboard scores from single-turn evals; independent researchers show that configuration choices flip those rankings, that multi-step agents drift past guardrails their one-shot tests never probe, and that “aligned” often means filtered rather than principled. This beat sits in that gap, treating alignment as an empirical claim that has to survive replication, not a marketing posture.
The same pattern repeats outside the model. Training-data pipelines depend on consent regimes that were never granted; default-on data collection settings turn enterprise tools into harvesters; shadow libraries underwrite frontier capability while their authors go uncompensated. Regulators respond unevenly: state laws fragment faster than federal frameworks consolidate, transparency rules hinge on tests like “average consumer” that courts will spend years defining, and disclosure obligations land on platforms with no safe harbor before the technical standards exist.
Coverage tracks the second-order effects too. Junior-developer pipelines hollow out when seniors lean on AI pair-programmers. Companion chatbots accrue real psychological weight, and model deprecations produce real grief. Content homogenization, detector arms races, and the steady automation of online discourse all sit downstream of decisions made in places that resist scrutiny. The throughline is principled skepticism, not panic. When a safety claim, a consent assumption, or a policy fix doesn’t survive contact with how systems actually behave, that gap is the story.
