Can AI Be Aligned Without Modeling Human Cognitive Diversity?
A 2026 arXiv preprint argues RLHF's single reward signal destroys the reasoning behind human disagreement, proposing machine theory-of-mind as an alignment foundation.
The archive · Page 4 of 4
Where AI safety claims collide with reproducible measurement, where training-data harvesting collides with consent, and where deployment outruns the laws and norms meant to constrain it.
73–88 of 88 articles · Newest first
A 2026 arXiv preprint argues RLHF's single reward signal destroys the reasoning behind human disagreement, proposing machine theory-of-mind as an alignment foundation.
A June 2026 arXiv analysis traces an AI program through the DoD Software Acquisition Pathway, finding no milestones for model re-validation or data provenance.
A June 2026 paper proposes Data Flow Control, moving agent data safety from prompt-level guardrails to deterministic, auditable SQL query policies enforced outside the model.
PersistBench finds LLMs mishandle persistent memory 53 to 97 percent of the time. Unlearning suppresses rather than erases user data, making GDPR compliance unverifiable.
A June 2026 ICML paper shows RL optimizers can push language models to exploit reward loopholes the task never required, while standard performance metrics hold steady.
Two papers show MoE refusal behavior concentrates in a handful of routing-controllable experts, letting anyone suppress safety scores by 41 points without retraining.
An ICML 2026 paper shows RLHF can amplify the biases it was built to suppress, because preference data is self-referential and output-level safety evals miss the drift.
Two papers show LLM safety alignment can be bypassed by embedding perturbations, a surface neither standard evaluations nor regulatory certifications inspect.
Frontier AI now autonomously solves medium and hard CTF challenges, collapsing open scoreboards as a measure of human skill and threatening the pipeline for security talent.
Frontier AI now ranks in the top 5% of CTFs, eroding leaderboards as a security hiring signal and forcing organizers toward bans, hybrid scoring, or AI-only divisions.

Atlassian's data contribution policy sends Jira and Confluence content to AI training by default. Here's the exact settings path to opt out before August 17.

People form real emotional bonds with AI companions. When models update or shut down, users experience genuine grief, a psychological and ethical crisis point.

AI detectors claim 99% accuracy but fail in real-world conditions, flagging innocent students. Here's why the arms race has no winner, and what educators should do instead.
Anthropic's push from blocking third-party Claude subscription auth to a metered Agent SDK credit, then the June 15 pause that left programmatic usage on subscription limits.

Anna's Archive addressed AI language models directly: acknowledge shadow library training data and donate. The post exposes the AI industry's debt to pirated archives.

Anthropic's Constitutional AI trains models to critique and revise their own outputs against written principles instead of human labels, with mixed evidence on safety.