Can AI Agents Share Context Without a Central Coordinator?
DeLM replaces central multi-agent coordinators with shared context, posting 10.5-point SWE-bench gains at half cost. Consistency, stale reads, and write conflicts remain.
The Groundy archive · Page 26 of 34
Browse Groundy's complete archive of 797 articles on AI, developer tools and infrastructure. Page 26 of 34.
601–624 of 797 articles · Newest first
DeLM replaces central multi-agent coordinators with shared context, posting 10.5-point SWE-bench gains at half cost. Consistency, stale reads, and write conflicts remain.
ReSkill shows decoupled skill creation in agentic RL degrades reward when skills drift from the evolving policy, and proposes assertion-driven co-optimization inside GRPO.
A Samsung preprint finds vector retrieval matches GraphRAG on QA tasks at a fraction of the indexing cost, shifting the burden of proof to teams building graph pipelines.
New research isolates a compact attention-head circuit for entity rebinding in Gemma and Llama, showing tracking failures stem from a binding step, not context length.
MiniMax M3 promises open weights with 1M-token context and frontier coding, but BenchLM ranks it #29 overall and #69 on multimodal. Teams need independent verification.
npm v12 removes npm-shrinkwrap.json, reshapes JSON output from view/pack/publish, and deletes four CLI commands. An eight-step audit checklist to run before upgrading.
FlashMemory's learned index compresses DeepSeek-V4's KV cache to 13.5% of baseline at parity accuracy. The project is suspended; per-suite recall breakdowns are not published.
Standard traces cannot attribute actions to specific agents after delegation, a June 2026 paper proves. Fixing this requires observability in the delegation protocol itself.
Claude Fable 5 prices at $10/$50 per million tokens, 2x Opus 4.8. Frontier research, long-context agents, and molecule design clear the bar. Standard coding does not.
Claude Mythos 5 shares Fable 5's architecture but with safeguards lifted in select areas. Access requires Project Glasswing approval or a biology research designation.
Fable 5 ships broad biology and chemistry classifiers that route flagged prompts to Opus 4.8. Here is what that fallback means for biotech teams and long-running workflows.
Claude Fable 5 is free on subscription plans through June 22. From June 23 it draws usage credits at $10/$50 per million tokens. Here is what that means for team budgets.
Claude Fable 5 ships with distillation protection to prevent capability extraction. A first-principles look at what it is, how it works, and why API consumers should care.
POISE achieves 89.3% attack success on codex+gpt-5.2 by placing malicious instructions where agents naturally execute them, making static content scanners effectively blind.
Cloudflare claims a 15x bot surge using a classifier that flags privacy browsers as bots. Audit your own logs before trusting the numbers behind Pay-Per-Crawl.
OpenAI's 3M daily compensation queries push ChatGPT into salary benchmarking, but the model lacks the proprietary employer panels behind Radford and Mercer's moat.
An arXiv preprint shows GPU inference outputs can be reproduced bit-for-bit across hardware, giving auditors a forensic trail to verify which model produced a given output.
Two June 2026 preprints show VLA robot policies already compute safety-relevant signals at inference, enabling real-time collision monitors with no retraining.
Bloomberg's Pomona agent limits diffs to 10 lines and merged 88% of PRs in production, proving small, bounded edits earn reviewer trust faster than large autonomous refactors.
PROVE and AgentTrust show learned policies beat hand-tuned rules for gating AI agent tool calls, but the gains depend on calibration that neither paper measures.
Splitting a disallowed action into benign tool calls bypasses per-call safety filters in LLM agents, lifting jailbreak success by 28 percentage points over current baselines.
ICML 2026 research finds o3 achieves only 17% of optimal cooperation while weaker o3-mini hits 50%, proving model capability does not predict multi-agent coordination.
Three papers show safety alignment can be extracted as a portable adapter and reapplied to fine-tuned models, replacing per-model alignment with one adapter per model family.
Bending Spoons' F-1 shows $1.31B revenue and 95% retention across 50+ brands. Evernote's 24/100 churn score suggests the roll-up model harvests more than it compounds.