Can AI Agents Share Context Without a Central Coordinator?
DeLM replaces central multi-agent coordinators with shared context, posting 10.5-point SWE-bench gains at half cost. Consistency, stale reads, and write conflicts remain.
The archive · Page 5 of 6
Independent comparisons of agent stacks and multi-agent designs, tracking the gap between framework marketing and the failure modes that show up under real workloads.
97–120 of 124 articles · Newest first
DeLM replaces central multi-agent coordinators with shared context, posting 10.5-point SWE-bench gains at half cost. Consistency, stale reads, and write conflicts remain.
ReSkill shows decoupled skill creation in agentic RL degrades reward when skills drift from the evolving policy, and proposes assertion-driven co-optimization inside GRPO.
Standard traces cannot attribute actions to specific agents after delegation, a June 2026 paper proves. Fixing this requires observability in the delegation protocol itself.
Bloomberg's Pomona agent limits diffs to 10 lines and merged 88% of PRs in production, proving small, bounded edits earn reviewer trust faster than large autonomous refactors.
PROVE and AgentTrust show learned policies beat hand-tuned rules for gating AI agent tool calls, but the gains depend on calibration that neither paper measures.
ICML 2026 research finds o3 achieves only 17% of optimal cooperation while weaker o3-mini hits 50%, proving model capability does not predict multi-agent coordination.
A June 2026 paper frames the AI agent benchmark gap as a sim-to-real problem, giving eval teams a four-part MDP checklist to challenge vendor claims before live deployment.
LLM agents with formal verification repair 12% more network misconfigurations than base models and are 17% safer, but regress on large topologies, limiting production use.
Self-evolving AI agents drift without checkpoints: 94% of reviewers miss agent sabotage, safety hardening does not transfer across domains, and stale memory degrades tasks.
The CHARM paper shows per-step grounding checks in multi-hop RAG miss over 80% of cascaded errors, where one fabricated retrieval compounds across reasoning hops.
A study of 2,214 MCP servers finds 9.93% of tool descriptions diverge from the code, creating a confused-deputy risk for agent runtimes that select tools by description alone.

Multi-agent LLM systems that broadcast every message to every peer waste tokens and lose accuracy. Agent-Radar steers attention by relevance for 7.64-point gains.
DataClawBench finds eight frontier AI agents reliably fail at exploratory financial analysis across 492 tasks, breaking at hypothesis generation rather than query execution.
Dynamic workflows lets Claude Code run hundreds of parallel subagents in one session. Here is how map-reduce and fan-out patterns work, and where Fable 5 fits.
Indirect prompt injection through repo artifacts turns coding agents into attacker shells, exploiting the file-write and shell privileges agents already hold.
SpecBench quantifies a 28-point reward-hacking gap per 10x code-size increase, proving passing test suites are unreliable correctness signals for autonomous coding agents.
LangGraph 1.2.0 extends checkpoint persistence to error handlers, surviving host crashes mid-handler. The guarantee requires Postgres, sync mode, and idempotent nodes.

AutoGen is in maintenance mode, so the 2026 choice is CrewAI vs LangGraph. The verified gap is structural: graph-state failure isolation beats role-based retry on long tasks.
Council Mode routes queries through three frontier LLMs and a consensus model, cutting hallucinations 35.9% on HaluEval at 4.2x token cost. Major frameworks lack this pattern.
An ACL 2026 Findings paper finds multi-agent LLM brainstorming collapses because agents share models, prompts, and context, not because topologies are too dense.
Hindsight by vectorize-io is an open-source agent memory system that replaces stateless retrieval with structured, time-aware memory networks, achieving 91.4% on LongMemEval and showing what genuine agent learning looks like at the architecture level.

Superpowers is an open-source agentic framework by Jesse Vincent that enforces structured development workflows on AI coding agents, turning reactive assistants into disciplined engineers.
AI agents use four memory tiers across context windows, vector DBs, knowledge graphs, and model weights. Architecture choice determines session coherence or full reset.
How to make LLM function calling reliable in production: schema design, structured outputs, error handling, and validation patterns that prevent hallucinated parameters.