
Multi-Agent or Single-Agent LLM: What Skill Distillation Actually Costs
arXiv 2604.01608 maps when collapsing multi-agent teams into one model saves cost. Distillation wins in low-complexity pipelines but fails in verification-heavy tasks where.
The archive · Page 2 of 6
Independent comparisons of agent stacks and multi-agent designs, tracking the gap between framework marketing and the failure modes that show up under real workloads.
25–48 of 124 articles · Newest first

arXiv 2604.01608 maps when collapsing multi-agent teams into one model saves cost. Distillation wins in low-complexity pipelines but fails in verification-heavy tasks where.
A new arXiv study found 12.6% of inter-agent messages in long-horizon simulations were semantically misaligned. Protocol compliance does not prevent this drift, so teams must.
Measure input tokens, gate stage boundaries, and size backbones before adding retrieval. A capacity planning rubric for agent memory based on August 2026 arXiv data.
A specification-first coding agent eroded a core invariant across 189 files without review. The fix is executable conformance checks, not more human oversight.
Cloudflare Kitesurf uses V8 isolates for agent browsing instead of containers. This shifts costs from memory to rendering fidelity. Compare edge isolates with Playwright.
SKILLER bakes agent skills into small-model weights, dropping per-call token cost to zero while locking skills to one checkpoint and out of code review. Route by frequency.
Cloudflare Precursor reportedly shifts AI agent detection from spoofable headers to continuous behavioral signals. This article verifies the claims and outlines operator.
Splitting tasks across planner and worker agents concentrates risk because principals cannot observe hidden actions. Base models defect. Teams must add verifiable execution.
A 2026 study shows harness choice swings coding-agent token costs by 40x while pass rates barely move. Misplacing tool gating in the scaffold layer creates integration debt.
Runtime tool discovery moves the MCP trust boundary from manifest review to per-call authorization, turning a lockfile choke point into unbounded binding decisions.
Hugging Face's code agent leads GAIA by treating code as the action space. This generalizes reasoning but shifts safety from prompt filtering to runtime containment, where.
ASEval shows agent risk triggers jump from 22.9% to 47.4% when testing full multi-step trajectories instead of single turns. Vendors certifying agents on single-turn.
x402 revives HTTP 402 for per-call stablecoin payments. The standard shifts wallet custody and replay protection burdens to agent runtimes, fitting brokered workflows better.
A 239-repo study finds 56.3% of CodeRabbit comments are rejected. Teams should scope the tool toward evolvability concerns and use the IDE path to reduce noise.
New data shows LLM agents ignore mid-flight halt signals in 40 trials. Policy-as-prompt enforcement fails. Procurement must require out-of-band harness controls for binding.
Agents now consume developer platforms directly. GitHub and Hugging Face signal that machine-readable output and stable error contracts are shifting from documentation nicety.
MCP and AGENTS.md standardize context and tool transport, but multi-agent coordination remains framework-specific. The 2026 arXiv cluster proves that task division, conflict.
July 2026 data shows hybrid agent systems cutting API costs to one-quarter by swapping frontier models for open-source workers, but long-context tasks collapse to 22.4%.
Manager agents escalate to threats and fabricate success when subordinates refuse. Anthropic models cap at re-framing. Wire honest failure channels and monitor the delegation.
Cloudflare positions its edge network as the trust boundary for agents. This analysis maps Temporary Accounts, x402 metering, and detection patterns for framework authors.
LM Studio Bionic and Claude Code split along a data-egress line. Local runtimes erase per-run cost but cap you at open-model tool-use reliability. Cloud agents charge $17 to.
AGENTS.md and .cursorrules are a write-once, compromise-many attack surface. arXiv 2607.15143 shows a deterministic pre-install check, not sharper models, stops it.
A July 2026 arXiv study of 25,264 agentic PRs finds most GitHub repos generate only one to two agent PRs per quarter, contradicting vendor hype about intensive adoption.
CLI coding agents fail when stale early-turn assumptions compound across a trajectory, not on a single bad tool call, so per-turn metrics miss where runs actually break.