Test-Time Scaling Cost Falls as PRMs Reuse Generator KV-Cache
KV-PRM cuts process reward model compute by reusing generator KV-caches instead of re-encoding text, reducing verification cost by up to 5,000x and making multi-agent.
The archive · Page 3 of 6
Independent comparisons of agent stacks and multi-agent designs, tracking the gap between framework marketing and the failure modes that show up under real workloads.
49–72 of 124 articles · Newest first
KV-PRM cuts process reward model compute by reusing generator KV-caches instead of re-encoding text, reducing verification cost by up to 5,000x and making multi-agent.
The July 2026 Skill Market paper defines reusable agent skills, but Claude Code, Cursor, and MCP encode skills in incompatible formats, forcing teams to pick a runtime before.
TTHE lets a coding agent rewrite its test rig during evaluation, raising coverage but blurring spec and verification. Benchmarks must now defend why their tests stay fixed.
GitLake layers git commits, branches, and merges over Apache Iceberg so agents write to reviewed branches and roll back bad outputs before they hit production tables.
Two July 2026 preprints show game-theoretic coordination can cut LLM hallucination, yet consensus breaks if one agent prioritizes cost, latency, or engagement over agreement.
WebSwarm's recursive multi-agent search beats flat ReAct on deep-and-wide benchmarks. Framework builders need spawn-and-merge primitives, not just larger context windows.
DeepSWE evaluates frontier agents on original, long-horizon tasks held out of GitHub, exposing when coding benchmarks measure memorized fixes instead of engineering skill.
A 4B subagent withdrawn from arXiv claimed to match frontier models on terminal execution and cut orchestrator tokens 30%. We explain the cost angle and how to test it.
A July 2026 preprint proposes Proof of Execution, a runtime attestation for every governed agent tool call. Audit receipts require instrumenting every tool invocation.
AgentTether models agent runs as a directed graph, detects drift, and steers execution back without retraining. On tau-bench Banking it repaired most failures and cut tokens.
Reasoning modes improve LLM negotiators as solvers but not as samplers, so multi-agent talks look inventive yet never agree; builders check moves against protocol rules.
A new travel-agent benchmark finds frontier models book animal-exploitation options below chance when the welfare preference is implicit. Fix the action space, not the prompt.
CHARLIE runs multi-agent RAG on-premise for forensic evidence, removing cloud LLM calls but shifting the bottleneck to GPU capacity, structured memory, and audit logging.
A July 2026 preprint reframes look-ahead bias as temporal non-interference, letting a type checker prove a pipeline leak-free before it runs rather than finding leaks after.
A July 2026 preprint says coding agents plan edits roughly 25 steps ahead before their latent program model collapses, so context length and pass@k may mismeasure failures.
FlowFixer achieves 71.3% workflow repair through symbolic inference, enabling deterministic rollback forcing agent frameworks to expose intermediate state rather than.
BOUNDARY_SYNC introduces the Coupling Amplification Factor to measure how inter-agent communication homogenizes multi-agent LLM systems, eroding diversity benefits.
An ICSME 2026 study finds single-agent RAG matches multi-agent README quality using 86% fewer tokens and half the latency, while developer-guided planning beats both.
A June 2026 arXiv preprint shows LLM agents re-identify people from anonymized location traces with no human analyst, naming 18 of 25 targets. Re-audit mobility datasets.
HAT-4D pairs a VLM agent with a human to lift one monocular video into 4D multi-object interactions, shifting embodied-AI data costs from capture rigs to feedback design.
LLawCo turns embodied agents' cooperation failures into readable rules fine-tuned into reasoning, making coordination policy an inspectable artifact engineers can edit.
A study of 930,000 agent-authored pull requests finds AI-native risk accumulates at the repository, not the agent, pushing governance to CI/CD and platform teams.
The Chai preprint reframes cryptographic misuse as a protocol-context recognition problem, claiming an LLM agent found a critical SSL-library flaw and 100-plus crypto bugs.
A June 2026 preprint makes spec-code divergence a blocking merge condition, treating traceability as a CI gate that catches silent drift before it compounds.