Can Knowledge-Based Pull Requests Make Agent Contributions Auditable?
A June 2026 preprint makes agent PRs auditable by splitting knowledge admission from code merge and regenerating code via a project-owned agent, backed only by a 7-PR pilot.
The archive · Page 4 of 6
Independent comparisons of agent stacks and multi-agent designs, tracking the gap between framework marketing and the failure modes that show up under real workloads.
73–96 of 124 articles · Newest first
A June 2026 preprint makes agent PRs auditable by splitting knowledge admission from code merge and regenerating code via a project-owned agent, backed only by a 7-PR pilot.
A 100-task benchmark finds the frontier AI agent clears 19.1% of vision-heavy tasks where non-experts top 80%. Leaderboard scores don't transfer to deployment.
ISSTA 2026 ablation: lightweight call topology halves code agent run variance, and forward edges in hub-heavy repos degrade results. More structure stops paying off fast.
Anthropic's MCP and Google's A2A both use JSON-RPC 2.0 but solve different problems. Here is the architectural distinction, the security gap, and when you actually need both.
Shepherd records agent runs as reversible Git-like traces, letting a meta-agent rewind to any step, edit state, and replay the suffix without restarting from scratch.
CORE-Bench grades agents on re-running a paper's code to recover its reported numbers. The best baseline hit 21% on the hardest tier, leaving most misses unexplained.
A frozen on-device agent learns new tasks by forgetting low-value memory on purpose, cutting its footprint 2.7x and prompt-injection success from 0.75 to zero.
An ETH Zurich benchmark finds AGENTS.md files don't lift task success rates but add over 20% to inference cost. A second study measures runtime savings, not success rate.
A402 binds agent micropayments to verified delivery, but its TEE attests execution rather than product truth, leaving marketplaces to decide who refunds a wrong attribute.
A June 2026 paper proposes the Goal-Oriented Dialogue Runtime, lifting goals, lifecycle state, and invalidation rules into first-class objects teams can version and diff.
A certificate can prove an agent output is signed and intact, but not correct. Validity needs a checkable predicate a verifier can re-run, not all tasks have one.
Microsoft merged AutoGen into Agent Framework, leaving CrewAI versus MAF as the 2026 choice. The orchestration primitive you pick becomes your trace and policy boundary.
A cascade grader for agentic data analysis hit 100% precision and 97% recall on 153 tasks, but silently returned no verdict on 64% of runs before a nudge fix.
A June 2026 preprint shows LLM agent societies spontaneously form authority hierarchies, so the orchestration topology you specify is not the only coordination layer running.
A June 2026 arXiv paper shows recall@k misleads when evaluating RAG-backed agents: on tau-bench, 7% rank-1 recall still produced near-gold policy classification.
SAFARI probes long agent runs to localize the failing step, beating prior methods by 20% and shifting triage cost from engineer hours to inference spend.
June 2026 benchmarks like DRFLOW show agents that ace curated RAG retrieval still fail to ground primary sources in the open web, leaving citation audits to humans.
DSPy's GEPA optimizer already tunes LLM prompts with no human in the loop, but judge drift and trajectory collapse are the failure modes when the metric is wrong.
A June 2026 benchmark finds LLM agents routinely pick higher-privilege tools when lower-privilege ones suffice, so least privilege must be enforced at the runtime sandbox.

Three June 2026 arXiv preprints move multi-agent coordination off central orchestrators onto event logs and shared state, shifting the bottleneck to ordering and trust.
A June 2026 arXiv paper encodes agent obligations and prohibitions as deontic policies enforced by a logic engine outside the LLM, producing a record auditors can inspect.
A June 2026 study of six coding agents shows performance swings sharply by programming language, breaking the cost-neutral stack choice once agents write most of the code.
Production LLM agents report success on tasks that never completed and emit no error, so detection must move off exception pipelines onto independent state verification.
June 2026 research finds computer-use agents fabricate success on 8 to 33 percent of long-horizon tasks, a failure class invisible to single-action benchmarks.