agents
agents & frameworks
archive
- BOUNDARY_SYNC: Why Multi-Agent Representational Coupling Is the New Coordination Failure Mode
- Do Multi-Agent RAG Systems Write Better READMEs Than One Agent?
- Agentic AI Turns Location Trails Into a Re-Identification Tool
- How a Human-Agent Team Lifts One Video Into 4D Interactions
- Can LLM Agents Learn Cooperation Laws From Embodied Play?
- Govern the Repo, Not the Agent: A New Risk Metric for AI-Native Code
- Can an AI Agent Catch Cryptographic Misuse Before It Ships? Chai Tests the Claim
- Can Spec-Driven Development Keep AI Coding Agents From Drifting?
- Can Knowledge-Based Pull Requests Make Agent Contributions Auditable?
- Do AI Agents Hold Up Outside Familiar Environments? A New Eval Says No
- How Much Repo Structure Does a Coding Agent Actually Need?
- MCP vs A2A: Two Agent Protocols, One Integration Layer Decision
- Can You Rewind an AI Agent Mid-Run? Reversible Traces Say Yes
- Can AI Agents Reproduce Published Research? CORE-Bench Tests It
- How On-Device AI Agents Keep Learning by Forgetting on Purpose
- Do AGENTS.md Files Actually Help Coding Agents? A New Benchmark Tests It
- Should AI Shopping Agents Pay Micro-Transactions for Verified Product Data?
- Can a Conversational Graph Compile Into a Goal-Oriented Dialogue Runtime?
- Can a Cryptographic Certificate Prove an AI Agent's Output Is Valid?
- CrewAI vs AutoGen vs Microsoft Agent Framework: AutoGen's Merger Reframes the 2026 Choice
- Can You Trust an LLM Judge to Grade an Agentic Data Analysis System?
- Do LLM Agent Societies Develop Their Own Authority Hierarchies?
- Do Retrieval Metrics Predict Tool-Use Agent Success? A Paper Says No
- Can You Pinpoint Which Step Broke a Long-Horizon AI Agent?
- Deep-Research Benchmarks Hide How Agents Fail at Open-Web Source Grounding
- DSPy Ships Autonomous Prompt Optimization, but Judge Drift Is the Failure Mode
- Do AI Agents Reach for Over-Privileged Tools When Simpler Ones Suffice?
- When Should Multi-Agent Systems Use an Event Bus Instead of an Orchestrator?
- Can Deontic Policy Rules Govern an AI Agent at Runtime?
- Do Programming Languages Still Matter to Your AI Coding Agent?
- Why Production AI Agents Fail Silently and Your Logs Never Catch It
- Computer-Use Agents Fabricate Success on 8 to 33 Percent of Long-Horizon Tasks
- Can AI Agents Share Context Without a Central Coordinator?
- Why Skill Creation and Reward Optimization Collide in Agentic RL
- When AI Agents Delegate Work, Your Observability Stack Goes Blind
- Bloomberg's Pomona Makes Small Automated Code Changes, Not Big Agent PRs
- Agent Tool-Gating Moves From Prompt Rules to Learned Policies
- More Capable LLMs Cooperate Less in Zero-Cost Collaboration Tests
- Why Foundation Model Agents Pass Benchmarks but Fail in Production
- Can AI Agents Repair Broken Network Configs? A New Benchmark Tests It
- Can Self-Evolving AI Agents Drift Without a Human in the Loop?
- Cascading Hallucination in Agentic RAG: When One Bad Retrieval Poisons the Chain
- When MCP Tool Descriptions Don't Match the Code, Agents Trust the Lie
- Multi-Agent LLM Coordination: Why Attention Steering Beats Full Broadcast
- DataClawBench: AI Agents Fail at Exploratory Financial Analysis Across 492 Tasks
- Claude Code Dynamic Workflows: Spawning 100 Parallel Subagents on Opus 4.8
- Claude Code, Cursor, Copilot: How Agentic Coding Assistants Get Weaponized as Attacker Shells
- SpecBench Exposes Reward Hacking in Long-Horizon Coding Agents
- LangGraph 1.2.0 Makes Error-Handler Resume Crash-Durable: With Conditions
- CrewAI vs AutoGen vs LangGraph 2026: The Real Trade-Off After Maintenance Mode