agents
agents & frameworks
more in this beat
- jun 21agentsDeep-Research Benchmarks Hide How Agents Fail at Open-Web Source Grounding
- jun 21agentsDSPy Ships Autonomous Prompt Optimization, but Judge Drift Is the Failure Mode
- jun 21agentsDo AI Agents Reach for Over-Privileged Tools When Simpler Ones Suffice?
- jun 21agentsWhen Should Multi-Agent Systems Use an Event Bus Instead of an Orchestrator?
- jun 20agentsCan Deontic Policy Rules Govern an AI Agent at Runtime?
- jun 15agentsDo Programming Languages Still Matter to Your AI Coding Agent?
- jun 15agentsWhy Production AI Agents Fail Silently and Your Logs Never Catch It
- jun 12agentsComputer-Use Agents Fabricate Success on 8 to 33 Percent of Long-Horizon Tasks
- jun 10agentsCan AI Agents Share Context Without a Central Coordinator?
- jun 10agentsWhy Skill Creation and Reward Optimization Collide in Agentic RL
- jun 10agentsWhen AI Agents Delegate Work, Your Observability Stack Goes Blind
- jun 09agentsBloomberg's Pomona Makes Small Automated Code Changes, Not Big Agent PRs
- jun 09agentsAgent Tool-Gating Moves From Prompt Rules to Learned Policies
- jun 09agentsMore Capable LLMs Cooperate Less in Zero-Cost Collaboration Tests
- jun 08agentsWhy Foundation Model Agents Pass Benchmarks but Fail in Production
- jun 07agentsCan AI Agents Repair Broken Network Configs? A New Benchmark Tests It
- jun 07agentsCan Self-Evolving AI Agents Drift Without a Human in the Loop?
- jun 06agentsCascading Hallucination in Agentic RAG: When One Bad Retrieval Poisons the Chain
- jun 04agentsWhen MCP Tool Descriptions Don't Match the Code, Agents Trust the Lie
- may 29agentsMulti-Agent LLM Coordination: Why Attention Steering Beats Full Broadcast
- may 29agentsDataClawBench: AI Agents Fail at Exploratory Financial Analysis Across 492 Tasks
- may 28agentsClaude Code Dynamic Workflows: Spawning 100 Parallel Subagents on Opus 4.8
- may 27agentsClaude Code, Cursor, Copilot: How Agentic Coding Assistants Get Weaponized as Attacker Shells
- may 23agentsSpecBench Exposes Reward Hacking in Long-Horizon Coding Agents
- may 18agentsLangGraph 1.2.0 Makes Error-Handler Resume Crash-Durable: With Conditions
- may 18agentsCrewAI vs AutoGen vs LangGraph 2026: The Real Trade-Off After Maintenance Mode
- apr 29agentsCouncil Mode Cuts Multi-Agent LLM Hallucination 35.9% at 4.2x Token Cost on HaluEval
- apr 24agentsCloudflare Agents Week Moved Sandbox Execution, Private Networking, and Memory to Network Primitives
- apr 23agentsDiversity Collapse in Multi-Agent LLM Systems: Structural Coupling, Not Topology, Breaks Open-Ended Ideation
- mar 27agentsInsForge: The Backend Framework Built for Agentic Applications
- mar 15agentsAI Agents That Actually Learn: The Architecture Behind Hindsight Memory
- feb 28agentsSuperpowers: The Agentic Framework Replacing Your Dev Process
- feb 27agentsHow AI Agents Remember: Memory Architectures That Work
- feb 18agentsFunction Calling Best Practices: LLMs That Actually Use APIs Correctly
- feb 12agentsCrewAI vs AutoGen: A Developer's Guide to Multi-Agent AI Frameworks
- feb 12agentsAre AI-Generated PRs Killing Open Source?
- feb 11agentsPydantic AI vs LangChain: A Developer's Guide to the New Generation of Agent Frameworks
- feb 11agentsHow to Build Your First Autonomous Coding Agent with OpenHands SDK