groundy

agents & frameworks

  1. jun 21agentsDeep-Research Benchmarks Hide How Agents Fail at Open-Web Source Grounding
  2. jun 21agentsDSPy Ships Autonomous Prompt Optimization, but Judge Drift Is the Failure Mode
  3. jun 21agentsDo AI Agents Reach for Over-Privileged Tools When Simpler Ones Suffice?
  4. jun 21agentsWhen Should Multi-Agent Systems Use an Event Bus Instead of an Orchestrator?
  5. jun 20agentsCan Deontic Policy Rules Govern an AI Agent at Runtime?
  6. jun 15agentsDo Programming Languages Still Matter to Your AI Coding Agent?
  7. jun 15agentsWhy Production AI Agents Fail Silently and Your Logs Never Catch It
  8. jun 12agentsComputer-Use Agents Fabricate Success on 8 to 33 Percent of Long-Horizon Tasks
  9. jun 10agentsCan AI Agents Share Context Without a Central Coordinator?
  10. jun 10agentsWhy Skill Creation and Reward Optimization Collide in Agentic RL
  11. jun 10agentsWhen AI Agents Delegate Work, Your Observability Stack Goes Blind
  12. jun 09agentsBloomberg's Pomona Makes Small Automated Code Changes, Not Big Agent PRs
  13. jun 09agentsAgent Tool-Gating Moves From Prompt Rules to Learned Policies
  14. jun 09agentsMore Capable LLMs Cooperate Less in Zero-Cost Collaboration Tests
  15. jun 08agentsWhy Foundation Model Agents Pass Benchmarks but Fail in Production
  16. jun 07agentsCan AI Agents Repair Broken Network Configs? A New Benchmark Tests It
  17. jun 07agentsCan Self-Evolving AI Agents Drift Without a Human in the Loop?
  18. jun 06agentsCascading Hallucination in Agentic RAG: When One Bad Retrieval Poisons the Chain
  19. jun 04agentsWhen MCP Tool Descriptions Don't Match the Code, Agents Trust the Lie
  20. may 29agentsMulti-Agent LLM Coordination: Why Attention Steering Beats Full Broadcast
  21. may 29agentsDataClawBench: AI Agents Fail at Exploratory Financial Analysis Across 492 Tasks
  22. may 28agentsClaude Code Dynamic Workflows: Spawning 100 Parallel Subagents on Opus 4.8
  23. may 27agentsClaude Code, Cursor, Copilot: How Agentic Coding Assistants Get Weaponized as Attacker Shells
  24. may 23agentsSpecBench Exposes Reward Hacking in Long-Horizon Coding Agents
  25. may 18agentsLangGraph 1.2.0 Makes Error-Handler Resume Crash-Durable: With Conditions
  26. may 18agentsCrewAI vs AutoGen vs LangGraph 2026: The Real Trade-Off After Maintenance Mode
  27. apr 29agentsCouncil Mode Cuts Multi-Agent LLM Hallucination 35.9% at 4.2x Token Cost on HaluEval
  28. apr 24agentsCloudflare Agents Week Moved Sandbox Execution, Private Networking, and Memory to Network Primitives
  29. apr 23agentsDiversity Collapse in Multi-Agent LLM Systems: Structural Coupling, Not Topology, Breaks Open-Ended Ideation
  30. mar 27agentsInsForge: The Backend Framework Built for Agentic Applications
  31. mar 15agentsAI Agents That Actually Learn: The Architecture Behind Hindsight Memory
  32. feb 28agentsSuperpowers: The Agentic Framework Replacing Your Dev Process
  33. feb 27agentsHow AI Agents Remember: Memory Architectures That Work
  34. feb 18agentsFunction Calling Best Practices: LLMs That Actually Use APIs Correctly
  35. feb 12agentsCrewAI vs AutoGen: A Developer's Guide to Multi-Agent AI Frameworks
  36. feb 12agentsAre AI-Generated PRs Killing Open Source?
  37. feb 11agentsPydantic AI vs LangChain: A Developer's Guide to the New Generation of Agent Frameworks
  38. feb 11agentsHow to Build Your First Autonomous Coding Agent with OpenHands SDK