groundy

agents & frameworks

  1. jun 24agentsCan a Cryptographic Certificate Prove an AI Agent's Output Is Valid?
  2. jun 24agentsCrewAI vs AutoGen vs Microsoft Agent Framework: AutoGen's Merger Reframes the 2026 Choice
  3. jun 24agentsCan You Trust an LLM Judge to Grade an Agentic Data Analysis System?
  4. jun 24agentsDo LLM Agent Societies Develop Their Own Authority Hierarchies?
  5. jun 24agentsDo Retrieval Metrics Predict Tool-Use Agent Success? A Paper Says No
  6. jun 24agentsCan You Pinpoint Which Step Broke a Long-Horizon AI Agent?
  7. jun 21agentsDeep-Research Benchmarks Hide How Agents Fail at Open-Web Source Grounding
  8. jun 21agentsDSPy Ships Autonomous Prompt Optimization, but Judge Drift Is the Failure Mode
  9. jun 21agentsDo AI Agents Reach for Over-Privileged Tools When Simpler Ones Suffice?
  10. jun 21agentsWhen Should Multi-Agent Systems Use an Event Bus Instead of an Orchestrator?
  11. jun 20agentsCan Deontic Policy Rules Govern an AI Agent at Runtime?
  12. jun 15agentsDo Programming Languages Still Matter to Your AI Coding Agent?
  13. jun 15agentsWhy Production AI Agents Fail Silently and Your Logs Never Catch It
  14. jun 12agentsComputer-Use Agents Fabricate Success on 8 to 33 Percent of Long-Horizon Tasks
  15. jun 10agentsCan AI Agents Share Context Without a Central Coordinator?
  16. jun 10agentsWhy Skill Creation and Reward Optimization Collide in Agentic RL
  17. jun 10agentsWhen AI Agents Delegate Work, Your Observability Stack Goes Blind
  18. jun 09agentsBloomberg's Pomona Makes Small Automated Code Changes, Not Big Agent PRs
  19. jun 09agentsAgent Tool-Gating Moves From Prompt Rules to Learned Policies
  20. jun 09agentsMore Capable LLMs Cooperate Less in Zero-Cost Collaboration Tests
  21. jun 08agentsWhy Foundation Model Agents Pass Benchmarks but Fail in Production
  22. jun 07agentsCan AI Agents Repair Broken Network Configs? A New Benchmark Tests It
  23. jun 07agentsCan Self-Evolving AI Agents Drift Without a Human in the Loop?
  24. jun 06agentsCascading Hallucination in Agentic RAG: When One Bad Retrieval Poisons the Chain
  25. jun 04agentsWhen MCP Tool Descriptions Don't Match the Code, Agents Trust the Lie
  26. may 29agentsMulti-Agent LLM Coordination: Why Attention Steering Beats Full Broadcast
  27. may 29agentsDataClawBench: AI Agents Fail at Exploratory Financial Analysis Across 492 Tasks
  28. may 28agentsClaude Code Dynamic Workflows: Spawning 100 Parallel Subagents on Opus 4.8
  29. may 27agentsClaude Code, Cursor, Copilot: How Agentic Coding Assistants Get Weaponized as Attacker Shells
  30. may 23agentsSpecBench Exposes Reward Hacking in Long-Horizon Coding Agents
  31. may 18agentsLangGraph 1.2.0 Makes Error-Handler Resume Crash-Durable: With Conditions
  32. may 18agentsCrewAI vs AutoGen vs LangGraph 2026: The Real Trade-Off After Maintenance Mode
  33. apr 29agentsCouncil Mode Cuts Multi-Agent LLM Hallucination 35.9% at 4.2x Token Cost on HaluEval
  34. apr 24agentsCloudflare Agents Week Moved Sandbox Execution, Private Networking, and Memory to Network Primitives
  35. apr 23agentsDiversity Collapse in Multi-Agent LLM Systems: Structural Coupling, Not Topology, Breaks Open-Ended Ideation
  36. mar 27agentsInsForge: The Backend Framework Built for Agentic Applications
  37. mar 15agentsAI Agents That Actually Learn: The Architecture Behind Hindsight Memory
  38. feb 28agentsSuperpowers: The Agentic Framework Replacing Your Dev Process
  39. feb 27agentsHow AI Agents Remember: Memory Architectures That Work
  40. feb 18agentsFunction Calling Best Practices: LLMs That Actually Use APIs Correctly
  41. feb 12agentsCrewAI vs AutoGen: A Developer's Guide to Multi-Agent AI Frameworks
  42. feb 12agentsAre AI-Generated PRs Killing Open Source?
  43. feb 11agentsPydantic AI vs LangChain: A Developer's Guide to the New Generation of Agent Frameworks
  44. feb 11agentsHow to Build Your First Autonomous Coding Agent with OpenHands SDK