groundy
articlessearch

agents & frameworks

  1. BOUNDARY_SYNC: Why Multi-Agent Representational Coupling Is the New Coordination Failure Mode
  2. Do Multi-Agent RAG Systems Write Better READMEs Than One Agent?
  3. Agentic AI Turns Location Trails Into a Re-Identification Tool
  4. How a Human-Agent Team Lifts One Video Into 4D Interactions
  5. Can LLM Agents Learn Cooperation Laws From Embodied Play?
  6. Govern the Repo, Not the Agent: A New Risk Metric for AI-Native Code
  7. Can an AI Agent Catch Cryptographic Misuse Before It Ships? Chai Tests the Claim
  8. Can Spec-Driven Development Keep AI Coding Agents From Drifting?
  9. Can Knowledge-Based Pull Requests Make Agent Contributions Auditable?
  10. Do AI Agents Hold Up Outside Familiar Environments? A New Eval Says No
  11. How Much Repo Structure Does a Coding Agent Actually Need?
  12. MCP vs A2A: Two Agent Protocols, One Integration Layer Decision
  13. Can You Rewind an AI Agent Mid-Run? Reversible Traces Say Yes
  14. Can AI Agents Reproduce Published Research? CORE-Bench Tests It
  15. How On-Device AI Agents Keep Learning by Forgetting on Purpose
  16. Do AGENTS.md Files Actually Help Coding Agents? A New Benchmark Tests It
  17. Should AI Shopping Agents Pay Micro-Transactions for Verified Product Data?
  18. Can a Conversational Graph Compile Into a Goal-Oriented Dialogue Runtime?
  19. Can a Cryptographic Certificate Prove an AI Agent's Output Is Valid?
  20. CrewAI vs AutoGen vs Microsoft Agent Framework: AutoGen's Merger Reframes the 2026 Choice
  21. Can You Trust an LLM Judge to Grade an Agentic Data Analysis System?
  22. Do LLM Agent Societies Develop Their Own Authority Hierarchies?
  23. Do Retrieval Metrics Predict Tool-Use Agent Success? A Paper Says No
  24. Can You Pinpoint Which Step Broke a Long-Horizon AI Agent?
  25. Deep-Research Benchmarks Hide How Agents Fail at Open-Web Source Grounding
  26. DSPy Ships Autonomous Prompt Optimization, but Judge Drift Is the Failure Mode
  27. Do AI Agents Reach for Over-Privileged Tools When Simpler Ones Suffice?
  28. When Should Multi-Agent Systems Use an Event Bus Instead of an Orchestrator?
  29. Can Deontic Policy Rules Govern an AI Agent at Runtime?
  30. Do Programming Languages Still Matter to Your AI Coding Agent?
  31. Why Production AI Agents Fail Silently and Your Logs Never Catch It
  32. Computer-Use Agents Fabricate Success on 8 to 33 Percent of Long-Horizon Tasks
  33. Can AI Agents Share Context Without a Central Coordinator?
  34. Why Skill Creation and Reward Optimization Collide in Agentic RL
  35. When AI Agents Delegate Work, Your Observability Stack Goes Blind
  36. Bloomberg's Pomona Makes Small Automated Code Changes, Not Big Agent PRs
  37. Agent Tool-Gating Moves From Prompt Rules to Learned Policies
  38. More Capable LLMs Cooperate Less in Zero-Cost Collaboration Tests
  39. Why Foundation Model Agents Pass Benchmarks but Fail in Production
  40. Can AI Agents Repair Broken Network Configs? A New Benchmark Tests It
  41. Can Self-Evolving AI Agents Drift Without a Human in the Loop?
  42. Cascading Hallucination in Agentic RAG: When One Bad Retrieval Poisons the Chain
  43. When MCP Tool Descriptions Don't Match the Code, Agents Trust the Lie
  44. Multi-Agent LLM Coordination: Why Attention Steering Beats Full Broadcast
  45. DataClawBench: AI Agents Fail at Exploratory Financial Analysis Across 492 Tasks
  46. Claude Code Dynamic Workflows: Spawning 100 Parallel Subagents on Opus 4.8
  47. Claude Code, Cursor, Copilot: How Agentic Coding Assistants Get Weaponized as Attacker Shells
  48. SpecBench Exposes Reward Hacking in Long-Horizon Coding Agents
  49. LangGraph 1.2.0 Makes Error-Handler Resume Crash-Durable: With Conditions
  50. CrewAI vs AutoGen vs LangGraph 2026: The Real Trade-Off After Maintenance Mode