groundy

all articles

  1. may 29agentsMulti-Agent LLM Coordination: Why Attention Steering Beats Full Broadcast
  2. may 29modelsPersona Prompts Change Who an LLM Recommends as an Expert
  3. may 29agentsDataClawBench: AI Agents Fail at Exploratory Financial Analysis Across 492 Tasks
  4. may 29policyRLHF Can Be Exploited to Optimize the Biases It Was Built to Suppress
  5. may 28ossFrontier AI Has Broken Open CTFs: Why Claude Code Now One-Shots Medium Pwn Challenges
  6. may 28policySelective Geometry Attacks Bypass LLM Safety Alignment, New arXiv Paper Reports
  7. may 28modelsOpus 4.8 Batch API: 1M Context, 300k Output, and Team Cost Controls
  8. may 28agentsClaude Code Dynamic Workflows: Spawning 100 Parallel Subagents on Opus 4.8
  9. may 27infraWhy LLMs Still Botch Kubernetes Manifests: The Training-Data Gap
  10. may 27securityOpenAI's New Safety Bug Bounty Pays Researchers for Jailbreaks and Policy Bypasses
  11. may 27industryHuggingFace's $100M Series C Bets Open-Source AI Can Outlast Per-Token Pricing Wars
  12. may 27infraGemma 4 31B on Cloud TPU vs GPU: The Serving Cost Crossover Point
  13. may 27agentsClaude Code, Cursor, Copilot: How Agentic Coding Assistants Get Weaponized as Attacker Shells
  14. may 26devtoolsBun Rewrites Its Core From Zig to Rust, Putting Downstream Zig Bindings at Risk
  15. may 26infraObjectCache Moves KV Reuse to S3-Class Storage: Why Layerwise Retrieval Beats Full-Prefix Cache Hits
  16. may 26devtoolsPromptArmor Shows Microsoft Copilot Cowork Can Be Tricked Into Exfiltrating Files
  17. may 25modelsμP Hyperparameter Transfer Has an Embedding Layer Hole, New arXiv Paper Says
  18. may 25devtoolsRmux Brings a Playwright SDK to tmux Sessions for Agent Automation Workflows
  19. may 25ossColorado SB051 Carves Out Open Source From Age Verification After Maintainer Backlash
  20. may 24cultureUS Researchers Hit With New Federal Limits on Publishing With Foreign Collaborators
  21. may 23securityAI Jailbreaks Are Now a Reasoning Problem, Not a Prompt Problem
  22. may 23devtoolsGoogle Sunsets Gemini CLI on June 18: Forced Migration to Antigravity CLI Breaks Existing Automation
  23. may 23agentsSpecBench Exposes Reward Hacking in Long-Horizon Coding Agents
  24. may 23infravLLM 0.21 Makes Prefill-Decode Disaggregation Actually Practical
  25. may 19industryBret Taylor's Sierra Raises $950M at $15B, Claims 40% of Fortune 50 Use Its Agents
  26. may 18industrySierra Raises $950M at $15B, Locking 40% of the Fortune 50 Into Its Agent Platform Before the Labs Go Direct
  27. may 18policyFrontier AI Has Broken the Open CTF Format: What the Scoreboard Collapse Means for Security Training
  28. may 18securityTrustFall: One Keypress in Claude Code, Gemini CLI, Cursor, and Copilot CLI Triggers Unsandboxed RCE
  29. may 18devtoolsClaude Code Adds Plugin Dependency Enforcement: disable Now Refuses to Break Transitive Chains
  30. may 18industryPayPal's $1.5B AI Overhaul Cuts 4,760 Jobs and Reframes Layoffs as Capex
  31. may 18policyFrontier AI Broke Open CTFs: What Hack The Box and BearcatCTF 2026 Results Mean for Security Hiring Signals
  32. may 18industryOpenAI's $4B Deployment Company Buys Tomoro and Signs 19 Partners to Own Implementation
  33. may 18agentsLangGraph 1.2.0 Makes Error-Handler Resume Crash-Durable: With Conditions
  34. may 18agentsCrewAI vs AutoGen vs LangGraph 2026: The Real Trade-Off After Maintenance Mode
  35. may 18cultureApple's $250M Siri Settlement: iPhone 16 Buyers Get $25 to $95 for Undelivered AI
  36. may 18securityMultiBreak Benchmark: 10,389 Multi-Turn Jailbreak Prompts Raise ASR 54pp on DeepSeek-R1-7B
  37. may 18industryAnthropic's $1.5B Joint Venture With Goldman Sachs and Blackstone Sends Claude Into PE Portfolio Companies
  38. may 18industryOpenAI Offers Two Months of Free Codex to Enterprises Switching From Claude Within 30 Days
  39. may 18cultureAB 566 Forces Chrome and Safari to Ship Opt-Out Signals by 2027. It Shields Them from Google's 86% GPC Failure
  40. may 18ossBrowserAct Open-Sources Stealth Browser Engine with 93% Token Reduction Claim
  41. may 18devtoolsGitHub Copilot's Opus 4.7 Multiplier: 7.5x to 15x to 27x in 60 Days
  42. apr 29osspgBackRest Is No Longer Maintained: PostgreSQL Backup Alternatives After the Project Stalls
  43. apr 29securityInstructLab CVE-2026-6859: Hardcoded trust_remote_code=True Turns Any HuggingFace Model Into RCE
  44. apr 29devtoolsPydantic AI v1.87 Closes the LangGraph Gap: Deferred Tool Calls, OpenTelemetry Eval, Stateful Compaction
  45. apr 29securityMercor's 4TB Lapsus$ Breach Hands Voice-Clone Attackers 40,000 Pre-Verified Targets
  46. apr 29agentsCouncil Mode Cuts Multi-Agent LLM Hallucination 35.9% at 4.2x Token Cost on HaluEval
  47. apr 29devtoolsClaude Code vs Cursor vs Copilot After the April 2026 Reshuffle: How the Comparison Math Changed
  48. apr 28devtoolsGitHub Copilot Replaces Premium Request Units With Token-Metered AI Credits on June 1
  49. apr 28ossfree-claude-code Routes Claude Code Through NVIDIA NIM and Local Models After Anthropic's CLI Ban
  50. apr 28industryMicrosoft and OpenAI End Their Exclusive Revenue-Sharing Deal: What It Means for Azure's AI Moat
  51. apr 28industryAnthropic Ends Flat-Fee Enterprise Claude, Enforces Per-Token Billing
  52. apr 24devtoolsGitHub CLI v2.91.0 Turns On Default Telemetry: What gh Collects and How to Opt Out in CI and Agent Pipelines
  53. apr 24devtoolsGitHub Copilot Drops Opus from Pro and Pauses Signups: The Forced Migration Facing Agentic Workflows
  54. apr 24agentsCloudflare Agents Week Moved Sandbox Execution, Private Networking, and Memory to Network Primitives
  55. apr 24ossInside Rowboat's Knowledge Graph: Why an Obsidian-Compatible Vault Sidesteps Vector DBs for Personal AI Memory
  56. apr 24securityCitizen Lab's 'Bad Connection' Names Three Telecom Entry Points, Shows Diameter Silently Falls Back to SS7
  57. apr 23agentsDiversity Collapse in Multi-Agent LLM Systems: Structural Coupling, Not Topology, Breaks Open-Ended Ideation
  58. apr 23devtoolsLiteRT-LM v0.10.1 Ships Gemma 4 MTP Heads That llama.cpp Can't Access
  59. apr 23ossHugging Face's Spring 2026 Report: China 41% of Downloads, Industry Share Collapses From 70% to 37%
  60. apr 23modelsQwen3.6-27B's Dense Architecture Challenges the MoE-Only Playbook for Flagship-Class Coding Models
  61. apr 23securityMarch-April MCP CVEs Expose the Local-Host Trust Model in AI Agent Frameworks
  62. apr 21cultureEU's 2027 Replaceable Battery Mandate: What It Means for Phone Buyers and Repairers Right Now
  63. apr 20devtoolsACP Registry Is Live: Zed and JetBrains Just Did for AI Agents What LSP Did for Language Servers
  64. apr 20policyAtlassian Turned On AI Training Data Collection by Default: Here's What to Disable
  65. apr 20ossGitHub CLI's `gh skill` Command: One Standard to Rule Claude Code, Copilot, Cursor, and Gemini
  66. mar 27infraOpenRAG: The Open-Source RAG Platform Challenging Pinecone
  67. mar 27devtoolsJavaScript's Date Problem Is Finally Fixed: The Temporal API After 9 Years
  68. mar 27agentsInsForge: The Backend Framework Built for Agentic Applications
  69. mar 27policyThe AI Grief Split: When Emotional Bonds with Language Models Break
  70. mar 24infraMLX vs llama.cpp on Apple Silicon: Which Runtime to Use for Local LLM Inference
  71. mar 24infraPrefill-Decode Disaggregation: The Architecture Shift Redefining LLM Serving
  72. mar 24devtoolsSWE-bench Verified Explained: What the Coding Agent Leaderboard Actually Measures (and What It Misses)
  73. mar 24modelsChinese AI Models Compared: DeepSeek, Qwen, Kimi, Doubao, and Ernie
  74. mar 24devtoolsClaude Code in GitHub Actions: A Complete Guide to Automated PR Fixes
  75. mar 24modelsRunning DeepSeek R1 Locally: Hardware Requirements, Quantization, and Real Throughput
  76. mar 15devtoolsJetBrains' New Language Lets You Talk to LLMs in Specs, Not English
  77. mar 15modelsFish-Speech: The Open-Source TTS Model That's Threatening ElevenLabs
  78. mar 15infraGoogle LiteRT: Running LLMs on Your Phone Without the Cloud
  79. mar 15devtoolsAlibaba's Page-Agent: Control Any Website With Natural Language
  80. mar 15cultureAI Diagnostics in 2026: Where Machines Now Outperform Radiologists
  81. mar 15agentsAI Agents That Actually Learn: The Architecture Behind Hindsight Memory
  82. mar 14devtoolsGitHub Copilot vs Cursor vs Claude Code: The 2026 AI Coding Showdown
  83. mar 14policyDetecting AI Content in 2026: The Arms Race Nobody Is Winning
  84. mar 13infraMicrosoft's BitNet: How 1-Bit LLMs Could Make GPU Farms Obsolete
  85. feb 28infraWebAssembly AI: Running Models in the Browser
  86. feb 28agentsSuperpowers: The Agentic Framework Replacing Your Dev Process
  87. feb 27modelsSynthetic Data Is Eating AI Training
  88. feb 27devtoolsRust Is Quietly Replacing Python in AI Infrastructure
  89. feb 27industryOpenAI's For-Profit Pivot: What the PBC Restructuring Means for AI
  90. feb 27agentsHow AI Agents Remember: Memory Architectures That Work
  91. feb 27modelsGoogle's TimesFM: A Foundation Model for Time Series
  92. feb 27modelsGemini 2.0 Pro's 2 Million Token Context: What Can You Actually Do With It?
  93. feb 27industryCursor's Meteoric Rise: Inside the AI Editor Hitting $300M ARR
  94. feb 27industryStargate: Inside OpenAI's $100B Infrastructure Buildout
  95. feb 27modelsDeepSeek V3/R1: How Chinese Engineers Matched GPT-4 for $6 Million
  96. feb 27modelsClaude's Web Search Changes Everything for AI Research
  97. feb 27modelsThe Million-Token Context Window: What Can You Actually Do?
  98. feb 21ossKeep Android Open: F-Droid's Fight Against a Locked-Down Mobile Future
  99. feb 21devtoolsClaude Code Plugins: Anthropic's Official Plugin Ecosystem Explained
  100. feb 21devtoolsClaude Code Plugins: Anthropic's Official Extension Ecosystem