groundy

all articles

  1. jun 06modelsDo Concept Bottleneck Model Benchmarks Measure Interpretability or Dataset Bias?
  2. jun 06agentsCascading Hallucination in Agentic RAG: When One Bad Retrieval Poisons the Chain
  3. jun 06securityVercel's Flags SDK Exposed Feature-Flag Definitions via CVE-2025-46332
  4. jun 05infraGenerating GPU Kernels for Moore Threads Silicon: Can LLMs Break CUDA Lock-In?
  5. jun 05devtoolsAlibaba's Open Code Review Moves AI Review Into the CLI, Not the PR
  6. jun 05infraMicrosoft's Azure Linux Goes General-Purpose: The Container Base-Image Play
  7. jun 05infraCloudflare Acquires VoidZero, the Company Behind Vite's Rust Toolchain
  8. jun 05securityJailbreak Suffixes Hit Harder at Specific Token Positions, New GCG Variant Shows
  9. jun 05policyWhen Should an LLM Forget You? A Benchmark for Deciding What Memory to Drop
  10. jun 05policyWhen RL Training Rewards Capability-Seeking: A New Alignment Risk
  11. jun 05securityActivation Steering Was Sold as LLM Control. New Work Makes It an Attack Surface
  12. jun 05cultureCan Teaching Logical Fallacies Inoculate People Against AI Misinformation?
  13. jun 05devtoolsVercel Ships Experimental Native CLI Binaries to Cut the Node Startup Tax
  14. jun 05policyRefusal Steering Targets Individual Experts in MoE LLMs
  15. jun 05infraPutting a Datacenter V100 in a Gaming PC: The Local LLM Math
  16. jun 05devtoolsVercel Rebuilds Its Marketplace CLI for Agents Instead of Humans
  17. jun 05securityThe 2026 npm Attacks Proved AI Coding Assistants Are a Supply-Chain Target
  18. jun 04securityChatGPT's New Lockdown Mode Borrows Apple's Name for a Prompt-Injection Kill Switch
  19. jun 04agentsWhen MCP Tool Descriptions Don't Match the Code, Agents Trust the Lie
  20. jun 04policyStacked Org Policies in LLM Chatbots Break Where Rules Collide
  21. jun 04securityStored Prompt Injection Now Persists Across AI Agent Sessions
  22. jun 04industryMiniMax M3 Bundles 1M Context and Native Multimodal Into One Open-Weight Model
  23. jun 04securityLLM Data Poisoning Survives the Data-Cleaning Defenses Built to Stop It
  24. jun 04devtoolsOpenAI Upgrades Codex Right as Teams Weigh Leaving Claude Code
  25. jun 03industryMorningstar's $780B SpaceX Mark Undercuts the IPO Target by Half
  26. jun 02ossAn Open-Source Home Camera That Encrypts End-to-End Instead of Trusting Ring
  27. jun 02devtoolsJetBrains Ships Codex Natively, Making Its IDE the Multi-Vendor AI Surface
  28. jun 01securityWhy Attack Success Rate Misleads LLM Jailbreak Benchmarks
  29. jun 01devtoolsTransformers.js v4 Moves Transformer Inference Into the Browser
  30. may 30ossAn Open-Source 80386 Rebuilt Around Intel's Original Microcode
  31. may 29cultureWikipedia's Foundation Is Running Big Tech's Anti-Labor Playbook, an Editor Argues
  32. may 29agentsMulti-Agent LLM Coordination: Why Attention Steering Beats Full Broadcast
  33. may 29modelsPersona Prompts Change Who an LLM Recommends as an Expert
  34. may 29agentsDataClawBench: AI Agents Fail at Exploratory Financial Analysis Across 492 Tasks
  35. may 29policyRLHF Can Be Exploited to Optimize the Biases It Was Built to Suppress
  36. may 28ossFrontier AI Has Broken Open CTFs: Why Claude Code Now One-Shots Medium Pwn Challenges
  37. may 28policySelective Geometry Attacks Bypass LLM Safety Alignment, New arXiv Paper Reports
  38. may 28modelsOpus 4.8 Batch API: 1M Context, 300k Output, and Team Cost Controls
  39. may 28agentsClaude Code Dynamic Workflows: Spawning 100 Parallel Subagents on Opus 4.8
  40. may 27infraWhy LLMs Still Botch Kubernetes Manifests: The Training-Data Gap
  41. may 27securityOpenAI's New Safety Bug Bounty Pays Researchers for Jailbreaks and Policy Bypasses
  42. may 27industryHuggingFace's $100M Series C Bets Open-Source AI Can Outlast Per-Token Pricing Wars
  43. may 27infraGemma 4 31B on Cloud TPU vs GPU: The Serving Cost Crossover Point
  44. may 27agentsClaude Code, Cursor, Copilot: How Agentic Coding Assistants Get Weaponized as Attacker Shells
  45. may 26devtoolsBun Rewrites Its Core From Zig to Rust, Putting Downstream Zig Bindings at Risk
  46. may 26infraObjectCache Moves KV Reuse to S3-Class Storage: Why Layerwise Retrieval Beats Full-Prefix Cache Hits
  47. may 26devtoolsPromptArmor Shows Microsoft Copilot Cowork Can Be Tricked Into Exfiltrating Files
  48. may 25modelsμP Hyperparameter Transfer Has an Embedding Layer Hole, New arXiv Paper Says
  49. may 25devtoolsRmux Brings a Playwright SDK to tmux Sessions for Agent Automation Workflows
  50. may 25ossColorado SB051 Carves Out Open Source From Age Verification After Maintainer Backlash
  51. may 24cultureUS Researchers Hit With New Federal Limits on Publishing With Foreign Collaborators
  52. may 23securityAI Jailbreaks Are Now a Reasoning Problem, Not a Prompt Problem
  53. may 23devtoolsGoogle Sunsets Gemini CLI on June 18: Forced Migration to Antigravity CLI Breaks Existing Automation
  54. may 23agentsSpecBench Exposes Reward Hacking in Long-Horizon Coding Agents
  55. may 23infravLLM 0.21 Makes Prefill-Decode Disaggregation Actually Practical
  56. may 19industryBret Taylor's Sierra Raises $950M at $15B, Claims 40% of Fortune 50 Use Its Agents
  57. may 18industrySierra Raises $950M at $15B, Locking 40% of the Fortune 50 Into Its Agent Platform Before the Labs Go Direct
  58. may 18policyFrontier AI Has Broken the Open CTF Format: What the Scoreboard Collapse Means for Security Training
  59. may 18securityTrustFall: One Keypress in Claude Code, Gemini CLI, Cursor, and Copilot CLI Triggers Unsandboxed RCE
  60. may 18devtoolsClaude Code Adds Plugin Dependency Enforcement: disable Now Refuses to Break Transitive Chains
  61. may 18industryPayPal's $1.5B AI Overhaul Cuts 4,760 Jobs and Reframes Layoffs as Capex
  62. may 18policyFrontier AI Broke Open CTFs: What Hack The Box and BearcatCTF 2026 Results Mean for Security Hiring Signals
  63. may 18industryOpenAI's $4B Deployment Company Buys Tomoro and Signs 19 Partners to Own Implementation
  64. may 18agentsLangGraph 1.2.0 Makes Error-Handler Resume Crash-Durable: With Conditions
  65. may 18agentsCrewAI vs AutoGen vs LangGraph 2026: The Real Trade-Off After Maintenance Mode
  66. may 18cultureApple's $250M Siri Settlement: iPhone 16 Buyers Get $25 to $95 for Undelivered AI
  67. may 18securityMultiBreak Benchmark: 10,389 Multi-Turn Jailbreak Prompts Raise ASR 54pp on DeepSeek-R1-7B
  68. may 18industryAnthropic's $1.5B Joint Venture With Goldman Sachs and Blackstone Sends Claude Into PE Portfolio Companies
  69. may 18industryOpenAI Offers Two Months of Free Codex to Enterprises Switching From Claude Within 30 Days
  70. may 18cultureAB 566 Forces Chrome and Safari to Ship Opt-Out Signals by 2027. It Shields Them from Google's 86% GPC Failure
  71. may 18ossBrowserAct Open-Sources Stealth Browser Engine with 93% Token Reduction Claim
  72. may 18devtoolsGitHub Copilot's Opus 4.7 Multiplier: 7.5x to 15x to 27x in 60 Days
  73. apr 29osspgBackRest Is No Longer Maintained: PostgreSQL Backup Alternatives After the Project Stalls
  74. apr 29securityInstructLab CVE-2026-6859: Hardcoded trust_remote_code=True Turns Any HuggingFace Model Into RCE
  75. apr 29devtoolsPydantic AI v1.87 Closes the LangGraph Gap: Deferred Tool Calls, OpenTelemetry Eval, Stateful Compaction
  76. apr 29securityMercor's 4TB Lapsus$ Breach Hands Voice-Clone Attackers 40,000 Pre-Verified Targets
  77. apr 29agentsCouncil Mode Cuts Multi-Agent LLM Hallucination 35.9% at 4.2x Token Cost on HaluEval
  78. apr 29devtoolsClaude Code vs Cursor vs Copilot After the April 2026 Reshuffle: How the Comparison Math Changed
  79. apr 28devtoolsGitHub Copilot Replaces Premium Request Units With Token-Metered AI Credits on June 1
  80. apr 28ossfree-claude-code Routes Claude Code Through NVIDIA NIM and Local Models After Anthropic's CLI Ban
  81. apr 28industryMicrosoft and OpenAI End Their Exclusive Revenue-Sharing Deal: What It Means for Azure's AI Moat
  82. apr 28industryAnthropic Ends Flat-Fee Enterprise Claude, Enforces Per-Token Billing
  83. apr 24devtoolsGitHub CLI v2.91.0 Turns On Default Telemetry: What gh Collects and How to Opt Out in CI and Agent Pipelines
  84. apr 24devtoolsGitHub Copilot Drops Opus from Pro and Pauses Signups: The Forced Migration Facing Agentic Workflows
  85. apr 24agentsCloudflare Agents Week Moved Sandbox Execution, Private Networking, and Memory to Network Primitives
  86. apr 24ossInside Rowboat's Knowledge Graph: Why an Obsidian-Compatible Vault Sidesteps Vector DBs for Personal AI Memory
  87. apr 24securityCitizen Lab's 'Bad Connection' Names Three Telecom Entry Points, Shows Diameter Silently Falls Back to SS7
  88. apr 23agentsDiversity Collapse in Multi-Agent LLM Systems: Structural Coupling, Not Topology, Breaks Open-Ended Ideation
  89. apr 23devtoolsLiteRT-LM v0.10.1 Ships Gemma 4 MTP Heads That llama.cpp Can't Access
  90. apr 23ossHugging Face's Spring 2026 Report: China 41% of Downloads, Industry Share Collapses From 70% to 37%
  91. apr 23modelsQwen3.6-27B's Dense Architecture Challenges the MoE-Only Playbook for Flagship-Class Coding Models
  92. apr 23securityMarch-April MCP CVEs Expose the Local-Host Trust Model in AI Agent Frameworks
  93. apr 21cultureEU's 2027 Replaceable Battery Mandate: What It Means for Phone Buyers and Repairers Right Now
  94. apr 20devtoolsACP Registry Is Live: Zed and JetBrains Just Did for AI Agents What LSP Did for Language Servers
  95. apr 20policyAtlassian Turned On AI Training Data Collection by Default: Here's What to Disable
  96. apr 20ossGitHub CLI's `gh skill` Command: One Standard to Rule Claude Code, Copilot, Cursor, and Gemini
  97. mar 27infraOpenRAG: The Open-Source RAG Platform Challenging Pinecone
  98. mar 27devtoolsJavaScript's Date Problem Is Finally Fixed: The Temporal API After 9 Years
  99. mar 27agentsInsForge: The Backend Framework Built for Agentic Applications
  100. mar 27policyThe AI Grief Split: When Emotional Bonds with Language Models Break