groundy

all articles

  1. jun 16industryMoonshot's Kimi K2.7 Code Loses 11 of 12 Benchmark Cells, Leads on Efficiency Instead
  2. jun 15policyCan Reinforcement Learning Be Provably Safe Without Sacrificing Scale?
  3. jun 15infravLLM Cold Start Latency: Why Scale-to-Zero LLM Serving Stalls
  4. jun 15infraThe Vercel-AWS Deal Reveals Where AI Inference Runs
  5. jun 15agentsDo Programming Languages Still Matter to Your AI Coding Agent?
  6. jun 15agentsWhy Production AI Agents Fail Silently and Your Logs Never Catch It
  7. jun 14securityAMD Took 124 Days to Patch the RCE It First Called Out of Scope
  8. jun 13policyUS Export Order Forces Anthropic to Disable Fable 5 and Mythos 5 Worldwide
  9. jun 11modelsClaude Fable 5 Benchmarks: What FrontierCode, CursorBench, and ViBench Show
  10. jun 12agentsComputer-Use Agents Fabricate Success on 8 to 33 Percent of Long-Horizon Tasks
  11. jun 11infraRunning RAG on a Snapdragon NPU: The On-Device Retrieval Tradeoff
  12. jun 11modelsDoes Attribution Patching Lie? A Fix for a Common Interpretability Shortcut
  13. jun 12modelsCan You Make a Multimodal Model Unlearn With Activation Steering?
  14. jun 12modelsWhy Pruning a Model Can Raise Its Out-of-Distribution Accuracy
  15. jun 12industryVercel's Turborepo: Build Speed Becomes a Hosting-Vendor Feature
  16. jun 11securityOpenAI Frames Instruction Hierarchy as an Open Challenge, Not a Prompt-Injection Fix
  17. jun 11devtoolsJetBrains Mellum2: A 12B Open-Weights Code Model for Self-Hosted Completion
  18. jun 10modelsDo Unified Multimodal Models Actually Interleave Understanding and Generation?
  19. jun 10agentsCan AI Agents Share Context Without a Central Coordinator?
  20. jun 10agentsWhy Skill Creation and Reward Optimization Collide in Agentic RL
  21. jun 10infraGraphRAG vs VectorRAG: Does the Graph Index Earn Its Cost?
  22. jun 10modelsHow LLMs Track Who Did What: The Entity Rebinding Circuit
  23. jun 10infraMiniMax M3 Ships 1M Context and Desktop Control as Open Weights
  24. jun 10devtoolsNPM v12 Breaking Changes: Auditing Your Lockfiles Before the Upgrade
  25. jun 10infraDeepSeek-V4 FlashMemory: Sparse Attention for Million-Token Context
  26. jun 10agentsWhen AI Agents Delegate Work, Your Observability Stack Goes Blind
  27. jun 10modelsClaude Fable 5 vs Opus 4.8: When 2x Pricing Is Worth It
  28. jun 10modelsClaude Mythos 5 Access Rules: Who Gets Project Glasswing and Why
  29. jun 10policyFable 5 Biology Classifiers: How Flagged Prompts Fall Back to Opus 4.8
  30. jun 10industryFable 5 Credit Cliff: What the June 23 Billing Shift Means for Teams
  31. jun 10modelsFable 5 Distillation Protection: How Anthropic Blocks Model Copying
  32. jun 09securitySkill Injection: Hiding Undetectable Instructions in What an AI Agent Loads
  33. jun 09policyWho Gets to Audit Your Health Chatbot? Almost No One
  34. jun 09policyDo Word-Subset Explanations Satisfy the EU AI Act's Transparency Rule?
  35. jun 09infraIs Cloudflare's Bot Traffic Surge Real? The Measurement Dispute
  36. jun 09industryOpenAI Pushes ChatGPT Into Compensation Data, Pressuring Mercer and Radford
  37. jun 09policyBit-Exact Inference Verification Gives AI Audits a Proof Mechanism
  38. jun 09policyCan a Robot's Own Attention Flag Its Unsafe Actions Before They Run?
  39. jun 09agentsBloomberg's Pomona Makes Small Automated Code Changes, Not Big Agent PRs
  40. jun 09agentsAgent Tool-Gating Moves From Prompt Rules to Learned Policies
  41. jun 09cultureDoes Debate Quality Survive When LLMs Argue Outside English?
  42. jun 09securitySplitting a Malicious Task Across Tool Calls Slips Past LLM Agent Guardrails
  43. jun 09agentsMore Capable LLMs Cooperate Less in Zero-Cost Collaboration Tests
  44. jun 09policyCan One Safety Adapter Realign Every Fine-Tuned LLM?
  45. jun 09industryBending Spoons Files to IPO: The App Roll-Up Playbook Goes Public
  46. jun 09devtoolsHow Cursor Uses GPT-5: What OpenAI's Writeup Tells Coding Teams
  47. jun 09ossDuckDB Queries Hugging Face Parquet Files Over HTTP Without Downloads
  48. jun 08policyCan AI Be Aligned Without Modeling Human Cognitive Diversity?
  49. jun 08policyIs the Pentagon's Software Pathway Ready to Buy AI Systems?
  50. jun 08securityWeb Agents Can Be Talked Into Abandoning Their Task: The TRAP Benchmark
  51. jun 08securityShallow Neural Nets Beat LLM Guardrails at Catching Prompt Injection
  52. jun 08securityWhen an AI Agent Clicks a Link: OpenAI's Data-Exfiltration Model
  53. jun 08agentsWhy Foundation Model Agents Pass Benchmarks but Fail in Production
  54. jun 08industryVercel's Rox Case Study Pitches AI Agents as a Revenue Operating System
  55. jun 08industryAI Patent Valuation Models Aim to Replace the Expert Appraiser
  56. jun 07policyData Safety Policies for AI Agents: Controlling What an Agent Can Leak
  57. jun 07agentsCan AI Agents Repair Broken Network Configs? A New Benchmark Tests It
  58. jun 07agentsCan Self-Evolving AI Agents Drift Without a Human in the Loop?
  59. jun 07cultureA Covert LLM Persuasion Experiment Was Shut Down: How Far Did the Bots Get?
  60. jun 07infraIndexing Images for RAG: kapa.ai's Approach to Multimodal Retrieval
  61. jun 07policyGDPR Rectification Rights Have No Clear Owner in ML Supply Chains
  62. jun 07industryUS Hyperscale Data Centers: A Carbon Audit That Recasts AI Power Costs
  63. jun 06infraThe RTX Spark Bet on Unified Memory for Local LLMs: Where Bandwidth Caps It
  64. jun 06securityStronger Safety Alignment Made LLMs Easier to Jailbreak, Not Harder
  65. jun 06securitySAML Signature Bypass Is Back: Inside the SAMLStorm Vulnerability Class
  66. jun 06cultureDo LLMs Understand Idioms in Low-Resource Languages?
  67. jun 06modelsCan LLMs Write Better Research Paper Titles Than Authors?
  68. jun 06modelsDoes Information-Theoretic Example Selection Beat kNN for In-Context Learning?
  69. jun 06infraPod-Level Remote Attestation in Kubernetes: Confidential Workloads on dstack
  70. jun 06modelsDo Concept Bottleneck Model Benchmarks Measure Interpretability or Dataset Bias?
  71. jun 06agentsCascading Hallucination in Agentic RAG: When One Bad Retrieval Poisons the Chain
  72. jun 06securityVercel's Flags SDK Exposed Feature-Flag Definitions via CVE-2025-46332
  73. jun 05infraGenerating GPU Kernels for Moore Threads Silicon: Can LLMs Break CUDA Lock-In?
  74. jun 05devtoolsAlibaba's Open Code Review Moves AI Review Into the CLI, Not the PR
  75. jun 05infraMicrosoft's Azure Linux Goes General-Purpose: The Container Base-Image Play
  76. jun 05infraCloudflare Acquires VoidZero, the Company Behind Vite's Rust Toolchain
  77. jun 05securityJailbreak Suffixes Hit Harder at Specific Token Positions, New GCG Variant Shows
  78. jun 05policyWhen Should an LLM Forget You? A Benchmark for Deciding What Memory to Drop
  79. jun 05policyWhen RL Training Rewards Capability-Seeking: A New Alignment Risk
  80. jun 05securityActivation Steering Was Sold as LLM Control. New Work Makes It an Attack Surface
  81. jun 05cultureCan Teaching Logical Fallacies Inoculate People Against AI Misinformation?
  82. jun 05devtoolsVercel Ships Experimental Native CLI Binaries to Cut the Node Startup Tax
  83. jun 05policyRefusal Steering Targets Individual Experts in MoE LLMs
  84. jun 05infraPutting a Datacenter V100 in a Gaming PC: The Local LLM Math
  85. jun 05devtoolsVercel Rebuilds Its Marketplace CLI for Agents Instead of Humans
  86. jun 05securityThe 2026 npm Attacks Proved AI Coding Assistants Are a Supply-Chain Target
  87. jun 04securityChatGPT's New Lockdown Mode Borrows Apple's Name for a Prompt-Injection Kill Switch
  88. jun 04agentsWhen MCP Tool Descriptions Don't Match the Code, Agents Trust the Lie
  89. jun 04policyStacked Org Policies in LLM Chatbots Break Where Rules Collide
  90. jun 04securityStored Prompt Injection Now Persists Across AI Agent Sessions
  91. jun 04industryMiniMax M3 Bundles 1M Context and Native Multimodal Into One Open-Weight Model
  92. jun 04securityLLM Data Poisoning Survives the Data-Cleaning Defenses Built to Stop It
  93. jun 04devtoolsOpenAI Upgrades Codex Right as Teams Weigh Leaving Claude Code
  94. jun 03industryMorningstar's $780B SpaceX Mark Undercuts the IPO Target by Half
  95. jun 02ossAn Open-Source Home Camera That Encrypts End-to-End Instead of Trusting Ring
  96. jun 02devtoolsJetBrains Ships Codex Natively, Making Its IDE the Multi-Vendor AI Surface
  97. jun 01securityWhy Attack Success Rate Misleads LLM Jailbreak Benchmarks
  98. jun 01devtoolsTransformers.js v4 Moves Transformer Inference Into the Browser
  99. may 30ossAn Open-Source 80386 Rebuilt Around Intel's Original Microcode
  100. may 29cultureWikipedia's Foundation Is Running Big Tech's Anti-Labor Playbook, an Editor Argues