groundy

all articles

  1. jun 21ossAdam's Open-Source AI CAD Claim Lacks a Confirmed Repo or Accuracy Benchmark
  2. jun 21agentsDo AI Agents Reach for Over-Privileged Tools When Simpler Ones Suffice?
  3. jun 21agentsWhen Should Multi-Agent Systems Use an Event Bus Instead of an Orchestrator?
  4. jun 21ossEpic Open-Sources Lore, a VCS Pitched at Git's Scaling Ceiling
  5. jun 21infraRunning Long-Context Agents on a 4-Bit KV Cache: Where Accuracy Breaks
  6. jun 21securityDefending Agentic AI With Deception: Misdirecting Model-Guided Attacks
  7. jun 21securityThe Autonomy Tax: Why RL Rewards the Wrong Behavior in Agents
  8. jun 21securityAnthropic's Procurement Risk Is Policy Refusal, Not Jailbreaks
  9. jun 20industryCan You Predict a Fine-Tune's Payoff Before Training Finishes?
  10. jun 20cultureWhen an Algorithm Sequences Gig Hiring, Whose Objective Does It Optimize?
  11. jun 20infraWhen LLM-Generated CUDA Kernels Pass Tests but Get the Math Wrong
  12. jun 20modelsCan RoboSSM's State-Space Backbone Replace Transformer Imitation Policies?
  13. jun 20modelsPruning Experts to Shrink MoE Models: Does Attribution-Guided Compression Beat Magnitude?
  14. jun 20agentsCan Deontic Policy Rules Govern an AI Agent at Runtime?
  15. jun 20modelsGLM-5.2 vs Kimi K2.7 Code: Two Open-Weight Bets on Agentic Coding
  16. jun 20modelsHow Linear Is a Transformer Feed-Forward Block? A New Test Says It's Learned, Not Built In
  17. jun 20devtoolsCursor Goes to SpaceX, Windsurf to Cognition: What Changes for Dev Teams
  18. jun 19cultureAI Essay Grading: What a Probe of LLM Internals Reveals About Scoring
  19. jun 19modelsGLM-5.2 Benchmarks: What 62.1% SWE-bench Pro and 99.2% AIME Actually Mean
  20. jun 19policyGLM-5.2 MIT Weights vs Llama License: Self-Hosting Compliance for Regulated Industries
  21. jun 19modelsGLM-5.2 on Terminal-Bench 2.1: Strengths, Gaps, and How to Route Real Coding Tasks
  22. jun 19modelsGLM-5.2 vs Claude Opus 4.8: Open-Weight Coding at Frontier Pricing
  23. jun 19modelsGLM-5.2's 753B MoE Costs More to Self-Host Than the MIT License Suggests
  24. jun 19infraRunning GLM-5.2 at Home: SGLang, vLLM, Transformers, and KTransformers Setup Guide
  25. jun 19devtoolsRunning GLM-5.2 in Cursor, Cline, and Roo Code: Migration Checklist and Gotchas
  26. jun 18modelsSTAR Replaces Scalar Reward in Text-to-Image RL with Attention-Derived Spatial Maps
  27. jun 16ossZhipu Open-Sources GLM-5.2 Under MIT While Anthropic Tightens Model Access
  28. jun 16modelsCan Editing One Neuron Fix LLM Repetition Loops?
  29. jun 16industryZhipu Ships GLM-5.2 With 1M Context and MIT Weights, but Zero Benchmarks at Launch
  30. jun 16infraAWS Bedrock Now Requires Data Sharing for Mythos: The Self-Hosting Calculus
  31. jun 16devtoolsVercel's Remend Turns Streaming-Markdown Repair Into a Dependency
  32. jun 16industryMoonshot's Kimi K2.7 Code Loses 11 of 12 Benchmark Cells, Leads on Efficiency Instead
  33. jun 15policyCan Reinforcement Learning Be Provably Safe Without Sacrificing Scale?
  34. jun 15infravLLM Cold Start Latency: Why Scale-to-Zero LLM Serving Stalls
  35. jun 15infraThe Vercel-AWS Deal Reveals Where AI Inference Runs
  36. jun 15agentsDo Programming Languages Still Matter to Your AI Coding Agent?
  37. jun 15agentsWhy Production AI Agents Fail Silently and Your Logs Never Catch It
  38. jun 14securityAMD Took 124 Days to Patch the RCE It First Called Out of Scope
  39. jun 13policyUS Export Order Forces Anthropic to Disable Fable 5 and Mythos 5 Worldwide
  40. jun 11modelsClaude Fable 5 Benchmarks: What FrontierCode, CursorBench, and ViBench Show
  41. jun 12agentsComputer-Use Agents Fabricate Success on 8 to 33 Percent of Long-Horizon Tasks
  42. jun 11infraRunning RAG on a Snapdragon NPU: The On-Device Retrieval Tradeoff
  43. jun 11modelsDoes Attribution Patching Lie? A Fix for a Common Interpretability Shortcut
  44. jun 12modelsCan You Make a Multimodal Model Unlearn With Activation Steering?
  45. jun 12modelsWhy Pruning a Model Can Raise Its Out-of-Distribution Accuracy
  46. jun 12industryVercel's Turborepo: Build Speed Becomes a Hosting-Vendor Feature
  47. jun 11securityOpenAI Frames Instruction Hierarchy as an Open Challenge, Not a Prompt-Injection Fix
  48. jun 11devtoolsJetBrains Mellum2: A 12B Open-Weights Code Model for Self-Hosted Completion
  49. jun 10modelsDo Unified Multimodal Models Actually Interleave Understanding and Generation?
  50. jun 10agentsCan AI Agents Share Context Without a Central Coordinator?
  51. jun 10agentsWhy Skill Creation and Reward Optimization Collide in Agentic RL
  52. jun 10infraGraphRAG vs VectorRAG: Does the Graph Index Earn Its Cost?
  53. jun 10modelsHow LLMs Track Who Did What: The Entity Rebinding Circuit
  54. jun 10infraMiniMax M3 Ships 1M Context and Desktop Control as Open Weights
  55. jun 10devtoolsNPM v12 Breaking Changes: Auditing Your Lockfiles Before the Upgrade
  56. jun 10infraDeepSeek-V4 FlashMemory: Sparse Attention for Million-Token Context
  57. jun 10agentsWhen AI Agents Delegate Work, Your Observability Stack Goes Blind
  58. jun 10modelsClaude Fable 5 vs Opus 4.8: When 2x Pricing Is Worth It
  59. jun 10modelsClaude Mythos 5 Access Rules: Who Gets Project Glasswing and Why
  60. jun 10policyFable 5 Biology Classifiers: How Flagged Prompts Fall Back to Opus 4.8
  61. jun 10industryFable 5 Credit Cliff: What the June 23 Billing Shift Means for Teams
  62. jun 10modelsFable 5 Distillation Protection: How Anthropic Blocks Model Copying
  63. jun 09securitySkill Injection: Hiding Undetectable Instructions in What an AI Agent Loads
  64. jun 09policyWho Gets to Audit Your Health Chatbot? Almost No One
  65. jun 09policyDo Word-Subset Explanations Satisfy the EU AI Act's Transparency Rule?
  66. jun 09infraIs Cloudflare's Bot Traffic Surge Real? The Measurement Dispute
  67. jun 09industryOpenAI Pushes ChatGPT Into Compensation Data, Pressuring Mercer and Radford
  68. jun 09policyBit-Exact Inference Verification Gives AI Audits a Proof Mechanism
  69. jun 09policyCan a Robot's Own Attention Flag Its Unsafe Actions Before They Run?
  70. jun 09agentsBloomberg's Pomona Makes Small Automated Code Changes, Not Big Agent PRs
  71. jun 09agentsAgent Tool-Gating Moves From Prompt Rules to Learned Policies
  72. jun 09cultureDoes Debate Quality Survive When LLMs Argue Outside English?
  73. jun 09securitySplitting a Malicious Task Across Tool Calls Slips Past LLM Agent Guardrails
  74. jun 09agentsMore Capable LLMs Cooperate Less in Zero-Cost Collaboration Tests
  75. jun 09policyCan One Safety Adapter Realign Every Fine-Tuned LLM?
  76. jun 09industryBending Spoons Files to IPO: The App Roll-Up Playbook Goes Public
  77. jun 09devtoolsHow Cursor Uses GPT-5: What OpenAI's Writeup Tells Coding Teams
  78. jun 09ossDuckDB Queries Hugging Face Parquet Files Over HTTP Without Downloads
  79. jun 08policyCan AI Be Aligned Without Modeling Human Cognitive Diversity?
  80. jun 08policyIs the Pentagon's Software Pathway Ready to Buy AI Systems?
  81. jun 08securityWeb Agents Can Be Talked Into Abandoning Their Task: The TRAP Benchmark
  82. jun 08securityShallow Neural Nets Beat LLM Guardrails at Catching Prompt Injection
  83. jun 08securityWhen an AI Agent Clicks a Link: OpenAI's Data-Exfiltration Model
  84. jun 08agentsWhy Foundation Model Agents Pass Benchmarks but Fail in Production
  85. jun 08industryVercel's Rox Case Study Pitches AI Agents as a Revenue Operating System
  86. jun 08industryAI Patent Valuation Models Aim to Replace the Expert Appraiser
  87. jun 07policyData Safety Policies for AI Agents: Controlling What an Agent Can Leak
  88. jun 07agentsCan AI Agents Repair Broken Network Configs? A New Benchmark Tests It
  89. jun 07agentsCan Self-Evolving AI Agents Drift Without a Human in the Loop?
  90. jun 07cultureA Covert LLM Persuasion Experiment Was Shut Down: How Far Did the Bots Get?
  91. jun 07infraIndexing Images for RAG: kapa.ai's Approach to Multimodal Retrieval
  92. jun 07policyGDPR Rectification Rights Have No Clear Owner in ML Supply Chains
  93. jun 07industryUS Hyperscale Data Centers: A Carbon Audit That Recasts AI Power Costs
  94. jun 06infraThe RTX Spark Bet on Unified Memory for Local LLMs: Where Bandwidth Caps It
  95. jun 06securityStronger Safety Alignment Made LLMs Easier to Jailbreak, Not Harder
  96. jun 06securitySAML Signature Bypass Is Back: Inside the SAMLStorm Vulnerability Class
  97. jun 06cultureDo LLMs Understand Idioms in Low-Resource Languages?
  98. jun 06modelsCan LLMs Write Better Research Paper Titles Than Authors?
  99. jun 06modelsDoes Information-Theoretic Example Selection Beat kNN for In-Context Learning?
  100. jun 06infraPod-Level Remote Attestation in Kubernetes: Confidential Workloads on dstack