groundy

all articles

  1. modelsGLM-5.2 vs Kimi K2.7 Code: Two Open-Weight Bets on Agentic Coding
  2. modelsHow Linear Is a Transformer Feed-Forward Block? A New Test Says It's Learned, Not Built In
  3. devtoolsCursor Goes to SpaceX, Windsurf to Cognition: What Changes for Dev Teams
  4. cultureAI Essay Grading: What a Probe of LLM Internals Reveals About Scoring
  5. modelsGLM-5.2 Benchmarks: What 62.1% SWE-bench Pro and 99.2% AIME Actually Mean
  6. policyGLM-5.2 MIT Weights vs Llama License: Self-Hosting Compliance for Regulated Industries
  7. modelsGLM-5.2 on Terminal-Bench 2.1: Strengths, Gaps, and How to Route Real Coding Tasks
  8. modelsGLM-5.2 vs Claude Opus 4.8: Open-Weight Coding at Frontier Pricing
  9. modelsGLM-5.2's 753B MoE Costs More to Self-Host Than the MIT License Suggests
  10. infraRunning GLM-5.2 at Home: SGLang, vLLM, Transformers, and KTransformers Setup Guide
  11. devtoolsRunning GLM-5.2 in Cursor, Cline, and Roo Code: Migration Checklist and Gotchas
  12. modelsSTAR Replaces Scalar Reward in Text-to-Image RL with Attention-Derived Spatial Maps
  13. ossZhipu Open-Sources GLM-5.2 Under MIT While Anthropic Tightens Model Access
  14. modelsCan Editing One Neuron Fix LLM Repetition Loops?
  15. industryZhipu Ships GLM-5.2 With 1M Context and MIT Weights, but Zero Benchmarks at Launch
  16. infraAWS Bedrock Now Requires Data Sharing for Mythos: The Self-Hosting Calculus
  17. devtoolsVercel's Remend Turns Streaming-Markdown Repair Into a Dependency
  18. industryMoonshot's Kimi K2.7 Code Loses 11 of 12 Benchmark Cells, Leads on Efficiency Instead
  19. policyCan Reinforcement Learning Be Provably Safe Without Sacrificing Scale?
  20. infravLLM Cold Start Latency: Why Scale-to-Zero LLM Serving Stalls
  21. infraThe Vercel-AWS Deal Reveals Where AI Inference Runs
  22. agentsDo Programming Languages Still Matter to Your AI Coding Agent?
  23. agentsWhy Production AI Agents Fail Silently and Your Logs Never Catch It
  24. securityAMD Took 124 Days to Patch the RCE It First Called Out of Scope
  25. policyUS Export Order Forces Anthropic to Disable Fable 5 and Mythos 5 Worldwide
  26. modelsClaude Fable 5 Benchmarks: What FrontierCode, CursorBench, and ViBench Show
  27. agentsComputer-Use Agents Fabricate Success on 8 to 33 Percent of Long-Horizon Tasks
  28. infraRunning RAG on a Snapdragon NPU: The On-Device Retrieval Tradeoff
  29. modelsDoes Attribution Patching Lie? A Fix for a Common Interpretability Shortcut
  30. modelsCan You Make a Multimodal Model Unlearn With Activation Steering?
  31. modelsWhy Pruning a Model Can Raise Its Out-of-Distribution Accuracy
  32. industryVercel's Turborepo: Build Speed Becomes a Hosting-Vendor Feature
  33. securityOpenAI Frames Instruction Hierarchy as an Open Challenge, Not a Prompt-Injection Fix
  34. devtoolsJetBrains Mellum2: A 12B Open-Weights Code Model for Self-Hosted Completion
  35. modelsDo Unified Multimodal Models Actually Interleave Understanding and Generation?
  36. agentsCan AI Agents Share Context Without a Central Coordinator?
  37. agentsWhy Skill Creation and Reward Optimization Collide in Agentic RL
  38. infraGraphRAG vs VectorRAG: Does the Graph Index Earn Its Cost?
  39. modelsHow LLMs Track Who Did What: The Entity Rebinding Circuit
  40. infraMiniMax M3 Ships 1M Context and Desktop Control as Open Weights
  41. devtoolsNPM v12 Breaking Changes: Auditing Your Lockfiles Before the Upgrade
  42. infraDeepSeek-V4 FlashMemory: Sparse Attention for Million-Token Context
  43. agentsWhen AI Agents Delegate Work, Your Observability Stack Goes Blind
  44. modelsClaude Fable 5 vs Opus 4.8: When 2x Pricing Is Worth It
  45. modelsClaude Mythos 5 Access Rules: Who Gets Project Glasswing and Why
  46. policyFable 5 Biology Classifiers: How Flagged Prompts Fall Back to Opus 4.8
  47. industryFable 5 Credit Cliff: What the June 23 Billing Shift Means for Teams
  48. modelsFable 5 Distillation Protection: How Anthropic Blocks Model Copying
  49. securitySkill Injection: Hiding Undetectable Instructions in What an AI Agent Loads
  50. policyWho Gets to Audit Your Health Chatbot? Almost No One
  51. policyDo Word-Subset Explanations Satisfy the EU AI Act's Transparency Rule?
  52. infraIs Cloudflare's Bot Traffic Surge Real? The Measurement Dispute
  53. industryOpenAI Pushes ChatGPT Into Compensation Data, Pressuring Mercer and Radford
  54. policyBit-Exact Inference Verification Gives AI Audits a Proof Mechanism
  55. policyCan a Robot's Own Attention Flag Its Unsafe Actions Before They Run?
  56. agentsBloomberg's Pomona Makes Small Automated Code Changes, Not Big Agent PRs
  57. agentsAgent Tool-Gating Moves From Prompt Rules to Learned Policies
  58. cultureDoes Debate Quality Survive When LLMs Argue Outside English?
  59. securitySplitting a Malicious Task Across Tool Calls Slips Past LLM Agent Guardrails
  60. agentsMore Capable LLMs Cooperate Less in Zero-Cost Collaboration Tests
  61. policyCan One Safety Adapter Realign Every Fine-Tuned LLM?
  62. industryBending Spoons Files to IPO: The App Roll-Up Playbook Goes Public
  63. devtoolsHow Cursor Uses GPT-5: What OpenAI's Writeup Tells Coding Teams
  64. ossDuckDB Queries Hugging Face Parquet Files Over HTTP Without Downloads
  65. policyCan AI Be Aligned Without Modeling Human Cognitive Diversity?
  66. policyIs the Pentagon's Software Pathway Ready to Buy AI Systems?
  67. securityWeb Agents Can Be Talked Into Abandoning Their Task: The TRAP Benchmark
  68. securityShallow Neural Nets Beat LLM Guardrails at Catching Prompt Injection
  69. securityWhen an AI Agent Clicks a Link: OpenAI's Data-Exfiltration Model
  70. agentsWhy Foundation Model Agents Pass Benchmarks but Fail in Production
  71. industryVercel's Rox Case Study Pitches AI Agents as a Revenue Operating System
  72. industryAI Patent Valuation Models Aim to Replace the Expert Appraiser
  73. policyData Safety Policies for AI Agents: Controlling What an Agent Can Leak
  74. agentsCan AI Agents Repair Broken Network Configs? A New Benchmark Tests It
  75. agentsCan Self-Evolving AI Agents Drift Without a Human in the Loop?
  76. cultureA Covert LLM Persuasion Experiment Was Shut Down: How Far Did the Bots Get?
  77. infraIndexing Images for RAG: kapa.ai's Approach to Multimodal Retrieval
  78. policyGDPR Rectification Rights Have No Clear Owner in ML Supply Chains
  79. industryUS Hyperscale Data Centers: A Carbon Audit That Recasts AI Power Costs
  80. infraThe RTX Spark Bet on Unified Memory for Local LLMs: Where Bandwidth Caps It
  81. securityStronger Safety Alignment Made LLMs Easier to Jailbreak, Not Harder
  82. securitySAML Signature Bypass Is Back: Inside the SAMLStorm Vulnerability Class
  83. cultureDo LLMs Understand Idioms in Low-Resource Languages?
  84. modelsCan LLMs Write Better Research Paper Titles Than Authors?
  85. modelsDoes Information-Theoretic Example Selection Beat kNN for In-Context Learning?
  86. infraPod-Level Remote Attestation in Kubernetes: Confidential Workloads on dstack
  87. modelsDo Concept Bottleneck Model Benchmarks Measure Interpretability or Dataset Bias?
  88. agentsCascading Hallucination in Agentic RAG: When One Bad Retrieval Poisons the Chain
  89. securityVercel's Flags SDK Exposed Feature-Flag Definitions via CVE-2025-46332
  90. infraGenerating GPU Kernels for Moore Threads Silicon: Can LLMs Break CUDA Lock-In?
  91. devtoolsAlibaba's Open Code Review Moves AI Review Into the CLI, Not the PR
  92. infraMicrosoft's Azure Linux Goes General-Purpose: The Container Base-Image Play
  93. infraCloudflare Acquires VoidZero, the Company Behind Vite's Rust Toolchain
  94. securityJailbreak Suffixes Hit Harder at Specific Token Positions, New GCG Variant Shows
  95. policyWhen Should an LLM Forget You? A Benchmark for Deciding What Memory to Drop
  96. policyWhen RL Training Rewards Capability-Seeking: A New Alignment Risk
  97. securityActivation Steering Was Sold as LLM Control. New Work Makes It an Attack Surface
  98. cultureCan Teaching Logical Fallacies Inoculate People Against AI Misinformation?
  99. devtoolsVercel Ships Experimental Native CLI Binaries to Cut the Node Startup Tax
  100. policyRefusal Steering Targets Individual Experts in MoE LLMs