groundy

all articles

  1. jun 25modelsPV-TAM Corrects Decoding Drift and Boundary-Marker Bias in VLM Localization Scoring
  2. jun 25agentsDo AGENTS.md Files Actually Help Coding Agents? A New Benchmark Tests It
  3. jun 25agentsShould AI Shopping Agents Pay Micro-Transactions for Verified Product Data?
  4. jun 25modelsMeituan's General 365 Benchmark: Top Models All Score Under 63%
  5. jun 25modelsLLM Surrogates in A/B Tests: The 39% Recovery Gap and the Silent Bias Risk
  6. jun 25modelsLLM Token Pricing vs Compute Cost: What the Tokenomics Math Shows
  7. jun 25modelsDo LLM Judges Favor Their Own Output? A Sanity Check on Self-Preference
  8. jun 24agentsCan a Conversational Graph Compile Into a Goal-Oriented Dialogue Runtime?
  9. jun 24securityAuto-Reproducing Text-to-Image Jailbreaks From Papers: The PixJail Pipeline
  10. jun 24agentsCan a Cryptographic Certificate Prove an AI Agent's Output Is Valid?
  11. jun 24infraVercel on the AWS Marketplace: What the Listing Does to Procurement and Lock-In
  12. jun 24policyMachine-Readable AI Usage Terms: Does ODRL's Permission Model Hold Up?
  13. jun 24agentsCrewAI vs AutoGen vs Microsoft Agent Framework: AutoGen's Merger Reframes the 2026 Choice
  14. jun 24devtoolsVercel Now Deploys Long-Running Node Servers: The Serverless Boundary Shifts
  15. jun 24policyWho Audits the Safety Rules an LLM Agent Evolves for Itself?
  16. jun 24agentsCan You Trust an LLM Judge to Grade an Agentic Data Analysis System?
  17. jun 24agentsDo LLM Agent Societies Develop Their Own Authority Hierarchies?
  18. jun 24infraServing Cold MoE Models: CrossPool Disaggregates KV Cache and Weights
  19. jun 24securityVercel BotID's Telemetry Is a Threat Intelligence Feed Most Teams Discard
  20. jun 24policyWhen Vibe-Coded Software Is Safety-Critical, Who Verifies It?
  21. jun 24securityExtracting Unseen Training Data From an LLM by Poisoning Its Loss Landscape
  22. jun 24agentsDo Retrieval Metrics Predict Tool-Use Agent Success? A Paper Says No
  23. jun 24infraVercel's In-Function Concurrency: What It Does to Cold Starts and Billing
  24. jun 24policyCan You Trust an AI Robustness Certificate? A Paper Says Verify It
  25. jun 24agentsCan You Pinpoint Which Step Broke a Long-Horizon AI Agent?
  26. jun 24industryVercel's Series D Thesis Hardened Into a Whole-Stack Lock-In
  27. jun 24devtoolsmake-look-scanned Simulates Scans in an Offline WASM File, Exposing PDF Provenance as a Pixel Check
  28. jun 24infraPoisoning a RAG Retriever: How Conflict-Aware Edits Inject False Knowledge
  29. jun 24modelsCan AI Write CAD Programs? CADBench Measures the Gap
  30. jun 24infraVercel Raised Its CDN Origin Timeout to Two Minutes: What Breaks First
  31. jun 24infraGradio-Lite Runs Model Inference in the Browser via Pyodide, No Server
  32. jun 24devtoolsVercel's Billing Usage API: Wiring Cost Data Into CI Cost Gates
  33. jun 24infraCloudflare AI Gateway Adds Spend Limits to Cap the Runaway Inference Bill
  34. jun 24infraVercel Now Honors stale-if-error: Serving Stale Cache When the Origin Dies
  35. jun 24modelsByteDance's Doubao 2.1 Pro vs GPT-5.5: Reading Self-Reported Benchmarks
  36. jun 23policyCan a Benchmark Catch When AI Discharge Summaries Drop Care Steps?
  37. jun 23devtoolsVercel CLI Now Scopes Commands to the Local Directory: Audit Your CI Scripts
  38. jun 23securityReact Router CVE-2025-31137: Vercel's Edge Fix Is Not the Patch
  39. jun 23infraVercel's Manual CDN Purge API: Cache Control Without a Redeploy
  40. jun 23industrySamsung Picks OpenAI's Codex for Its Engineers, Pressuring GitHub Copilot
  41. jun 23devtoolsVercel Sandbox Snapshot Retention: What Custom Windows Change for Agent Runtimes
  42. jun 23industryPotion.so Sold After 4,000 Vercel Deploys: The Micro-SaaS Exit Playbook
  43. jun 23policyDo LLM Personality Tests Measure Anything? A New Paper Says No
  44. jun 23securityReported React Server Components Leak Is Unconfirmed: Audit the Payload
  45. jun 23devtoolsGenerating Vercel Firewall Rules From Natural Language: What to Audit
  46. jun 23devtoolsGLM-5.2 Coding Plan vs Claude Opus 4.8: Picking a Model for Coding Agents
  47. jun 23securityVercel's Secure AI Agent Guidance Pushes Defense Into the Sandbox
  48. jun 23securityNx Supply-Chain Attack Used Developers' Own AI CLIs to Hunt Secrets
  49. jun 23industryVercel Folds Backends, Agent Tooling, and Operations Into Its Deploy Platform
  50. jun 23infraCloudflare Now Routes Public Traffic to Private Apps via DNS, No VPN
  51. jun 23ossOpenAI's Patch the Planet Is Security Capacity for Nine Projects, Not Sustainability Funding
  52. jun 23ossMiniMax M3 Claims GPT-5.5-Beating Code With 1M Context and Open Weights
  53. jun 23industryGeorge Hotz Says Only AGI Doom Justifies Today's AI Valuations
  54. jun 23infraGitHub's AI Capacity Crunch Pushes Microsoft to Rent AWS Compute
  55. jun 23policyCommunity LoRA Mining Raises a Consent Gap for Style Generation
  56. jun 22cultureWhy Audio Deepfake Detectors Keep Losing the Voice-Cloning Arms Race
  57. jun 21securityMixed Compliance Data Makes Safety Fine-Tuning a Curation Problem
  58. jun 21policyWhen an LLM Narrates a Solver, the Explanation Drifts From the Math
  59. jun 21infraCloudflare's Temporary Accounts Give AI Agents Disposable Credentials
  60. jun 21policyGrading DiffusionGemma: How an Open-Weight Diffusion Model Scores on Transparency
  61. jun 21policyWho Owns Editorial Authority When LLMs Mediate Knowledge?
  62. jun 21ossLithuania's Open-Source Drone-Detection Network Signals an Air-Defense Shift
  63. jun 21cultureWhy AI Misreads Nigerian English: A Register Gap in Public Discourse
  64. jun 21agentsDeep-Research Benchmarks Hide How Agents Fail at Open-Web Source Grounding
  65. jun 21policyVector Database Access Control Is Missing, and RAG Pipelines Pay for It
  66. jun 21agentsDSPy Ships Autonomous Prompt Optimization, but Judge Drift Is the Failure Mode
  67. jun 21cultureWhat YouTube's Coding Tutorials Teach About Who Belongs in Software
  68. jun 21industryFinance Agent Benchmarks Expose Where Lending Automation Breaks
  69. jun 21ossNLnet's Grant Model Diverges From VC-Backed Open Source
  70. jun 21ossAdam's Open-Source AI CAD Claim Lacks a Confirmed Repo or Accuracy Benchmark
  71. jun 21agentsDo AI Agents Reach for Over-Privileged Tools When Simpler Ones Suffice?
  72. jun 21agentsWhen Should Multi-Agent Systems Use an Event Bus Instead of an Orchestrator?
  73. jun 21ossEpic Open-Sources Lore, a VCS Pitched at Git's Scaling Ceiling
  74. jun 21infraRunning Long-Context Agents on a 4-Bit KV Cache: Where Accuracy Breaks
  75. jun 21securityDefending Agentic AI With Deception: Misdirecting Model-Guided Attacks
  76. jun 21securityThe Autonomy Tax: Why RL Rewards the Wrong Behavior in Agents
  77. jun 21securityAnthropic's Procurement Risk Is Policy Refusal, Not Jailbreaks
  78. jun 20industryCan You Predict a Fine-Tune's Payoff Before Training Finishes?
  79. jun 20cultureWhen an Algorithm Sequences Gig Hiring, Whose Objective Does It Optimize?
  80. jun 20infraWhen LLM-Generated CUDA Kernels Pass Tests but Get the Math Wrong
  81. jun 20modelsCan RoboSSM's State-Space Backbone Replace Transformer Imitation Policies?
  82. jun 20modelsPruning Experts to Shrink MoE Models: Does Attribution-Guided Compression Beat Magnitude?
  83. jun 20agentsCan Deontic Policy Rules Govern an AI Agent at Runtime?
  84. jun 20modelsGLM-5.2 vs Kimi K2.7 Code: Two Open-Weight Bets on Agentic Coding
  85. jun 20modelsHow Linear Is a Transformer Feed-Forward Block? A New Test Says It's Learned, Not Built In
  86. jun 20devtoolsCursor Goes to SpaceX, Windsurf to Cognition: What Changes for Dev Teams
  87. jun 19cultureAI Essay Grading: What a Probe of LLM Internals Reveals About Scoring
  88. jun 19modelsGLM-5.2 Benchmarks: What 62.1% SWE-bench Pro and 99.2% AIME Actually Mean
  89. jun 19policyGLM-5.2 MIT Weights vs Llama License: Self-Hosting Compliance for Regulated Industries
  90. jun 19modelsGLM-5.2 on Terminal-Bench 2.1: Strengths, Gaps, and How to Route Real Coding Tasks
  91. jun 19modelsGLM-5.2 vs Claude Opus 4.8: Open-Weight Coding at Frontier Pricing
  92. jun 19modelsGLM-5.2's 753B MoE Costs More to Self-Host Than the MIT License Suggests
  93. jun 19infraRunning GLM-5.2 at Home: SGLang, vLLM, Transformers, and KTransformers Setup Guide
  94. jun 19devtoolsRunning GLM-5.2 in Cursor, Cline, and Roo Code: Migration Checklist and Gotchas
  95. jun 18modelsSTAR Replaces Scalar Reward in Text-to-Image RL with Attention-Derived Spatial Maps
  96. jun 16ossZhipu Open-Sources GLM-5.2 Under MIT While Anthropic Tightens Model Access
  97. jun 16modelsCan Editing One Neuron Fix LLM Repetition Loops?
  98. jun 16industryZhipu Ships GLM-5.2 With 1M Context and MIT Weights, but Zero Benchmarks at Launch
  99. jun 16infraAWS Bedrock Now Requires Data Sharing for Mythos: The Self-Hosting Calculus
  100. jun 16devtoolsVercel's Remend Turns Streaming-Markdown Repair Into a Dependency