groundy

Groundy — independent coverage of developer tools, infrastructure, and platforms

models

Deepfake KYC Fraud: Tamper-Resilient Watermarks That Recover the Original Face

VeriFi preprint claims watermarks recover original faces after tampering, not just flag fakes. Robustness is untested against commercial tools and compression, so treat as a R

6 min
devtools

AI Autofix Won't Clear Your Backlog: Patching Is a Capacity Problem

infra

Getting an AI Agent Into Cloudflare's BotBase: What Operators Must Verify

policy

LLM Judge Scores Change Run to Run: Why AI Audits Need Variance Reporting

  1. industryOpenAI's Cursor Decision Is a Vendor Exit Drill for AI Coding Tools
  2. policyHugging Face's AI Action Plan Reply: Open Weights vs the Frontier Lab Lobby
  3. devtoolsTeaching Junior Developers When AI Writes the First Draft
  4. modelsCan LLMs Train on Their Own Problems? What Zero-Data Self-Play Changes
  5. devtoolsDiagnosing LLM Prompt Injection Detectors Before You Gate an Agent on Them
  6. infraAlibaba's ScaleSense: When Learned Autoscaling Beats Provisioning Rules
  7. modelsWhy LLM Agent Benchmarks Move When the Harness Changes
  8. policyWhy Temperature 0 Won't Save Your Financial AI Audit
  9. infraAWS Cognito Postmortem: The Real Cost of Free Managed Auth
  10. modelsCan You Serve LLMs on 2-Bit Weights? What Ultra-Low-Bit Quantization Costs
  11. devtoolsLocal Coding LLMs Hallucinate Packages: Slopsquatting Defenses Compared
  12. policyPayPal Blocks GrapheneOS: A Payment Contingency Guide for Open Source
  13. infraHow Cloudflare Saved 100 TB in 1.1.1.1's DNS Cache and What Operators Can Copy
  14. devtoolsTesting Cloud APIs Without a Cloud Account: LocalStack vs Emulator Synthesis
  15. modelsDo LLMs Still Need BPE? What RL-Trained Tokenizers Change
  16. policyDoes the EU AI Act Exempt Open-Weight Models Like Qwen and Llama?
  17. agentsQwen-Agent Stretches 8k to 1M Context: Do You Need a Long-Context Model?
  18. industryLLM Uncertainty Methods Compared: What Actually Catches Hallucinations
  19. infraWhere Simple RAG Breaks: Multi-Document QA Needs Hierarchy, Not More Chunks
  20. agentsCan AI Agents Do Root Cause Analysis? What Cloud-OpsBench Measures
  21. agentsWhich AI Agents Behave Badly: Cloudflare's Agentic Internet Data
  22. policyAI Agents Outgrow OAuth: What Task-Scoped Authorization Requires
  23. modelsWhy Prompt Caching Can Change Model Outputs: A Prefix Invariance Audit
  24. industryAI Video Generators as World Simulators: What VGI-Bench Actually Measures
browse all 739 articles →