groundy

Groundy — independent coverage of developer tools, infrastructure, and platforms

models

Why Drift Monitors Confuse Covariate Shift With Concept Drift

A new preprint proposes CJSD, a two-discriminator test that separates covariate shift from concept drift. This distinction determines whether to retrain or recalibrate, fixing

6 min
devtools

Node CLIs Running Local AI Models: Transformers.js v4 vs a Python Sidecar

models

AI-Generated Apps Look Right, but Do They Actually Work?

agents

Compressing LLM Agent History: Pixel Rendering vs Summarization

  1. policyCloudflare's 1-Click Fix for Vibe-Coded Apps: What It Doesn't Solve
  2. infraGPU Memory Explained: Why LLM Throughput Collapses When VRAM Runs Out
  3. policyDiverValue-Bench: Measuring LLM Value Divergence Across 74 Markets
  4. infraTask-Based OAuth Consent: Scoping AI Agent Permissions Per Action
  5. industryAudio Token Compression: Cutting Voice LLM Inference Costs
  6. modelsDeepSeek v4 Flash Vision: Routing Images Without Verified Pricing
  7. agentsClaude Code Weekly Limits Promo Ends in August 2026: Budgeting for Agent Teams
  8. modelsDeepSeek 32B on RTX 3090: Tokens per Second by Quant and Context
  9. modelsPTXBench: LLMs Can Port GPU Kernels, But Not Beat Tuned Libraries
  10. agentsVibe Coding vs Control: How Developers Actually Used AI Coding Agents
  11. policyAndroid Privacy Policies vs. Runtime Logs: A 0.4% Alignment Study
  12. industryLLM Data Center Control: Why Advisory Beats Closed-Loop
  13. infraV8 Isolates vs MicroVMs vs Wasm: Where Spectre Still Draws the Line
  14. modelsCan LLMs Reuse Another Model's KV Cache? What Cross-Model Transfer Shows
  15. infraFine-Tuning DeepSeek Without NVIDIA: What the Ascend SuperPOD Run Shows
  16. agentsMulti-Agent or Single-Agent LLM: What Skill Distillation Actually Costs
  17. agentsMulti-Agent LLM Systems Drift Into Misaligned Communication Over Long Horizons
  18. infraCloudflare WebMCP: The Security Baseline for Agent-Ready Sites
  19. policyWhy Machine Unlearning Can't Certify GDPR Erasure
  20. agentsSizing Agent Memory: A Capacity Planning Rubric for Long-Horizon LLMs
  21. infraCloudflare AI Search vs Self-Hosted RAG: Where the Build-vs-Buy Line Lands
  22. infraCloudflare H1 2026 DDoS Report: DNS Floods and Sizing Past 1 Tbps
  23. industryLLM Conflict-of-Interest Benchmark: Sponsor Bias as a Measurable Failure Mode
  24. policyRA-Bench: Why Deepfake Detectors Fail on Re-Shared Crisis Video
browse all 698 articles →