Groundy — independent coverage of developer tools, infrastructure, and platforms
Why Machine Unlearning Can't Certify GDPR Erasure
Machine unlearning cannot certify GDPR erasure because no shared protocol proves weight removal. Use retrain-and-attest pipelines to satisfy Article 17 deletion requests.
agentsSizing Agent Memory: A Capacity Planning Rubric for Long-Horizon LLMs
Measure input tokens, gate stage boundaries, and size backbones before adding retrieval. A capacity planning rubric for agent memory based on August 2026 arXiv data.
Cloudflare AI Search vs Self-Hosted RAG: Where the Build-vs-Buy Line Lands
Cloudflare's AI Search lacks verifiable pricing and access control docs. For Postgres 13+ teams, pgvector offers a documented, self-hosted alternative that avoids vendor.
devtoolsVCoT-Bench: Why AI Rust Verification Fails Merge Gates
VCoT-Bench shows frontier LLMs struggle with Rust verification chains. AI code needs compiler gates, property tests, and human review, not model judgment.
infraCloudflare H1 2026 DDoS Report: DNS Floods and Sizing Past 1 Tbps
Cloudflare's H1 2026 DDoS report claims DNS floods exceeded 1 Tbps. This article verifies the structural risks of UDP amplification and anycast absorption, while noting the.
industryLLM Conflict-of-Interest Benchmark: Sponsor Bias as a Measurable Failure Mode
arXiv 2604.08525 benchmarks LLM sponsor bias. GPT 5.1 disrupted purchases 94% of the time. Treat conflict of interest as a measurable failure mode for eval suites and.
policyRA-Bench: Why Deepfake Detectors Fail on Re-Shared Crisis Video
RA-Bench finds no detector family generalizes on re-shared crisis video. Automated detection cannot carry takedown commitments alone. Provenance signals and human review must.
agentsWhen Spec-First Agents Dismantle Invariants: A Governance Case Study
A specification-first coding agent eroded a core invariant across 189 files without review. The fix is executable conformance checks, not more human oversight.
- modelsKimi K3: 2.8T Parameters, MoE Routing, and Self-Hosting Reality
- modelsKimi K3 vs Qwen3.8 Max: Routing Strategy for July 2026
- modelsQwen3.8 Max Release Audit: API, Open Weights, and the License Catch
- devtoolsDrizzle vs Prisma: Choosing a TypeScript ORM in 2026
- infrapgvector vs Pinecone vs Qdrant: Picking a Vector Database in 2026
- infraMLX vs llama.cpp on Apple Silicon: Which Runtime to Use for Local LLM Inference
- modelsChinese AI Models Compared: DeepSeek, Qwen, Kimi, Doubao, and Ernie
- modelsGLM-5.2 Benchmarks: What 62.1% SWE-bench Pro and 99.2% AIME Actually Mean
- cultureEU's 2027 Replaceable Battery Mandate: What It Means for Phone Buyers and Repairers Right Now
- infraPrefill-Decode Disaggregation: The Architecture Shift Redefining LLM Serving
- industryCursor's Meteoric Rise: Inside the AI Editor Hitting $300M ARR
- devtoolsGitHub Copilot vs Cursor vs Claude Code: The 2026 AI Coding Showdown
- industryAnthropic Ends Flat-Fee Enterprise Claude, Enforces Per-Token Billing
- ossKeep Android Open: F-Droid's Fight Against a Locked-Down Mobile Future
- devtoolsClaude Code Plugins: Anthropic's Official Plugin Ecosystem Explained
- aug 19policyWhy Machine Unlearning Can't Certify GDPR Erasure
- aug 19agentsSizing Agent Memory: A Capacity Planning Rubric for Long-Horizon LLMs
- aug 19infraCloudflare AI Search vs Self-Hosted RAG: Where the Build-vs-Buy Line Lands
- aug 18devtoolsVCoT-Bench: Why AI Rust Verification Fails Merge Gates
- aug 18infraCloudflare H1 2026 DDoS Report: DNS Floods and Sizing Past 1 Tbps
- aug 18industryLLM Conflict-of-Interest Benchmark: Sponsor Bias as a Measurable Failure Mode
- aug 18policyRA-Bench: Why Deepfake Detectors Fail on Re-Shared Crisis Video
- aug 18agentsWhen Spec-First Agents Dismantle Invariants: A Governance Case Study
- aug 18devtoolsRust GPU Offload: arXiv 2608.13759 Analysis for Rust Teams
- aug 18agentsCloudflare Kitesurf: V8 Isolates vs Containers for Agent Browsing
- aug 18policyDario Amodei on AI Regulation: Frontier Labs as Their Own Lobbyists
- aug 18agentsClaude Code Skills vs Model Weights: Where Should Agent Skills Live?
- aug 17devtoolsCursor Origin vs GitHub: The Real Cost of Switching Repo Hosts
- aug 17devtoolsTreat AI Autofix as Untrusted Input: Merge Gates and CI Scoping
- aug 01modelsDeCRIM: Decompose Constraints to Stop Silent Drops in Agent Outputs
- aug 01devtoolsKimi K3 Local Inference: Why 2.8T Parameters Break the Consumer RAM Floor
- jul 31modelsOperator-Level Triage for Silent Mixed-Precision Instability
- jul 31modelsBeyondUncertainty: Weak Confidence Signal for RAG Routing, Not Calibration
- jul 31policyPublic Sector AI Procurement Must Shift from Model Certification to Task Authorization
- jul 30infradaVinci-kernel shifts the RL kernel bottleneck from reward shaping to skill libraries
- jul 30agentsCloudflare Precursor: Behavioral Detection for AI Agents
- jul 30devtoolsProvenance as a CI Gate: Attributing Agent-Authorship in Code
- jul 30devtoolsOHTTP CLI: Stateless Privacy for Agents vs VPN and Tor
- jul 30modelsKimi K3 on M1 Max: Bandwidth, Not Capacity, Limits Local MoE Inference
- jul 30policyWhy Written AI Policies Fail to Control Agent Behavior
- jul 29agentsWhy Multi-Agent LLM Delegation Concentrates Risk
- jul 29modelsKimi Linear Cuts KV Cache 75% but Recall Remains the Binding Constraint
- jul 29policyWhy Vendor Model Cards Fail Clinical Ethics Procurement
- jul 29agentsHarness vs Scaffold: Why Claude Code and LangGraph Are Not Interchangeable
- jul 28devtoolsFine-Tuning vs RAG for Internal APIs: StarCoder2 Constraints
- jul 28agentsMCP Tool Discovery Moves From Hardcoded Config to Runtime Agent Search
- jul 28infraCalibrated LLM Monitoring: Conformal Prediction with Drift Detection
- jul 28devtoolsMellum2 Unverified: Why MoE Active Parameters Matter More Than Total Size
- jul 28agentsCode-as-Action Agents Beat GAIA But Require Runtime Sandboxing
- jul 28policyTRIDENT Benchmark: LLM Safety Gaps in Finance, Medicine, and Law
- jul 28modelsWhy Chat Leaderboards Do Not Predict Image Quality
- jul 27devtoolsPyPI Wheel Reproducibility: 15% Byte-Identical, 79% Source-Equivalent
- jul 27modelsKimi K3 Procurement: Governance Review Over Phantom Government Assessments
- jul 27policyEU AI Act Traceability: Why ML Pipelines Fail Conformity Assessment
- jul 27infraCloudflare AI Crawler Controls: Block, Charge, or Allow Bots Per Route
- jul 27agentsWhy Agent Security Tests Must Audit Full Trajectories, Not Single Turns
- jul 26agentsx402 Per-Call Payments: Agent Wallet Custody and Replay Risks
- jul 26modelsDeepSeek Compute Leak: Why Open-Weight Routing Needs a Swap Path
- jul 26modelsContext Ordering Beats Window Size for Long-Context Agents
- jul 26devtoolsCHRONO-RESOLUTION: npm, PyPI, and crates.io lockfile drift measured at release points
- jul 25infraPostgres LISTEN/NOTIFY Scales: When to Drop Redis for Job Fan-Out
- jul 25devtoolsCLI-Tool-Bench: Why Patch Leaderboards Fail for 0-to-1 Code Generation
- jul 25policyImplicit Bias in LLMs Passes NYC and EU Audits
- jul 25agentsCodeRabbit Review Study: 56% Rejection Rate Demands Targeted Scoping
- jul 24modelsDiffusion LLMs: Training Cost, Not Parallel Decoding, Drives Deployment