Groundy — independent coverage of developer tools, infrastructure, and platforms
DeepSeek 32B on RTX 3090: Tokens per Second by Quant and Context
Bandwidth math predicts 32-39 tok/s for Q4_K_M DeepSeek 32B on an RTX 3090 at 8k context. 32k context overflows 24GB VRAM. No measured benchmarks exist, so every figure is.
modelsPTXBench: LLMs Can Port GPU Kernels, But Not Beat Tuned Libraries
PTXBench shows LLMs can generate architecture-specific PTX for H100 and B200, but no model matches frontier libraries. Use it to cut hot-loop porting costs, not to replace.
Vibe Coding vs Control: How Developers Actually Used AI Coding Agents
New data shows pro devs use AI coding agents as collaborators, not delegates. Weight review surfaces and checkpoints over autonomous run length when buying agent tooling.
policyAndroid Privacy Policies vs. Runtime Logs: A 0.4% Alignment Study
A study of 1,000 Android apps found only 0.4% align privacy policies with runtime logs. 67.6% leak sensitive data. Shift compliance from document review to CI log-inventory.
industryLLM Data Center Control: Why Advisory Beats Closed-Loop
SemPlan data shows LLM planners fail on 74% of operational queries. Keep actuation in classical MPC, use LLMs only for advisory, and demand third-party verification before.
infraV8 Isolates vs MicroVMs vs Wasm: Where Spectre Still Draws the Line
Cloudflare's 2026 Spectre revisit is unverified. This guide maps V8 isolates, Firecracker microVMs, and Wasm sandboxes to define a data-classification rule for edge secrets.
modelsCan LLMs Reuse Another Model's KV Cache? What Cross-Model Transfer Shows
A new preprint proposes a lightweight reader to reuse KV caches across models, but evidence suggests caches remain model-bound. Learn when recompute beats adaptation.
infraFine-Tuning DeepSeek Without NVIDIA: What the Ascend SuperPOD Run Shows
SLAI T-Rex reports 34.22% MFU for full-parameter DeepSeek-V4 post-training on Ascend. The single-cluster result narrows the CUDA moat for fine-tuning, but cost data remains.
- modelsKimi K3: 2.8T Parameters, MoE Routing, and Self-Hosting Reality
- modelsKimi K3 vs Qwen3.8 Max: Routing Strategy for July 2026
- modelsQwen3.8 Max Release Audit: API, Open Weights, and the License Catch
- devtoolsDrizzle vs Prisma: Choosing a TypeScript ORM in 2026
- infrapgvector vs Pinecone vs Qdrant: Picking a Vector Database in 2026
- infraMLX vs llama.cpp on Apple Silicon: Which Runtime to Use for Local LLM Inference
- modelsChinese AI Models Compared: DeepSeek, Qwen, Kimi, Doubao, and Ernie
- modelsGLM-5.2 Benchmarks: What 62.1% SWE-bench Pro and 99.2% AIME Actually Mean
- cultureEU's 2027 Replaceable Battery Mandate: What It Means for Phone Buyers and Repairers Right Now
- devtoolsGitHub Copilot vs Cursor vs Claude Code: The 2026 AI Coding Showdown
- industryCursor's Meteoric Rise: Inside the AI Editor Hitting $300M ARR
- infraPrefill-Decode Disaggregation: The Architecture Shift Redefining LLM Serving
- devtoolsClaude Code in GitHub Actions: A Complete Guide to Automated PR Fixes
- industryAnthropic Ends Flat-Fee Enterprise Claude, Enforces Per-Token Billing
- ossKeep Android Open: F-Droid's Fight Against a Locked-Down Mobile Future
- aug 21modelsDeepSeek 32B on RTX 3090: Tokens per Second by Quant and Context
- aug 21modelsPTXBench: LLMs Can Port GPU Kernels, But Not Beat Tuned Libraries
- aug 21agentsVibe Coding vs Control: How Developers Actually Used AI Coding Agents
- aug 21policyAndroid Privacy Policies vs. Runtime Logs: A 0.4% Alignment Study
- aug 21industryLLM Data Center Control: Why Advisory Beats Closed-Loop
- aug 21infraV8 Isolates vs MicroVMs vs Wasm: Where Spectre Still Draws the Line
- aug 20modelsCan LLMs Reuse Another Model's KV Cache? What Cross-Model Transfer Shows
- aug 20infraFine-Tuning DeepSeek Without NVIDIA: What the Ascend SuperPOD Run Shows
- aug 20agentsMulti-Agent or Single-Agent LLM: What Skill Distillation Actually Costs
- aug 20devtoolsTraining a Personal Coding Assistant: GPU Cost vs a Copilot Seat
- aug 20agentsMulti-Agent LLM Systems Drift Into Misaligned Communication Over Long Horizons
- aug 19infraCloudflare WebMCP: The Security Baseline for Agent-Ready Sites
- aug 19policyWhy Machine Unlearning Can't Certify GDPR Erasure
- aug 19agentsSizing Agent Memory: A Capacity Planning Rubric for Long-Horizon LLMs
- aug 19infraCloudflare AI Search vs Self-Hosted RAG: Where the Build-vs-Buy Line Lands
- aug 18devtoolsVCoT-Bench: Why AI Rust Verification Fails Merge Gates
- aug 18infraCloudflare H1 2026 DDoS Report: DNS Floods and Sizing Past 1 Tbps
- aug 18industryLLM Conflict-of-Interest Benchmark: Sponsor Bias as a Measurable Failure Mode
- aug 18policyRA-Bench: Why Deepfake Detectors Fail on Re-Shared Crisis Video
- aug 18agentsWhen Spec-First Agents Dismantle Invariants: A Governance Case Study
- aug 18devtoolsRust GPU Offload: arXiv 2608.13759 Analysis for Rust Teams
- aug 18agentsCloudflare Kitesurf: V8 Isolates vs Containers for Agent Browsing
- aug 18policyDario Amodei on AI Regulation: Frontier Labs as Their Own Lobbyists
- aug 18agentsClaude Code Skills vs Model Weights: Where Should Agent Skills Live?
- aug 17devtoolsCursor Origin vs GitHub: The Real Cost of Switching Repo Hosts
- aug 17devtoolsTreat AI Autofix as Untrusted Input: Merge Gates and CI Scoping
- aug 01modelsDeCRIM: Decompose Constraints to Stop Silent Drops in Agent Outputs
- aug 01devtoolsKimi K3 Local Inference: Why 2.8T Parameters Break the Consumer RAM Floor
- jul 31modelsOperator-Level Triage for Silent Mixed-Precision Instability
- jul 31modelsBeyondUncertainty: Weak Confidence Signal for RAG Routing, Not Calibration
- jul 31policyPublic Sector AI Procurement Must Shift from Model Certification to Task Authorization
- jul 30infradaVinci-kernel shifts the RL kernel bottleneck from reward shaping to skill libraries
- jul 30agentsCloudflare Precursor: Behavioral Detection for AI Agents
- jul 30devtoolsProvenance as a CI Gate: Attributing Agent-Authorship in Code
- jul 30devtoolsOHTTP CLI: Stateless Privacy for Agents vs VPN and Tor
- jul 30modelsKimi K3 on M1 Max: Bandwidth, Not Capacity, Limits Local MoE Inference
- jul 30policyWhy Written AI Policies Fail to Control Agent Behavior
- jul 29agentsWhy Multi-Agent LLM Delegation Concentrates Risk
- jul 29modelsKimi Linear Cuts KV Cache 75% but Recall Remains the Binding Constraint
- jul 29policyWhy Vendor Model Cards Fail Clinical Ethics Procurement
- jul 29agentsHarness vs Scaffold: Why Claude Code and LangGraph Are Not Interchangeable
- jul 28devtoolsFine-Tuning vs RAG for Internal APIs: StarCoder2 Constraints
- jul 28agentsMCP Tool Discovery Moves From Hardcoded Config to Runtime Agent Search
- jul 28infraCalibrated LLM Monitoring: Conformal Prediction with Drift Detection
- jul 28devtoolsMellum2 Unverified: Why MoE Active Parameters Matter More Than Total Size
- jul 28agentsCode-as-Action Agents Beat GAIA But Require Runtime Sandboxing
- jul 28policyTRIDENT Benchmark: LLM Safety Gaps in Finance, Medicine, and Law
- jul 28modelsWhy Chat Leaderboards Do Not Predict Image Quality
- jul 27devtoolsPyPI Wheel Reproducibility: 15% Byte-Identical, 79% Source-Equivalent
- jul 27modelsKimi K3 Procurement: Governance Review Over Phantom Government Assessments