Groundy — independent coverage of developer tools, infrastructure, and platforms
daVinci-kernel shifts the RL kernel bottleneck from reward shaping to skill libraries
The daVinci-kernel preprint argues RL-tuned CUDA kernels plateau because skill libraries are mis-architected, not because rewards are mis-shaped. It co-evolves skill.
agentsCloudflare Precursor: Behavioral Detection for AI Agents
Cloudflare Precursor reportedly shifts AI agent detection from spoofable headers to continuous behavioral signals. This article verifies the claims and outlines operator.
Provenance as a CI Gate: Attributing Agent-Authorship in Code
Git credits the merger, not the model. F(AI)2R proposes a PROV-O provenance graph gated by CI to record agent authorship and human verification, preventing blame collapse in.
devtoolsOHTTP CLI: Stateless Privacy for Agents vs VPN and Tor
Cloudflare's OHTTP protocol splits request identity from content across two parties. This CLI approach offers stateless privacy for agents, shifting key management burden to.
modelsKimi K3 on M1 Max: Bandwidth, Not Capacity, Limits Local MoE Inference
Running Kimi K3 on an M1 Max proves local MoE feasibility via expert offloading, but unified memory bandwidth caps throughput. Treat this as a prototyping tool, not a serving.
policyWhy Written AI Policies Fail to Control Agent Behavior
HANDBOOK.md benchmark shows frontier agents follow 124-page policies on only 36.2% of trials. Governance requires runtime enforcement, not documentation.
agentsWhy Multi-Agent LLM Delegation Concentrates Risk
Splitting tasks across planner and worker agents concentrates risk because principals cannot observe hidden actions. Base models defect. Teams must add verifiable execution.
modelsKimi Linear Cuts KV Cache 75% but Recall Remains the Binding Constraint
Kimi Linear cuts KV cache 75% and boosts decode 6x for 1M-token contexts, but recall stays the binding constraint. Route memory-bound jobs only after benchmarking in-context.
- modelsKimi K3: 2.8T Parameters, MoE Routing, and Self-Hosting Reality
- modelsKimi K3 vs Qwen3.8 Max: Routing Strategy for July 2026
- modelsQwen3.8 Max Preview: Missing Benchmarks, Weights, and Pricing
- devtoolsDrizzle vs Prisma: Choosing a TypeScript ORM in 2026
- infrapgvector vs Pinecone vs Qdrant: Picking a Vector Database in 2026
- modelsGLM-5.2 Benchmarks: What 62.1% SWE-bench Pro and 99.2% AIME Actually Mean
- modelsChinese AI Models Compared: DeepSeek, Qwen, Kimi, Doubao, and Ernie
- infraMLX vs llama.cpp on Apple Silicon: Which Runtime to Use for Local LLM Inference
- modelsGLM-5.2 on Terminal-Bench 2.1: Strengths, Gaps, and How to Route Real Coding Tasks
- industryCursor's Meteoric Rise: Inside the AI Editor Hitting $300M ARR
- infraPrefill-Decode Disaggregation: The Architecture Shift Redefining LLM Serving
- devtoolsRunning GLM-5.2 in Cursor, Cline, and Roo Code: Migration Checklist and Gotchas
- devtoolsClaude Code in GitHub Actions: A Complete Guide to Automated PR Fixes
- agentsSuperpowers: The Agentic Framework Replacing Your Dev Process
- infraRunning GLM-5.2 at Home: SGLang, vLLM, Transformers, and KTransformers Setup Guide
- jul 30infradaVinci-kernel shifts the RL kernel bottleneck from reward shaping to skill libraries
- jul 30agentsCloudflare Precursor: Behavioral Detection for AI Agents
- jul 30devtoolsProvenance as a CI Gate: Attributing Agent-Authorship in Code
- jul 30devtoolsOHTTP CLI: Stateless Privacy for Agents vs VPN and Tor
- jul 30modelsKimi K3 on M1 Max: Bandwidth, Not Capacity, Limits Local MoE Inference
- jul 30policyWhy Written AI Policies Fail to Control Agent Behavior
- jul 29agentsWhy Multi-Agent LLM Delegation Concentrates Risk
- jul 29modelsKimi Linear Cuts KV Cache 75% but Recall Remains the Binding Constraint
- jul 29policyWhy Vendor Model Cards Fail Clinical Ethics Procurement
- jul 29agentsHarness vs Scaffold: Why Claude Code and LangGraph Are Not Interchangeable
- jul 28devtoolsFine-Tuning vs RAG for Internal APIs: StarCoder2 Constraints
- jul 28agentsMCP Tool Discovery Moves From Hardcoded Config to Runtime Agent Search
- jul 28infraCalibrated LLM Monitoring: Conformal Prediction with Drift Detection
- jul 28devtoolsMellum2 Unverified: Why MoE Active Parameters Matter More Than Total Size
- jul 28agentsCode-as-Action Agents Beat GAIA But Require Runtime Sandboxing
- jul 28policyTRIDENT Benchmark: LLM Safety Gaps in Finance, Medicine, and Law
- jul 28modelsWhy Chat Leaderboards Do Not Predict Image Quality
- jul 27devtoolsPyPI Wheel Reproducibility: 15% Byte-Identical, 79% Source-Equivalent
- jul 27modelsKimi K3 Procurement: Governance Review Over Phantom Government Assessments
- jul 27policyEU AI Act Traceability: Why ML Pipelines Fail Conformity Assessment
- jul 27infraCloudflare AI Crawler Controls: Block, Charge, or Allow Bots Per Route
- jul 27agentsWhy Agent Security Tests Must Audit Full Trajectories, Not Single Turns
- jul 26agentsx402 Per-Call Payments: Agent Wallet Custody and Replay Risks
- jul 26modelsDeepSeek Compute Leak: Why Open-Weight Routing Needs a Swap Path
- jul 26modelsContext Ordering Beats Window Size for Long-Context Agents
- jul 26devtoolsCHRONO-RESOLUTION: npm, PyPI, and crates.io lockfile drift measured at release points
- jul 25infraPostgres LISTEN/NOTIFY Scales: When to Drop Redis for Job Fan-Out
- jul 25devtoolsCLI-Tool-Bench: Why Patch Leaderboards Fail for 0-to-1 Code Generation
- jul 25policyImplicit Bias in LLMs Passes NYC and EU Audits
- jul 25agentsCodeRabbit Review Study: 56% Rejection Rate Demands Targeted Scoping
- jul 24modelsDiffusion LLMs: Training Cost, Not Parallel Decoding, Drives Deployment
- jul 24infraTailscale on Azure: Measure Direct vs DERP Routing to Control Latency and Egress
- jul 24infraAccelerate vs Megatron Core: The Model Size Curve for Distributed Training
- jul 24agentsLLM Agents Ignore Mid-Flight Halt Signals: 0 of 40 Trials Stopped
- jul 24modelsOpen-Weight Routers vs Fable 5: The Routing Math That Actually Matters
- jul 24agentsAgent-First CLIs: Why GitHub, npm, and PyPI Must Publish Machine-Readable Contracts
- jul 23policyEU Driver Monitoring: GDPR Compliance Without Consent
- jul 23infraWhy cgroups, not permission prompts, bound AI agent CPU and memory
- jul 23modelsDeepSeek-V4 1M Context vs RAG: Why Retrieval Stays
- jul 23modelsQwen-Image-3.0 Does Not Exist: Why Self-Hosting Image Models Is Premature
- jul 22policyEU AI Act bans emotion AI in schools, but permits it where models fail
- jul 22agentsMCP and AGENTS.md Standardize Context, Not Agent Coordination
- jul 22industryAI Search's One-Answer Rule: When Better Content Makes Search Worse
- jul 22agentsCursor's Swarm Math: When Cheap Agents Save Money and When They Fail
- jul 21agentsRuntime monitoring beats alignment for agent-to-agent coercion
- jul 21infravLLM Configs Shift Energy, Latency, and Accuracy: A 9,000-Run Study
- jul 21policyHuggingFace vs GitHub Models vs Replicate: Policy Compliance for Uploaders
- jul 20modelsKimi K3: 2.8T Parameters, MoE Routing, and Self-Hosting Reality
- jul 20modelsKimi K3 vs Qwen3.8 Max: Routing Strategy for July 2026
- jul 20agentsCloudflare's Agent Stack: Edge Trust, Identity, and Metering