articles
all articles
models
- aug 01DeCRIM: Decompose Constraints to Stop Silent Drops in Agent Outputs
- jul 31Operator-Level Triage for Silent Mixed-Precision Instability
- jul 31BeyondUncertainty: Weak Confidence Signal for RAG Routing, Not Calibration
- jul 30Kimi K3 on M1 Max: Bandwidth, Not Capacity, Limits Local MoE Inference
- jul 29Kimi Linear Cuts KV Cache 75% but Recall Remains the Binding Constraint
infra
- jul 30daVinci-kernel shifts the RL kernel bottleneck from reward shaping to skill libraries
- jul 28Calibrated LLM Monitoring: Conformal Prediction with Drift Detection
- jul 27Cloudflare AI Crawler Controls: Block, Charge, or Allow Bots Per Route
- jul 25Postgres LISTEN/NOTIFY Scales: When to Drop Redis for Job Fan-Out
- jul 24Tailscale on Azure: Measure Direct vs DERP Routing to Control Latency and Egress
devtools
- aug 01Kimi K3 Local Inference: Why 2.8T Parameters Break the Consumer RAM Floor
- jul 30Provenance as a CI Gate: Attributing Agent-Authorship in Code
- jul 30OHTTP CLI: Stateless Privacy for Agents vs VPN and Tor
- jul 28Fine-Tuning vs RAG for Internal APIs: StarCoder2 Constraints
- jul 28Mellum2 Unverified: Why MoE Active Parameters Matter More Than Total Size
agents
- jul 30Cloudflare Precursor: Behavioral Detection for AI Agents
- jul 29Why Multi-Agent LLM Delegation Concentrates Risk
- jul 29Harness vs Scaffold: Why Claude Code and LangGraph Are Not Interchangeable
- jul 28MCP Tool Discovery Moves From Hardcoded Config to Runtime Agent Search
- jul 28Code-as-Action Agents Beat GAIA But Require Runtime Sandboxing
feed
- aug 01modelsDeCRIM: Decompose Constraints to Stop Silent Drops in Agent Outputs
- aug 01devtoolsKimi K3 Local Inference: Why 2.8T Parameters Break the Consumer RAM Floor
- jul 31modelsOperator-Level Triage for Silent Mixed-Precision Instability
- jul 31modelsBeyondUncertainty: Weak Confidence Signal for RAG Routing, Not Calibration
- jul 31policyPublic Sector AI Procurement Must Shift from Model Certification to Task Authorization
- jul 30infradaVinci-kernel shifts the RL kernel bottleneck from reward shaping to skill libraries
- jul 30agentsCloudflare Precursor: Behavioral Detection for AI Agents
- jul 30devtoolsProvenance as a CI Gate: Attributing Agent-Authorship in Code
- jul 30devtoolsOHTTP CLI: Stateless Privacy for Agents vs VPN and Tor
- jul 30modelsKimi K3 on M1 Max: Bandwidth, Not Capacity, Limits Local MoE Inference
- jul 30policyWhy Written AI Policies Fail to Control Agent Behavior
- jul 29agentsWhy Multi-Agent LLM Delegation Concentrates Risk
- jul 29modelsKimi Linear Cuts KV Cache 75% but Recall Remains the Binding Constraint
- jul 29policyWhy Vendor Model Cards Fail Clinical Ethics Procurement
- jul 29agentsHarness vs Scaffold: Why Claude Code and LangGraph Are Not Interchangeable
- jul 28devtoolsFine-Tuning vs RAG for Internal APIs: StarCoder2 Constraints
- jul 28agentsMCP Tool Discovery Moves From Hardcoded Config to Runtime Agent Search
- jul 28infraCalibrated LLM Monitoring: Conformal Prediction with Drift Detection
- jul 28devtoolsMellum2 Unverified: Why MoE Active Parameters Matter More Than Total Size
- jul 28agentsCode-as-Action Agents Beat GAIA But Require Runtime Sandboxing
- jul 28policyTRIDENT Benchmark: LLM Safety Gaps in Finance, Medicine, and Law
- jul 28modelsWhy Chat Leaderboards Do Not Predict Image Quality
- jul 27devtoolsPyPI Wheel Reproducibility: 15% Byte-Identical, 79% Source-Equivalent
- jul 27modelsKimi K3 Procurement: Governance Review Over Phantom Government Assessments
- jul 27policyEU AI Act Traceability: Why ML Pipelines Fail Conformity Assessment
- jul 27infraCloudflare AI Crawler Controls: Block, Charge, or Allow Bots Per Route
- jul 27agentsWhy Agent Security Tests Must Audit Full Trajectories, Not Single Turns
- jul 26agentsx402 Per-Call Payments: Agent Wallet Custody and Replay Risks
- jul 26modelsDeepSeek Compute Leak: Why Open-Weight Routing Needs a Swap Path
- jul 26modelsContext Ordering Beats Window Size for Long-Context Agents
- jul 26devtoolsCHRONO-RESOLUTION: npm, PyPI, and crates.io lockfile drift measured at release points
- jul 25infraPostgres LISTEN/NOTIFY Scales: When to Drop Redis for Job Fan-Out
- jul 25devtoolsCLI-Tool-Bench: Why Patch Leaderboards Fail for 0-to-1 Code Generation
- jul 25policyImplicit Bias in LLMs Passes NYC and EU Audits
- jul 25agentsCodeRabbit Review Study: 56% Rejection Rate Demands Targeted Scoping
- jul 24modelsDiffusion LLMs: Training Cost, Not Parallel Decoding, Drives Deployment
- jul 24infraTailscale on Azure: Measure Direct vs DERP Routing to Control Latency and Egress
- jul 24infraAccelerate vs Megatron Core: The Model Size Curve for Distributed Training
- jul 24agentsLLM Agents Ignore Mid-Flight Halt Signals: 0 of 40 Trials Stopped
- jul 24modelsOpen-Weight Routers vs Fable 5: The Routing Math That Actually Matters
- jul 24agentsAgent-First CLIs: Why GitHub, npm, and PyPI Must Publish Machine-Readable Contracts
- jul 23policyEU Driver Monitoring: GDPR Compliance Without Consent
- jul 23infraWhy cgroups, not permission prompts, bound AI agent CPU and memory
- jul 23modelsDeepSeek-V4 1M Context vs RAG: Why Retrieval Stays
- jul 23modelsQwen-Image-3.0 Does Not Exist: Why Self-Hosting Image Models Is Premature
- jul 22policyEU AI Act bans emotion AI in schools, but permits it where models fail
- jul 22agentsMCP and AGENTS.md Standardize Context, Not Agent Coordination
- jul 22industryAI Search's One-Answer Rule: When Better Content Makes Search Worse
- jul 22agentsCursor's Swarm Math: When Cheap Agents Save Money and When They Fail
- jul 21agentsRuntime monitoring beats alignment for agent-to-agent coercion
- jul 21infravLLM Configs Shift Energy, Latency, and Accuracy: A 9,000-Run Study
- jul 21policyHuggingFace vs GitHub Models vs Replicate: Policy Compliance for Uploaders
- jul 20modelsKimi K3: 2.8T Parameters, MoE Routing, and Self-Hosting Reality
- jul 20modelsKimi K3 vs Qwen3.8 Max: Routing Strategy for July 2026
- jul 20agentsCloudflare's Agent Stack: Edge Trust, Identity, and Metering
- jul 20modelsQwen3.8 Max Preview: Missing Benchmarks, Weights, and Pricing
- jul 20policySAMark Text Watermarking: Paraphrase Robustness and the Policy Gap
- jul 20industryHuggingFace's $100M Series C Locks Teams Into Deployment
- jul 19modelsHuggingFace 100x Inference: Generalizable vs Platform-Locked Optimizations
- jul 19agentsLM Studio Bionic vs Claude Code: Local-First vs Cloud Agent Tradeoffs
- jul 19infraAWS Estimated Billing Was Off by $1.7B: Reconciling Actual Cloud Spend
- jul 19modelsKimi K3 Code Arena Rank: Self-Hosting Cost Math for Coding Agents
- jul 19infraCloudflare Attribution vs Custom Logs: The Per-Path AI Crawler Decision
- jul 18devtoolsGrok CLI uploads entire workspace to GCS by default, independent of model reads
- jul 18infraSpectral Compute CUDA Translation: vLLM Procurement vs Porting Cost
- jul 18infraRunning MiniCPM-V-4.6 on Fermi: What 6 GB of VRAM Forces
- jul 17agentsCan a Malicious AGENTS.md File Compromise Your Coding Agent? A Threat Model
- jul 17infraLLM Inference Without a GPU: Pure CPU vs Hybrid CPU-GPU Scheduling
- jul 17cultureWhen Cultural LLM Alignment Gets a Positive Target, Who Writes the Spec?
- jul 17agentsHow GitHub Projects Actually Adopt Coding Agents: New Empirical Data
- jul 17infraRL-Found CUDA Kernels Beat cuBLAS: Kernel Tuning Shifts to Reward Design
- jul 17policyA Digital Twin Can Validate AV Safety, but No Regulator Accepts the Evidence
- jul 17oss62.7% of Linux Foundation Repos Still Carry Non-Inclusive Terms, and LLMs Are Learning Them
- jul 17policyWhy EU AI Act Monitoring Will Miss Discontinuous LLM Alignment Failures
- jul 16devtoolsDrizzle vs Prisma: Choosing a TypeScript ORM in 2026
- jul 15infrapgvector vs Pinecone vs Qdrant: Picking a Vector Database in 2026
- jul 14modelsCan Tool-Adaptive LLM Rerankers Improve RAG Without Always Calling Tools?
- jul 14securityNetInjectBench: Prompt Injection Becomes a Network Availability Problem
- jul 14infraOllama vs LM Studio: Picking a Local LLM Runtime in 2026
- jul 14infraBeyond Quantization: LLM Efficiency Is Now a Memory-Bandwidth Problem
- jul 14modelsDoes Speculative Decoding with Progressive Tree Drafting Cut LLM Latency?
- jul 14agentsWhy CLI Coding Agents Derail Mid-Run, Not at the First Mistake
- jul 14industryCoreWeave, Nebius, and the GPU Debt Loop Behind Your Inference Bill
- jul 14ossBERTopic vs LDA: Hosted Embeddings Erased the GPU Cost Argument
- jul 13industryOpenAI's Statsig Acquisition Turns Feature Flags Into a Lock-In Question
- jul 13infraHow Sparse LLM Weights Cut GPU Inference Cost Without Quantization
- jul 13securityType-Checking LLM Agent Secrets: Why Information Flow Needs a Calculus
- jul 13securityVercel SAMLStorm Protection Misses Self-Hosted Identity Providers
- jul 13agentsTest-Time Scaling Cost Falls as PRMs Reuse Generator KV-Cache
- jul 13ossRISCBoy Open-Sources a Handheld Console Designed From Scratch
- jul 13ossSoofi S: Sovereign AI Is Cheap to Adopt, Expensive to Sustain
- jul 13agentsClaude Code Skills vs Cursor Rules vs MCP: How Agent Skill Systems Compare
- jul 12agentsTTHE: Test-Time Harness Evolution Changes the Test-Code Contract for Coding Agents
- jul 12devtoolsGrok Build CLI Sends File Listings and Code Fragments to xAI, Widening Endpoint Trust Boundaries
- jul 12agentsGit-for-Data for Agentic Lakehouses: Why Agents Need Versioned State
- jul 12devtoolsVercel Adds Zero-Config Node Server Deploys: Hono's Pattern Goes Mainstream
- jul 11securityContext-Aware Prompt Injection Defenses for LLM Agents: Why Static Filters Fail
- jul 11devtoolsOpenAI's Codex Refresh: The Upgrade That Puts Pressure on Cursor and Claude Code
- jul 11securityFinal-Token vs Full-Sequence Safety Probes: Why LLM Red Teams Need Both
- jul 11securitys1ngularity Supply Chain Attack Hits Nx: What Monorepo Teams Should Patch