infra
infrastructure & runtime
archive
- Vercel In-Function Concurrency: What It Changes for Stateful Node.js
- Running LLMs on AMD GPUs With ROCm: What Actually Works
- Cloudflare Meerkat: What Globally Distributed Consensus Costs at the Edge
- Vercel CDN Now Honors External Origin Cache-Control: Audit Your Headers
- AI Found Real Bugs in Cloudflare's Circl Crypto Library
- Pruning RAG Context: What to Cut Before the LLM Sees It
- Cloudflare's x402 Gateway: What Per-Request API Billing Actually Needs
- Doubao 2.1 Pro: What 180 Trillion Daily Tokens Means for Inference Infrastructure
- Every CUDA Kernel Pays a Launch Tax: The Host-to-Device Walkthrough
- Vercel Montreal Region: Audit Residency Before You Migrate
- GLM-5.2 on vLLM and Ascend: Open Weights Beyond NVIDIA
- How Vercel Runs Its Own CDN in Front of Discourse: A Self-Dogfooding Case Study
- Vercel Runtime Logs Surface CDN Cache Hits, Not the Eviction Cause
- Multimodal Knowledge Graph RAG vs Vector RAG: What MKG-RAG-Bench Shows
- Vercel Observability Now Tracks Redirects and Rewrites Beside Function Errors
- Cloudflare Workflows Saga Rollbacks: Compensating Actions in Serverless Orchestration
- Static Corpus RAG: The Bible Case for Separating Churn from Algorithm Complexity
- Vercel's KIKO Milano Black Friday Case Study: What the Scaling Claims Skip
- Vercel Postgres vs Neon vs Supabase: When the Bundled DB Wins
- Fine-Tuning a 20B LLM With RLHF on a 24GB GPU: What Fits
- Vercel Flat Rate CDN Beta: Break-Even Math for Spiky Workloads, Tax for the Rest
- Where DeepSeek Weights Actually Run on Vercel's AI Gateway
- Vercel's Anti-Lock-In Pitch: What the Open-Source Bet Still Locks In
- Vercel Adds Tag-Based CDN Cache Invalidation: Surrogate Keys at the Edge
- GLM 5.2 Fast on Vercel AI Gateway: What Routing Through Wafer Actually Buys
- Vercel CDN Cache Tags vs Path Purging: When Tag Invalidation Wins
- Prisma Joins the Vercel Marketplace: The ORM Becomes the Database Vendor
- OpenAI on AWS Bedrock: Routing Math to Run Before You Move Traffic
- Vercel's Function Observability: What Native Metrics Replace and What They Don't
- AWS Databases on the Vercel Marketplace: The Cross-Cloud Latency Tax
- Turso on the Vercel Marketplace: Edge SQLite vs the Serverless Connection Pool
- Vercel on the AWS Marketplace: What the Listing Does to Procurement and Lock-In
- Serving Cold MoE Models: CrossPool Disaggregates KV Cache and Weights
- Vercel's In-Function Concurrency: What It Does to Cold Starts and Billing
- Poisoning a RAG Retriever: How Conflict-Aware Edits Inject False Knowledge
- Vercel Raised Its CDN Origin Timeout to Two Minutes: What Breaks First
- Gradio-Lite Runs Model Inference in the Browser via Pyodide, No Server
- Cloudflare AI Gateway Adds Spend Limits to Cap the Runaway Inference Bill
- Vercel Now Honors stale-if-error: Serving Stale Cache When the Origin Dies
- Vercel's Manual CDN Purge API: Cache Control Without a Redeploy
- Cloudflare Now Routes Public Traffic to Private Apps via DNS, No VPN
- GitHub's AI Capacity Crunch Pushes Microsoft to Rent AWS Compute
- Cloudflare's Temporary Accounts Give AI Agents Disposable Credentials
- Running Long-Context Agents on a 4-Bit KV Cache: Where Accuracy Breaks
- When LLM-Generated CUDA Kernels Pass Tests but Get the Math Wrong
- Running GLM-5.2 at Home: SGLang, vLLM, Transformers, and KTransformers Setup Guide
- AWS Bedrock Now Requires Data Sharing for Mythos: The Self-Hosting Calculus
- vLLM Cold Start Latency: Why Scale-to-Zero LLM Serving Stalls
- The Vercel-AWS Deal Reveals Where AI Inference Runs
- Running RAG on a Snapdragon NPU: The On-Device Retrieval Tradeoff