groundy

infrastructure & runtime

  1. jun 25infraTurso on the Vercel Marketplace: Edge SQLite vs the Serverless Connection Pool
  2. jun 24infraVercel on the AWS Marketplace: What the Listing Does to Procurement and Lock-In
  3. jun 24infraServing Cold MoE Models: CrossPool Disaggregates KV Cache and Weights
  4. jun 24infraVercel's In-Function Concurrency: What It Does to Cold Starts and Billing
  5. jun 24infraPoisoning a RAG Retriever: How Conflict-Aware Edits Inject False Knowledge
  6. jun 24infraVercel Raised Its CDN Origin Timeout to Two Minutes: What Breaks First
  7. jun 24infraGradio-Lite Runs Model Inference in the Browser via Pyodide, No Server
  8. jun 24infraCloudflare AI Gateway Adds Spend Limits to Cap the Runaway Inference Bill
  9. jun 24infraVercel Now Honors stale-if-error: Serving Stale Cache When the Origin Dies
  10. jun 23infraVercel's Manual CDN Purge API: Cache Control Without a Redeploy
  11. jun 23infraCloudflare Now Routes Public Traffic to Private Apps via DNS, No VPN
  12. jun 23infraGitHub's AI Capacity Crunch Pushes Microsoft to Rent AWS Compute
  13. jun 21infraCloudflare's Temporary Accounts Give AI Agents Disposable Credentials
  14. jun 21infraRunning Long-Context Agents on a 4-Bit KV Cache: Where Accuracy Breaks
  15. jun 20infraWhen LLM-Generated CUDA Kernels Pass Tests but Get the Math Wrong
  16. jun 19infraRunning GLM-5.2 at Home: SGLang, vLLM, Transformers, and KTransformers Setup Guide
  17. jun 16infraAWS Bedrock Now Requires Data Sharing for Mythos: The Self-Hosting Calculus
  18. jun 15infravLLM Cold Start Latency: Why Scale-to-Zero LLM Serving Stalls
  19. jun 15infraThe Vercel-AWS Deal Reveals Where AI Inference Runs
  20. jun 11infraRunning RAG on a Snapdragon NPU: The On-Device Retrieval Tradeoff
  21. jun 10infraGraphRAG vs VectorRAG: Does the Graph Index Earn Its Cost?
  22. jun 10infraMiniMax M3 Ships 1M Context and Desktop Control as Open Weights
  23. jun 10infraDeepSeek-V4 FlashMemory: Sparse Attention for Million-Token Context
  24. jun 09infraIs Cloudflare's Bot Traffic Surge Real? The Measurement Dispute
  25. jun 07infraIndexing Images for RAG: kapa.ai's Approach to Multimodal Retrieval
  26. jun 06infraThe RTX Spark Bet on Unified Memory for Local LLMs: Where Bandwidth Caps It
  27. jun 06infraPod-Level Remote Attestation in Kubernetes: Confidential Workloads on dstack
  28. jun 05infraGenerating GPU Kernels for Moore Threads Silicon: Can LLMs Break CUDA Lock-In?
  29. jun 05infraMicrosoft's Azure Linux Goes General-Purpose: The Container Base-Image Play
  30. jun 05infraCloudflare Acquires VoidZero, the Company Behind Vite's Rust Toolchain
  31. jun 05infraPutting a Datacenter V100 in a Gaming PC: The Local LLM Math
  32. may 27infraWhy LLMs Still Botch Kubernetes Manifests: The Training-Data Gap
  33. may 27infraGemma 4 31B on Cloud TPU vs GPU: The Serving Cost Crossover Point
  34. may 26infraObjectCache Moves KV Reuse to S3-Class Storage: Why Layerwise Retrieval Beats Full-Prefix Cache Hits
  35. may 23infravLLM 0.21 Makes Prefill-Decode Disaggregation Actually Practical
  36. mar 27infraOpenRAG: The Open-Source RAG Platform Challenging Pinecone
  37. mar 24infraMLX vs llama.cpp on Apple Silicon: Which Runtime to Use for Local LLM Inference
  38. mar 24infraPrefill-Decode Disaggregation: The Architecture Shift Redefining LLM Serving
  39. mar 15infraGoogle LiteRT: Running LLMs on Your Phone Without the Cloud
  40. mar 13infraMicrosoft's BitNet: How 1-Bit LLMs Could Make GPU Farms Obsolete
  41. feb 28infraWebAssembly AI: Running Models in the Browser
  42. feb 19infraTailscale Peer Relays: The Missing Piece for True P2P Networking
  43. feb 19infraDNS-Persist-01 Validation: Let's Encrypt's Model for Permanent ACME Certificate Authorization
  44. feb 12infraThe Complete Guide to Local LLMs