infra
infrastructure & runtime
archive
- GraphRAG vs VectorRAG: Does the Graph Index Earn Its Cost?
- MiniMax M3 Ships 1M Context and Desktop Control as Open Weights
- DeepSeek-V4 FlashMemory: Sparse Attention for Million-Token Context
- Is Cloudflare's Bot Traffic Surge Real? The Measurement Dispute
- Indexing Images for RAG: kapa.ai's Approach to Multimodal Retrieval
- The RTX Spark Bet on Unified Memory for Local LLMs: Where Bandwidth Caps It
- Pod-Level Remote Attestation in Kubernetes: Confidential Workloads on dstack
- Generating GPU Kernels for Moore Threads Silicon: Can LLMs Break CUDA Lock-In?
- Microsoft's Azure Linux Goes General-Purpose: The Container Base-Image Play
- Cloudflare Acquires VoidZero, the Company Behind Vite's Rust Toolchain
- Putting a Datacenter V100 in a Gaming PC: The Local LLM Math
- Why LLMs Still Botch Kubernetes Manifests: The Training-Data Gap
- Gemma 4 31B on Cloud TPU vs GPU: The Serving Cost Crossover Point
- ObjectCache Moves KV Reuse to S3-Class Storage: Why Layerwise Retrieval Beats Full-Prefix Cache Hits
- vLLM 0.21 Makes Prefill-Decode Disaggregation Actually Practical
- OpenRAG: The Open-Source RAG Platform Challenging Pinecone
- MLX vs llama.cpp on Apple Silicon: Which Runtime to Use for Local LLM Inference
- Prefill-Decode Disaggregation: The Architecture Shift Redefining LLM Serving
- Google LiteRT: Running LLMs on Your Phone Without the Cloud
- Microsoft's BitNet: How 1-Bit LLMs Could Make GPU Farms Obsolete
- WebAssembly AI: Running Models in the Browser
- Tailscale Peer Relays: The Missing Piece for True P2P Networking
- DNS-Persist-01 Validation: Let's Encrypt's Model for Permanent ACME Certificate Authorization
- The Complete Guide to Local LLMs