Cloudflare Now Routes Public Traffic to Private Apps via DNS, No VPN
Cloudflare's private-origins DNS routing lets public hostnames reach RFC 1918 apps without a VPN, but the flag only routes traffic; it does not authenticate callers.
The archive · Page 5 of 6
The serving stack, network fabric, and cloud-account substrate beneath production AI, where every throughput claim collides with rebuild windows, egress invoices, and control-plane risk.
97–120 of 130 articles · Newest first
Cloudflare's private-origins DNS routing lets public hostnames reach RFC 1918 apps without a VPN, but the flag only routes traffic; it does not authenticate callers.
Microsoft reportedly rents AWS compute to keep GitHub's AI inference running after Azure fell behind, signaling that owned infrastructure no longer self-supplies AI load.
Cloudflare's temporary accounts give agents auto-expiring 60-minute credentials, but the launch is an onboarding shortcut, not a scoped security control.
UltraQuant cuts agent time-to-first-token 3.47x with 4-bit KV caching on AMD CDNA4, but its June 2026 preprint omits the accuracy numbers operators need to ship it.
LLM-written CUDA kernels compile, run, and pass smoke tests while returning wrong numerics, so crash-free execution is not enough to trust AI-generated GPU code.

GLM-5.2 weights are live on HuggingFace under MIT license: 753B MoE, 1M-token context, FP8 and BF16 variants. How to pick a deployment framework and model the hardware cost.
AWS Bedrock's provider_data_share gate for Mythos-class models removes the in-AWS data boundary regulated teams bought it for, pushing them toward self-hosted serving.

A June 2026 MLSys paper breaks vLLM cold start into six CPU-bound boot phases, showing why scale-to-zero serving forces operators back into warm GPU pools.
Vercel's May 2026 AWS databases integration clarifies where its AI workloads actually run: inference stays behind external APIs while the stateful tier moves to AWS regions.
End-to-end RAG on the Snapdragon X Elite Hexagon NPU delivers 4x lower latency and 4x less energy than CPU with no quality loss, but soldered memory caps your index size.
A Samsung preprint finds vector retrieval matches GraphRAG on QA tasks at a fraction of the indexing cost, shifting the burden of proof to teams building graph pipelines.
MiniMax M3 promises open weights with 1M-token context and frontier coding, but BenchLM ranks it #29 overall and #69 on multimodal. Teams need independent verification.
FlashMemory's learned index compresses DeepSeek-V4's KV cache to 13.5% of baseline at parity accuracy. The project is suspended; per-suite recall breakdowns are not published.
Cloudflare claims a 15x bot surge using a classifier that flags privacy browsers as bots. Audit your own logs before trusting the numbers behind Pay-Per-Crawl.
kapa.ai's data shows indexing image captions at ingestion adds 1-6% query overhead versus 27-51% for raw query-time vision, shifting recall risk to caption fidelity.

LLM decode is memory-bandwidth-bound, not capacity-bound. A 70B model on the DGX Spark's 273 GB/s hits roughly 2.7 tok/s. Count GB/s, not GB, when sizing inference hardware.
dstack-capsule binds pod identity into Intel TDX hardware quotes, enabling multi-pod confidential VMs without the per-VM density tax of Confidential Containers.
MusaCoder trains a 9B model to emit native GPU kernels for Moore Threads' MUSA architecture, claiming parity with frontier models on vendor-controlled benchmarks.
Microsoft's Azure Linux 4.0 extends the internal CBL-Mariner into a Fedora-based server OS for VMs. Preview gaps remain, and AKS teams should test now but wait for GA.
Cloudflare acquired VoidZero, putting Vite, Rolldown, and Oxc maintainers on a deploy-target vendor's payroll. MIT licensing stays. Roadmap neutrality is the open question.

A used V100 looks like cheap VRAM for local inference, but no bf16, no FlashAttention, and CUDA 13 deprecation lock buyers into a software stack that is actively contracting.
A 1.5B-parameter model hits 91.5% on Kubernetes YAML generation, but the remaining failures are syntactically valid manifests that deploy and quietly violate cluster intent.
TPU v6e Flex-start delivers 308M tokens per dollar for Gemma 4 31B prefill, undercutting H100 rates for open-weight serving, but production decode costs remain unquantified.
ObjectCache retrieves KV cache per-layer from S3, adding 5.6% TTFT at 64K context but 56-75 ms at 4K. Long-context deployments where DRAM is the bottleneck benefit most.