Vercel Makes WAF Mitigated Traffic Free: Recompute Your Edge Cost Model
Vercel now waives CDN and bandwidth charges for WAF-mitigated traffic, removing the penalty for aggressive blocking and shifting the bottleneck to rule tuning.
The archive · Page 3 of 6
The serving stack, network fabric, and cloud-account substrate beneath production AI, where every throughput claim collides with rebuild windows, egress invoices, and control-plane risk.
49–72 of 130 articles · Newest first
Vercel now waives CDN and bandwidth charges for WAF-mitigated traffic, removing the penalty for aggressive blocking and shifting the bottleneck to rule tuning.
GLM-5.2 self-hosting on 8 B200s costs $1.74 per million output tokens at batch 50 but $15.68 for one request. Bursty loads make a managed gateway cheaper than idle GPU node.
Cloudflare DMARC Management is now generally available and free for Cloudflare DNS customers, but p=reject still breaks on SPF lookup limits and authentication misalignment.
Claude Code's per-tool prompts are consent controls, not isolation. A July 2026 arXiv preprint and MCP's prompt-injection bugs show why agent runtimes need real isolation.
TF-Engram pages LLM memory to SSD, moving the bottleneck from HBM to storage bandwidth, latency, and prefetch accuracy. It wins only when reads arrive before decode stalls.
A July 2026 preprint shows Triton and TileLang kernels can pass correctness checks yet run hundreds of times slower than library baselines, moving validation cost to adopters.
Vercel Edge Config reads feature flags in under 15ms, but its 10-second global write window means flipped flags can still route traffic to disabled regions.
Cloud Run GPU's 19-second cold start for Gemma 3 4B means scale-to-zero beats dedicated GPUs only for spiky, batch workloads below roughly 40 to 50 percent utilization.
Vercel's in-function concurrency lets one instance run multiple Node.js or Python requests, cutting idle billing and cold starts but forcing handling of shared state, leaks.
HuggingFace's Optimum-AMD recipe and ROCm 7.2.4's vLLM fixes make the MI300X a serviceable inference target, but newer attention kernels and training still trail CUDA.
Cloudflare's Meerkat runs QuePaxa consensus at the edge, so every write waits on a cross-region quorum. The write-latency tax suits control-plane state, not transactions.
Since April 6, 2026, new Vercel projects honor Cache-Control from external origins by default, so operators must audit rewrite headers or risk stale, unintended responses.
zkSecurity's AI agent found seven real bugs in Cloudflare's CIRCL crypto library, all fixed upstream, showing AI-assisted review is now a baseline layer for crypto code.
Pruning RAG context is a ranking decision. Reranking and compression keep only answer-changing chunks, shifting cost from the prompt onto retrieval and shrinking the cite set.
Cloudflare's Monetization Gateway prices resources per request in stablecoins via HTTP 402, shifting fee friction to callers and splitting pricing into metered and flat plans.

Doubao 2.1 Pro ships at ¥6/¥30 per million tokens. The family's 180 trillion daily tokens reset what Western inference stacks must assume about price and capacity.

Every CUDA kernel pays a fixed driver-queue tax before its first FLOP runs. The fusion, graphs, and batching sold as bandwidth wins mostly hide the launch overhead.
A Vercel Montreal region only earns its cost when a legal rule forces data to stay in Canada. For every other workload, the real work is a residency audit, not a migration.

GLM-5.2 ships MIT-licensed with same-week serving recipes for NVIDIA vLLM and Huawei Ascend NPUs, breaking the open-weights-but-NVIDIA-only trade-off for self-hosters.
Vercel fronts its own Discourse forum with its CDN, but the edge cache only serves anonymous reads. Logged-in pages fall to the origin by design.
Vercel's Runtime Logs now show CDN cache key, tags, and revalidation reason per request, but the eviction cause stays hidden, forcing manual header inspection.
MKG-RAG-Bench isolates retrieval in multimodal knowledge graph RAG and finds it is the bottleneck. Adding images and graph edges costs more without guaranteed accuracy.
Vercel added redirect and rewrite telemetry to Observability, putting per-route errors beside function errors and shifting routing incidents toward on-call engineers.
Cloudflare Workflows saga rollbacks move compensation into declared per-step undo handlers, exposing the idempotency assumption every saga platform quietly offloads.