DeepSeek 32B on RTX 3090: Tokens per Second by Quant and Context
Bandwidth math predicts 32-39 tok/s for Q4_K_M DeepSeek 32B on an RTX 3090 at 8k context. 32k context overflows 24GB VRAM. No measured benchmarks exist, so every figure is.
The archive · Page 2 of 6
Where architecture, training tricks, and eval methodology meet the marketing layer — separating durable progress in foundation models from leaderboard theater that quietly falls apart under load.
25–48 of 129 articles · Newest first
Bandwidth math predicts 32-39 tok/s for Q4_K_M DeepSeek 32B on an RTX 3090 at 8k context. 32k context overflows 24GB VRAM. No measured benchmarks exist, so every figure is.
PTXBench shows LLMs can generate architecture-specific PTX for H100 and B200, but no model matches frontier libraries. Use it to cut hot-loop porting costs, not to replace.
A new preprint proposes a lightweight reader to reuse KV caches across models, but evidence suggests caches remain model-bound. Learn when recompute beats adaptation.
DeCRIM shows that decomposing multi-constraint instructions into individually checkable units reduces silent drops by 7-8% on benchmarks, shifting reliability work from.
arXiv 2607.25494 introduces a single-pass tool to localize silent numerical instability in bf16 and fp16 training. Treat it as a triage layer for unexplained loss spikes, not.
arXiv:2607.25600 shows verbalized confidence is a weak but real routing signal for RAG. It saves 20.4% retrieval calls for 28.2% token overhead. The probe is poorly.
Running Kimi K3 on an M1 Max proves local MoE feasibility via expert offloading, but unified memory bandwidth caps throughput. Treat this as a prototyping tool, not a serving.
Kimi Linear cuts KV cache 75% and boosts decode 6x for 1M-token contexts, but recall stays the binding constraint. Route memory-bound jobs only after benchmarking in-context.
A July 22 visual test comparing GPT-5.6, Claude, Gemini, and Grok highlights a structural gap: chat models lack dedicated image backends. Teams must route by subtask using.
Moonshot AI's Kimi K3 release triggers governance review for regulated teams. Beijing jurisdiction, Anthropic accusations, and missing government assessment require internal.
A leaked transcript claims DeepSeek faces a compute gap, but V4 shipped in April. Teams should treat this as a resilience prompt to pin Qwen as a swap-in spine, not panic.
ARBIGRAPH shows tool agents lose 33.3% accuracy on dependent chains despite fitting context. Context ordering, truncation, and caching now beat raw window size for agent.
Parallel decoding has not cut serving costs for diffusion LLMs. arXiv:2605.13026 closes training gaps by 4x, but KV-cache absence keeps inference slower than autoregressive.
Fable 5 pricing and caching set the bar for open-weight routers. The Echo claim lacks verification. Routing wins only for high-volume, cache-unfriendly routine traffic.
DeepSeek-V4 markets a 1M token window for agents, but no independent benchmarks verify depth recall. CEO-Bench and ProGraph research show structured memory outperforms raw.
No primary source confirms Qwen-Image-3.0 exists. Independent benchmarks show open-weight models fail 85% of precise image tasks. Self-hosting branded templates remains risky.

Kimi K3 pairs 2.8T parameters with 16-of-896 expert routing and a 1M context window. Hosted API is live, but full weights arrive July 27, 2026. Self-hosting requires.
Kimi K3 offers concrete pricing for production bake-offs while Qwen3.8 Max remains a shadow evaluation. Compare both against Fable 5, GPT-5.6 Sol, and GLM-5.2 by workload.
Qwen3.8 Max now has a stable API, $2/$6 pricing, 1M context, and downloadable 2.4T weights. The release is real, but the API and open checkpoint are not the same product.

HuggingFace claims 100x inference speedup, but independent benchmarks show its TGI engine trails vLLM by 24x on throughput. This article decomposes the claim to separate.
Kimi K3 is frontier-class, but its 2.8T MXFP4 checkpoint needs 64-plus accelerators. The real deployment math reveals what an arena rank cannot.
TALRanker folds the tool-call decision into the reranker's scoring policy, turning tool latency from a fixed per-query tax into a budget the model spends only when uncertain.
Progressive Tree Drafting grows draft tokens as a pruned tree to claim a 2× speedup, but the gain hinges on verifier acceptance and shrinks on code and reasoning.
A new arXiv preprint replaces sampling-based averaging in Bayesian deep ensembles with closed-form Bayesian aggregation, cutting per-query inference cost and shifting the.