groundy
Models & Research

Running DeepSeek R1 Locally: Hardware, Quantization, and Real Throughput

Running DeepSeek R1 locally on real hardware: per-quant VRAM budgets, computed decode ceilings for a 24 GB card, measured quality loss, and sourced throughput figures.

Published Updated 15 references
A small green resin dinosaur steps out of a scuffed carrying enclosure beside a much larger dinosaur. Yellow dorsal scales and hard shadows contrast with the warm ivory background.
On this page12 sections

Running DeepSeek R1 locally is possible on a single consumer GPU, but only if you are realistic about which variant you target and how you count memory. The full 671B model needs hundreds of gigabytes: its FP8 weights come to roughly 670 GB, and even 4-bit quantization leaves about 370 GB before the KV cache, per China Unicom’s deployment analysis. The distilled checkpoints fit on one card, and the quantization you pick decides both how much context you can afford and, because decode is memory-bandwidth-bound, how fast tokens appear. The gap between “technically runs” and “usably fast” is where most guides fail practitioners, and it is where this one spends its time.

The Model Landscape: 671B vs. Distilled Variants

DeepSeek R1 is not a single model. It is a family spanning a 1.5B distilled checkpoint up to the full 671B Mixture-of-Experts (MoE) model trained with large-scale reinforcement learning. (DeepSeek AI. “DeepSeek-R1.” GitHub, January 2025) The economics behind it are unusual: how these models were trained for a fraction of typical frontier compute costs explains why a 671B open model exists at all.

The distilled models (1.5B, 7B, 8B, 14B, 32B, 70B) were produced by fine-tuning Qwen2.5 and Llama 3 series checkpoints on 800K reasoning samples generated by the full R1 model. They are dense transformers, not MoE, which makes them easier to quantize and deploy on single GPUs. The 671B full model activates 37B parameters per token but requires storing all 671B, with a 128K context window. (DeepSeek AI. “DeepSeek-R1.” Hugging Face) That distinction shapes every hardware decision below.

DeepSeek released R1-0528 in May 2025 as an update to the full model (model card). Architecture and local hardware requirements are unchanged, but the model roughly doubles its average chain-of-thought per problem (12K to 23K tokens on AIME), lifting AIME 2024 Pass@1 from 79.8% to 91.4% and AIME 2025 from 70.0% to 87.5%. The update also ships a new distilled variant, R1-0528-Qwen3-8B, trained from Qwen3 8B Base on traces from the updated full model. At 86.0% Pass@1 on AIME 2024 it beats the original R1-Distill-Qwen-7B’s 55.5% by 30 points and edges Qwen3-235B-A22B’s 85.7%, per the 0528 model card and the January R1 evaluation. Quantization guidance below applies to both generations.

R1 is no longer DeepSeek’s current API model. DeepSeek shipped DeepSeek-V4 as an open-weight public preview on April 24, 2026, with a 1M-token default context and DeepSeek Sparse Attention, in two sizes: V4-Pro (1.6T total, 49B active) and V4-Flash (284B total, 13B active). The legacy deepseek-chat and deepseek-reasoner names were retired after July 24, 2026; the API’s current names are deepseek-flash and deepseek-v4-pro, and requests to the old V4-Flash name are now served by a V4.1-Flash refresh. (DeepSeek API Docs, April 2026; DeepSeek API Docs) None of that obsoletes a local R1 deployment: the open weights stay downloadable, and even V4-Flash’s 284B parameters sit far beyond any consumer memory budget before quantization. The local DeepSeek decision in late 2026 is still R1-0528 and its distills, and this guide covers them. DeepSeek-V4 FlashMemory covers the new architecture itself.

Hardware Tiers: What Actually Runs What

Single-Card Consumer GPUs

Capacity arithmetic, not vendor marketing, sets the tiers. Weight footprints below are computed from parameter count times bits per weight, with the bit-widths taken from China Unicom’s quantization study (Q4_K_M averages 4.82 bits per weight, Q3_K_M 3.81):

  • 7B and 8B distills: about 4–5 GB at Q4_K_M. Comfortable on 8–12 GB cards.
  • 14B distill: about 8–9 GB at Q4_K_M. Fits 12 GB-class cards with room for context.
  • 32B distill: about 19–20 GB at Q4_K_M. The largest distill a 24 GB card holds, and only with a short context budget. At Q8_0 it needs roughly 33–35 GB, which does not fit.
  • 70B distill: about 42 GB at 4.82 bits. Any 24 GB configuration is an offloading configuration.

For workstation budgets, the RTX PRO 6000 Blackwell packs 96 GB of GDDR7 at 1,792 GB/s on one card, per NVIDIA’s specifications. That holds the 70B distill at Q8_0 (about 70–75 GB computed) with context headroom. It does not hold any 671B build: the smallest variant in China Unicom’s comparison is Unsloth’s dynamic 2-bit UD-Q2_K_XL at 212 GB (arXiv, Table 1).

What a 32B Distill Needs on a 24 GB Card

This is the most common planning question for 24 GB cards, and the answer splits into two parts: capacity arithmetic, which is settled, and a speed model you apply to your own hardware. The weight figures come from parameter count times the measured bit-widths in China Unicom’s study. The ceiling column normalizes to 1,000 GB/s of memory bandwidth so you can scale it to whatever card you own by looking up its rated bandwidth.

QuantizationBits/weightWeights (computed)Fits 24 GB?Free after weights (computed)Ceiling per 1,000 GB/s (computed)
Q8_0~8.5 nominal~33–35 GBNon/an/a
Q4_K_M4.82~19.5–20 GBYes, tight~4 GB for KV cache and runtime~50 t/s
Q3_K_M3.81~15.5–16 GBYes~8 GB~64 t/s

Every figure in the right three columns is arithmetic, not measurement: weights are parameter count times bits per weight, free memory subtracts them from 24 GB, and the ceiling divides bandwidth by bytes read per token. Measure your own build before promising anyone a number.

Context is the hidden budget line. A dense model’s FP16 KV cache costs 2 × layers × KV heads × head dimension × 2 bytes per token, so the per-token price depends on the base architecture’s constants. The distill model card names Qwen2.5-32B as the base but does not publish those constants; they live in the base model’s configuration file, and until you read them there, any per-token figure is a calculated estimate rather than a verified number (distill card). What the weight arithmetic does settle: at Q4_K_M a 24 GB card has roughly 4 GB left once the ~19.5–20 GB of weights load, and Q3_K_M’s smaller file leaves about twice that margin. That is why context length on this card is a deliberate budget decision, not a max-it-by-reflex setting.

For speed, use the roofline rather than a borrowed number. Decode on a dense model is bandwidth-bound: every generated token reads the full weight file once, so your ceiling is memory bandwidth divided by bytes per token. At Q4_K_M the 32B distill reads about 19.5–20 GB per token, so a card with 1,000 GB/s of memory bandwidth ceilings near 50 tokens/second; divide your card’s rated bandwidth by its quant’s bytes-per-token figure for the equivalent number. Sustained decode lands below the ceiling because KV-cache reads and kernel overhead are not free. The one published measurement that calibrates the discount: NVIDIA’s DGX Spark offers 273 GB/s and ran Llama 3.1 70B in FP8 (about 70 GB of weights) at 2.7 t/s single-stream, roughly 70% of its 3.9 t/s arithmetic ceiling, per LMSYS’s October 2025 review. Two relative effects hold without any measurement: Q3_K_M reads 21% fewer bytes per token than Q4_K_M (3.81 vs 4.82 bits) for a proportionally higher ceiling, and any configuration that spills layers to system RAM falls off a cliff.

Prompt processing is the less constraining half. Prefill is compute-bound and runs far above decode on the same hardware: LMSYS measured 2,074 t/s of prefill against 83.5 t/s of decode for the 14B distill on the Spark, and the KTransformers hybrid reaches 286 t/s of prefill on CPU-side experts. Decode is the number that decides whether interactive use feels fluid.

If the Q4_K_M ceiling sits below your threshold, the alternative on the same card is R1-0528-Qwen3-8B at Q4_K_M: about 5 GB of weights, a much higher bandwidth ceiling, and a state-of-the-art AIME 2024 score among open-source models (86.0%) at its May 2025 release (model card).

Apple Silicon

Apple’s unified memory architecture eliminates the GPU/CPU memory split that complicates NVIDIA workflows, which matters most for large quantized models. (Runtime choice matters too: see MLX vs llama.cpp on Apple Silicon for how the two compare on M-series throughput.)

The published ceiling is Dave Lee’s March 2025 test: an M3 Ultra Mac Studio with the maximum 512 GB of unified memory loaded a 4-bit 671B build occupying 404 GB, with 448 GB manually allocated as video memory, and delivered 17–18 tokens per second under 200 W. That configuration starts around $10,000, and 512 GB is the requirement, not a luxury. (MacRumors, March 17, 2025) Apple has since announced a new Mac Studio with M5 Max and M5 Ultra chips, and an unverified first M5 Ultra benchmark surfaced in mid-September 2026, per MacRumors; the 512 GB M3 Ultra remains the Apple configuration with a published 671B result.

At the small end, a 16 GB M2 Pro MacBook Pro runs the 7B distill through Ollama at 53.1 t/s for a single request, degrading to 9.1 t/s at 19 concurrent requests, per a developer benchmark published February 2025. The concurrency cliff, not single-stream speed, is what disqualifies a laptop as a shared server.

Hybrid, CPU-Only, and Datacenter Paths

Running the full 671B at usable speed takes either an 8-GPU node or a hybrid design that exploits the MoE structure. On the GPU-node path, China Unicom’s numbers are the planning baseline: Q4_K_M brings the 671B to 377 GB of weights and 71 GB per GPU across eight GPUs at 32K context (arXiv, Table 1). SGLang has run exactly this class of deployment, with prefill-decode disaggregation and expert parallelism on 96-GPU H100 clusters and GB200 NVL72 systems, per LMSYS; the R1 model card also publishes vLLM and SGLang serve commands for the distills (Hugging Face).

The middle path is KTransformers, the CPU/GPU hybrid runtime from the KVCache.AI group. It keeps attention layers and the KV cache on the GPU, streams experts from system RAM, and runs expert matmuls on Intel AMX kernels. Its documented best case pairs a 24 GB RTX 4090-class card with Xeon Gold 6454S sockets for the Q4_K_M 671B (382 GB of DRAM on a single socket, 1 TB on the dual-socket test machine): up to 286.55 tokens/second prefill (v0.3-preview, six-expert selection) and 13.69 tokens/second decode, against 10.31 prefill and 4.51 decode for llama.cpp on the same CPUs. (KTransformers tutorial) The cost moves to DRAM and an Intel server CPU (AMX is Intel-only, per the project), and generation remains constrained, in the project’s own analysis, by CPU compute speed and memory bandwidth rather than anything the GPU can fix. The same hybrid trick now serves other large MoE models, including GLM-5.2’s home-deployment path.

For CPU-only inference, llama.cpp’s own discussion forum puts expectations in single digits. A single 48-core EPYC 7K62 with 512 GB of RAM reached 4.2 t/s on the 461.81 GB Q5_K_S 671B build; a dual-socket 96-core machine was slower at 2.9 t/s because llama.cpp does not scale well across sockets. The practical guidance from the thread: one CPU with maximum memory channels. (GitHub llama.cpp Discussions) Usable for overnight batch jobs, not for interactive use.

Quantization: What the Measurements Say

Quantization reduces weight precision from 16-bit floats to lower bit-widths, trading accuracy for memory and bandwidth. The GGUF format, used by llama.cpp and Ollama, is the standard for local LLM deployment; llama.cpp supports 1.5-bit through 8-bit integer quantization plus CPU+GPU hybrid inference for models larger than VRAM. (ggml-org/llama.cpp)

For the 32B distill, China Unicom’s study supplies measured effects rather than rules of thumb (arXiv, Table 4):

FormatBits/weight32B weights (computed)Measured effect vs BF16
Q8_0~8.5 nominal~33–35 GB0.29% average drop
Q4_K_M4.82~19.5–20 GB0% average drop; MATH-500 93.90 vs 93.65; AIME 70.40 vs 69.59
Q3_K_M3.81~15.5–16 GB0.68% average drop; LiveCodeBench 55.20 vs 57.08

Q4_K_M is the practical default: zero measured average drop on the 32B distill, and a 0.68% weighted-average drop for the full 671B against the FP8 API. Q8_0 is near-lossless (0.29%) but, at the ~8.5 bits per weight it averages, reads about 76% more bytes per token than Q4_K_M’s 4.82, which on bandwidth-bound decode is speed you pay for on every token. Q3_K_M holds up on average but slips on harder code generation, the LiveCodeBench row above; that is the real cost of the extra context headroom, not a generic “quality loss.” Independent work points the same direction: Red Hat’s compression of the full distill suite found FP8 and INT8 near-lossless across its reasoning benchmarks, with INT4 recovering 97%+ accuracy for the 7B and larger models (Red Hat Developer).

For the 671B, the same study compares five builds, including Unsloth’s dynamic 2-bit UD-Q2_K_XL (arXiv, Tables 1 and 2):

BuildSizeBits/weightMemory per GPU (8 GPUs, 32K ctx)R1 weighted-average drop
Q4_K_M (llama.cpp)377 GB4.8271 GB0.68%
Q3_K_M (llama.cpp)298 GB3.8161 GB1.80%
DQ3_K_M (China Unicom)281 GB3.5959 GB0.34%
Q2_K_L (llama.cpp)228 GB2.9152 GBnot reported for R1
UD-Q2_K_XL (Unsloth)212 GB2.7050 GB0.94%

The pattern worth noticing: dynamic bit allocation beats static quantization at similar sizes. UD-Q2_K_XL at 212 GB drops 0.94% while static Q3_K_M at 298 GB drops 1.80%; China Unicom’s own DQ3_K_M lands within 0.34% of FP8 at 281 GB. Every one of these still needs a multi-GPU node or hundreds of gigabytes of RAM.

Reported Throughput Across Configurations

Published generation (decode) figures, each attributed:

HardwareModelQuantDecode t/sSource
M2 Pro 16 GBR1-Distill-7BOllama default tag (quant not stated)53.1 (9.1 at 19 concurrent)dev.to, Feb 2025
M3 Ultra 512 GBR1 671B4-bit17–18MacRumors, Mar 2025
DGX Spark 128 GBR1-Distill-14BFP8, batch 883.5 (2,074 prefill)LMSYS, Oct 2025
DGX Spark 128 GBLlama 3.1 70B (reference)FP82.7 (803 prefill)LMSYS, Oct 2025
RTX 4090-class 24 GB + Xeon Gold 6454S, 382 GB–1 TB DRAMR1/V3 671BQ4_K_M hybrid13.69 (286.55 prefill)KTransformers docs, 2025
Single EPYC 48-core, 512 GBR1 671BQ5_K_S4.2llama.cpp discussion, 2025
Dual EPYC 96-core, 1 TBR1 671BQ5_K_S2.9llama.cpp discussion, 2025

The Spark’s two rows bracket the whole bandwidth argument: the same 273 GB/s machine serves a 14B at 83.5 t/s when batched but a 70B at 2.7 t/s single-stream, because bytes per token set the ceiling. On usability, the MacBook benchmark’s author sets a personal floor around 20 tokens/second and an ideal near 100 (dev.to); there is no industry-standard threshold, so calibrate against your own tolerance.

Practical Deployment: Ollama vs. llama.cpp

Ollama provides the lowest-friction path:

Terminal window
# Install Ollama, then:
ollama run deepseek-r1:14b
# For the 32B variant:
ollama run deepseek-r1:32b

Ollama handles model download and server management, and public Ollama benchmarks exercise the distills at both q4_K_M and q8_0, including LMSYS’s Spark testing of the 14B (LMSYS). The tradeoff is that it does not expose all llama.cpp tuning parameters, and GPU utilization may sit below a hand-tuned invocation.

llama.cpp offers control over context length, batch size, and threading. Its current quick start reaches Hugging Face directly, so a GGUF can be pulled at invocation time rather than downloaded separately (ggml-org/llama.cpp). For the 32B on a 24 GB card with a local file:

Terminal window
llama cli \
-m DeepSeek-R1-Distill-Qwen-32B-Q4_K_M.gguf \
-ngl 99 \
-c 8192 \
-t 8 \
-p "Your prompt here"

The -ngl 99 flag offloads all layers to GPU; setting it lower begins CPU offloading, which tanks throughput. LM Studio wraps llama.cpp in a GUI for desktop use. For a middle ground between cloud APIs and full local deployment, WebAssembly AI: Running Models covers in-browser inference.

Common Pitfalls

Buying capacity instead of bandwidth. Decode on a dense distill is roughly memory bandwidth divided by model size in bytes, so a card with more VRAM but slower memory is not automatically faster. NVIDIA’s DGX Spark (GB10 Grace-Blackwell, 128 GB unified LPDDR5X at 273 GB/s) makes the point with numbers: it ran GPT-OSS 20B at 49.7 t/s decode where an RTX PRO 6000 Blackwell reached 215 t/s and an RTX 5090 reached 205 t/s, roughly four times faster on the same model, per LMSYS. The RTX Spark unified-memory tradeoff works the numbers in full. Count GB/s before GB.

Ignoring the reasoning-token tax. R1-0528 averages about 23K chain-of-thought tokens per hard AIME problem, nearly double the original R1’s 12K, per the model card. At 20 tokens/second that is roughly 19 minutes of generation before the answer arrives, most of it hidden thinking. A distill that benchmarks slightly lower but emits fewer reasoning tokens can feel faster, and it shrinks your KV cache.

Undersizing the KV cache. Weights are not the whole memory budget. A dense model’s KV cache grows with every token of context and must sit in memory alongside the weights, which is what tips a fits-in-24-GB plan into silent CPU offload. The per-token cost is base-model arithmetic: 2 × layers × KV heads × head dimension × 2 bytes at FP16. The distill card names Qwen2.5-32B as the base but does not publish those constants, so read them from the base model’s configuration before turning any per-token estimate into a memory plan (distill card). Set context deliberately rather than maxing it by reflex.

Mismatched chat templates and settings. DeepSeek’s usage notes recommend temperature 0.5–0.7 (0.6 suggested), no system prompt, and enforcing a <think> prefix, because the original R1 can otherwise skip its reasoning pattern (DeepSeek-R1, Hugging Face). Ollama and LM Studio bundle the correct template; hand-rolled invocations frequently do not, producing “the model got dumber after I switched runtimes” reports that are really template bugs. R1-0528 relaxes both rules: system prompts are supported and the think prefix is no longer required (R1-0528, Hugging Face).

What the Distilled Models Actually Score

DeepSeek’s January 2025 evaluation remains the reference for the original distills, with the May 2025 card supplying the Qwen3-8B row (Hugging Face; 0528 card):

ModelAIME 2024 Pass@1MATH-500GPQA-DiamondLiveCodeBenchCodeForces rating
R1-Distill-Qwen-7B55.5%92.8%49.1%37.6%1189
R1-0528-Qwen3-8B (May 2025)86.0%n/a61.1%60.5%n/a
R1-Distill-Qwen-14B69.7%93.9%59.1%53.1%1481
R1-Distill-Qwen-32B72.6%94.3%62.1%57.2%1691
R1-Distill-Llama-70B70.0%94.5%65.2%57.5%1633
OpenAI o1-mini (reference)63.6%90.0%60.0%53.8%1820

The 32B distill beats o1-mini on four of the five published columns: AIME 2024 (72.6 vs 63.6), MATH-500 (94.3 vs 90.0), GPQA-Diamond (62.1 vs 60.0) and LiveCodeBench (57.2 vs 53.8). Competitive programming is the exception; o1-mini’s 1820 CodeForces rating beats the distill’s 1691. DeepSeek’s own summary is worded accordingly, that the 32B “outperforms OpenAI-o1-mini across various benchmarks,” not all of them (model card). The R1-0528-Qwen3-8B row changed the calculus for budget hardware: 86.0% on AIME 2024 from an 8B model, ahead of Qwen3-235B-A22B’s 85.7% on that benchmark, per the May 2025 model card. Its LiveCodeBench figure uses the later 2408–2505 problem window, so it is not directly comparable to the original distills’ column.

The 671B Full Model: Who Actually Needs It?

The full model beats every distill on hard reasoning: 91.4% on AIME 2024 for R1-0528 per the 0528 card, against 72.6% for the 32B distill per the January R1 evaluation. But the hardware gap is enormous. Unless you have:

  • A 512 GB M3 Ultra Mac Studio, the Apple configuration with a published 671B result, or
  • An 8-GPU 80 GB node (71 GB per GPU at Q4_K_M, 32K context, per China Unicom), or
  • A hybrid KTransformers box with 382 GB+ of DRAM, or a CPU server with 500 GB of RAM and tolerance for the 2.8–4.2 tokens/second the llama.cpp thread reports

…the distilled models are the rational choice. The Llama-70B distill scores 94.5% on MATH-500 and the 32B scores 94.3% (DeepSeek-R1), and both run on hardware an individual can buy.

References

Follow the links in the article for context. The supporting material is collected here for further reading.

  1. DeepSeek AI. "DeepSeek-R1." Hugging Facehuggingface.coAccessed
  2. DeepSeek AI. "DeepSeek-R1-Distill-Qwen-32B." Hugging Facehuggingface.coAccessed
  3. DeepSeek AI. "DeepSeek-R1-0528." Hugging Face, May 2025huggingface.coAccessed
  4. ggml-org. "llama.cpp: LLM inference in C/C++." GitHubgithub.comAccessed
  5. NVIDIA. "RTX PRO 6000 Blackwell Workstation Edition."nvidia.comAccessed
  6. DeepSeek API Docs. "DeepSeek-V4 Preview Release." April 2026api-docs.deepseek.comAccessed
  7. DeepSeek API Docs. "Your First API Call."api-docs.deepseek.comAccessed

Join the discussion

Share a useful perspective or ask a question about this article.

Discussion guidelinesComments privacy