groundy
Models & Research

Can You Self-Host Mistral Large 4? What 1.05T Parameters Cost to Run

Mistral Large 4's 1.05T parameters make single-node self-hosting unrealistic at full precision, requiring multi-terabyte memory and leaving the API as the only current option.

Published 12 references
A bulky green resin dinosaur rests a yellow forepaw inside an open suitcase, while its broad body and curled tail remain outside against an ivory background.
On this page10 sections

Self-hosting Mistral Large 4 on a single node is not realistic at full or half precision, and the reason is a number most coverage will skip: 1.05 trillion total parameters. That is the figure Mistral’s own model docs published with the Public Preview on October 6, 2026. The same page lists 52B active parameters, a 1M-token context, and API pricing of $0.68 per million input tokens and $2.09 per million output tokens (model docs). Run the conventional byte-per-parameter arithmetic on that total and the weight set alone needs roughly 2.1 TB in BF16, 1.05 TB in FP8, and 525 GB at 4-bit precision, before KV cache. An 8×H100 node (640 GB HBM) holds only the 4-bit case, with thin headroom.

There is a second, more immediate constraint: as of October 7, 2026, the weights do not exist for download. Mistral’s launch post says “Weights drop end of this month,” after a red-teaming period with vetted partners. So the honest answer for the next few weeks is simple: the hosted API is the only path, and the self-hosting question is a planning exercise. It is still worth doing now, because the arithmetic says this model redraws the line where “open-weight” stops meaning “self-hostable” for mid-size platform teams.

Every ML4 number in this article is vendor-reported preview data. The launch post does relay results from independent evaluations, including the Artificial Analysis Cyber Index, vals.ai, and a Surge AI blind human study, but every figure arrives through Mistral’s own pages: no third party has published an ML4 result yet, so there is no measured serving throughput and no downloadable checkpoint to check against. Mistral’s pages also disagree with each other on the size (the docs say 52B active and 1.05T total; the launch post says “49 billion active parameters” and “1 trillion-parameter”). Where I do math, it is arithmetic on those vendor totals, clearly flagged as inference, not measurement.

What shipped on October 6

The model documentation describes Mistral Large 4 as “a state-of-the-art, open-weight, general-purpose multimodal model with a granular Mixture-of-Experts architecture,” with 52B active parameters, 1.05T total parameters, and a 1.6B vision encoder, in Public Preview v26.10. The launch post adds that the model “was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe,” and that “The public preview is served on that same infrastructure.” That last detail is quietly useful: the vendor’s reference serving environment is a Grace Blackwell fleet, not a single commodity node.

The docs also list API pricing with two columns: $0.68 and $1.36 per million input tokens, $2.09 and $4.18 per million output tokens (pricing table), and $0.07/$0.14 for cached input. What the doubled column means, standard versus batch or some other tier, is left unexplained by the docs, so any TCO math below uses the lower figures and flags the assumption.

Mistral’s launch post carries a long list of vendor-reported benchmarks: 82% on one Artificial Analysis Cyber Index test that asks a model to reproduce a real vulnerability in open-source software and then patch it (claimed as the highest of any model), and 93% on Cybench (launch post). The post also reports 61.7% on DeepSWE v1.1, 49.8% on a Coding Agent Index, and 59.9% on AutomationBench, and says that on Lakera’s B3 security benchmark ML4 resists 93.3% of attacks (launch post). None of these figures is published independently by the evaluators themselves; everything arrives through Mistral’s post, so treat the whole set as vendor-relayed claims until third-party results appear on their own channels.

Total parameters set the memory wall; active parameters set the compute bill

The sizing mistake the announcement invites is to read “52B active” and think of a 52B-class deployment. That number governs how much compute each token costs, not how much memory the model occupies. The MoE serving literature is blunt about this. The FineMoE paper (EuroSys ‘26) states: “Though certain model parameters remain inactive during inference, they must still reside in GPU memory to allow for potential future activation.” Every expert might be needed by the next token, so every expert’s weights must be reachable.

This is not a subtle effect. MoE-Lightning reports that Mixtral 8x22B’s expert FFN parameters alone require over 256 GB, 4–5× the memory of dense models needing similar inference FLOPs. And the expert layers are where the bytes live: in recent large-scale MoE models, “expert weights account for the dominant fraction of total parameters, often exceeding 90%” (OLED-MoE, Section 2.2, EuroSys ‘27). The paper says this of DeepSeek-style, Qwen-style, and LLaDA-style architectures generally, not of ML4; but if the pattern holds, roughly nine-tenths of the 1.05T total is expert weight, which matters because that is exactly the component quantization and offloading strategies target.

So the working rule for any MoE, and ML4 in particular: size memory from the total parameter count, size compute (and therefore per-token latency and throughput) from the active count. The Kimi K2 technical report is the closest published data point at this scale: K2 is a 32B-active, 1T-total MoE, and “storing the model parameters in BF16 and their gradient accumulation buffer in FP32 requires approximately 6 TB of GPU memory, distributed over a model-parallel group of 256 GPUs.” That figure includes training buffers, so it overstates pure inference, but it corroborates the order of magnitude: a trillion-parameter MoE is a multi-terabyte residency problem.

The memory floors, and what they exclude

Using the conventional 2 bytes per parameter at BF16, 1 byte at FP8, and 0.5 bytes at 4-bit weights, on Mistral’s stated 1.05T total:

PrecisionBytes/paramWeight-only floorFits 8×H100 (640 GB)?Fits 8×H200 (768 GB)?Fits 8×B200 (960 GB)?
BF162~2.1 TBNoNoNo
FP81~1.05 TBNoNoNo
4-bit (W4A16)0.5~525 GBYes, ~115 GB headroomYes, ~243 GB headroomYes, comfortably

The node capacities come from NanoFlow’s accelerator table as printed: H100 at 80 GB, B200 at 120 GB with 8,000 GB/s bandwidth, H200 at 96 GB, and AMD MI325X at 256 GB. One caution cuts across the table: NanoFlow’s figures are as-printed for 2023–2024-era parts, so re-verify against the vendor’s spec sheet before any purchase decision. The H200 entry (96 GB) may not match current HBM3e SKUs; a larger current H200 SKU would change the FP8 answer, and any configuration that only just clears the floor still leaves little room for KV cache. The B200 entry (120 GB) carries the same risk in the other direction: a larger current SKU would flip the FP8 cell above from No to a marginal Yes. The floors bind beyond accelerator nodes too: a machine with 512 GB of unified memory sits below the 4-bit weight floor of ~525 GB before any KV cache is counted.

“Weight-only floor” deserves emphasis, because everything excluded pushes the requirement up, not down. The floors above leave out KV cache (for a model advertising a 1M-token context, this is not a rounding error at long contexts), activations, the 1.6B vision encoder, and framework overhead. The 8×H100 4-bit row, with ~115 GB of headroom, is the thinnest margin in the table, and real deployments would want more for concurrency.

FP8 is worth treating as the realistic full-quality tier rather than a compromise. Hopper and Blackwell execute FP8 natively with higher matrix-multiply throughput than 16-bit formats, per MosaicQuant’s hardware survey, and FP8-Flow-MoE reports on a 671B-parameter MoE “up to 21% higher throughput and 16.5 GB lower memory usage per GPU compared to BF16 and naïve FP8 baselines.” That result is from a different model and a training-adjacent recipe, so read it as evidence that FP8 serving at MoE scale is practical, not as an ML4 prediction.

For the 4-bit tier, the accuracy risk ordering matters. QServe divides integer quantization into W8A8, W4A16, and W4A4, and reports that “the former two methods are considered nearly lossless in terms of accuracy” while W4A4 is not. The practical reading: 4-bit weights with 16-bit activations (W4A16) is the defensible floor; pushing activations to 4-bit to squeeze below the 525 GB weight floor computed above from Mistral’s stated totals is where accuracy risk becomes real, and nobody has measured this on a 1T-parameter model with a 1M context.

The offloading counter-case, honestly weighed

The strongest argument against “it doesn’t fit” is expert offloading: keep hot experts in GPU memory and stream cold ones from CPU RAM or disk. The literature here is genuinely encouraging. MoE-Lightning describes the standard approach of offloading weights and KV cache “to CPU memory or disk,” reaches GPU-memory-bound throughput ceilings with 2–3× less CPU memory than baselines, and OLED-MoE reports cutting time-per-output-token by 1.23×, 7.93× and improving expert-cache utilization 1.44×, 4.23× over prior offloading systems.

Two caveats decide whether this helps you. First, none of the published results is anywhere near ML4’s size, and one line of them is not even on autoregressive models. FineMoE evaluates Mixtral 8x7B, Qwen1.5-MoE, and Phi-3.5-MoE; MoE-Lightning evaluates Mixtral 8x7B, Mixtral 8x22B, and DBRX (132B); and OLED-MoE’s TPOT gains are measured on 16B and 100B LLaDA-style diffusion MoE models, a different decoding paradigm entirely. Nothing published demonstrates offloading a 1T-parameter autoregressive MoE, and the I/O volumes scale with the model. Second, offloading buys batch throughput, not latency. MoE-Lightning itself states there is no benefit in swapping weights to GPU for latency-oriented applications with one or two prompts, because such workloads stay bound by the CPU-to-GPU memory roof. If your workload is interactive chat or agentic loops with tight tail-latency budgets, offloading does not rescue a small-memory deployment. If your workload is overnight batch evaluation or document processing, it might, and FP8 or 4-bit weights with expert streaming over a large CPU memory pool becomes the plausible self-host architecture for a team without a B200 cluster.

I would not plan around offloading as the primary design, though. The evidence gap between 141B-class demonstrations and a 1.05T model is too wide, and ML4’s expert count and routing topology are unpublished, so you cannot yet model the cache behavior.

Self-host TCO versus the API

At the rates in Mistral’s pricing table of $0.68 per million input tokens and $2.09 per million output tokens (the lower of the two columns; the meaning of the doubled $1.36/$4.18 column in the same table is unresolved), the API converts a capital problem into an operating expense. A workload processing 100M input and 20M output tokens a month costs about $110 at Mistral’s listed rates ($68 for input plus about $42 for output). Even at ten times that volume, you are in the low four figures monthly, against a multi-GPU FP8 pool (an 8×B200 node at NanoFlow’s printed 120 GB per GPU does not even clear the FP8 weight floor) or a 4-bit 8×H100 deployment with the operational burden of an unproven model, an unpublished architecture, and no measured throughput.

The comparison that matters for most mid-size teams is not “API versus my cluster” but “API versus the open models I already serve.” The open-weight ladder documented in the Tele-FLM retrospective runs Mistral 141B, DeepSeek 236B, Grok 314B, Llama-3 beyond 400B. Models in the 200–400B class fit in BF16 or FP8 on one or two 8×H100 nodes; ML4 at 1.05T raises the memory floor several-fold over that class. If your serving plan was built around 200B-class open weights, ML4 is not an incremental addition to it. It is a different infrastructure commitment, closer to Kimi K2’s 256-GPU reference than to anything single-node.

The one precedent worth studying is Mistral’s own. Mistral Large 2, at 123B dense, was explicitly “designed for single-node inference,” and it shipped under the Mistral Research License, with commercial self-deployment requiring a separate Mistral Commercial License. Neither the model docs nor the launch post states ML4’s license terms. If the same structure applies, “open-weight” may mean you can download it, study it, and still owe Mistral a contract before a commercial deployment. The license may gate self-hosting as firmly as VRAM does, and nobody can price that until the weights land.

What to verify before committing

The decision window opens when weights drop at the end of October 2026, which Mistral says will arrive “along with more details on the architecture, additional benchmarks, and our post-training methodology.” Before sizing hardware or signing an API commitment, I would want five things resolved:

  1. License terms. Whether commercial self-deployment is permitted, and at what cost, following the Large 2 precedent.
  2. Expert count and routing top-k. These determine whether offloading and expert-caching systems can even be modeled for ML4, and how the expert-dominant share of total parameters reported by OLED-MoE maps onto its architecture.
  3. The second price column. Whether the $0.68/$2.09 column in Mistral’s pricing table is batch, standard, or something else changes every TCO comparison above.
  4. Independent serving throughput. Every ML4 performance number today arrives through Mistral’s own pages, with red-teaming still under way; wait for independently published measurements at FP8 and 4-bit before trusting the floors as deployment plans.
  5. The 52B versus 49B discrepancy. The docs and the launch post disagree; the difference does not change the memory-floor order of magnitude, but spec flux during preview is a reason to re-run this arithmetic on the final model card rather than the preview one.

The durable takeaway outlives this launch. For MoE models, “open-weight” and “self-hostable” are separate properties joined by a single number: total parameters. Active parameters tell you what a token costs to compute; total parameters tell you whether the model fits in your building. At 1.05T, with the weight floors computed from Mistral’s stated totals (roughly 2.1 TB at BF16, 1.05 TB at FP8, and 525 GB at 4-bit), Mistral Large 4 fits in very few buildings. Until the weights, license, and independently published benchmarks arrive, the API at Mistral’s listed pricing of $0.68/$2.09 per million tokens is the realistic path, and any self-host plan should be drafted around FP8 or 4-bit precision with the explicit understanding that it is arithmetic on vendor-stated totals, not measurement.

Frequently Asked Questions

When will Mistral Large 4 weights be available for download?

as of October 7, 2026, the weights do not exist for download. Mistral’s launch post says “Weights drop end of this month,” after a red-teaming period with vetted partners.

What is the minimum memory required to run Mistral Large 4 at 4-bit precision?

Run the conventional byte-per-parameter arithmetic on that total and the weight set alone needs roughly 2.1 TB in BF16, 1.05 TB in FP8, and 525 GB at 4-bit precision, before KV cache.

What is the API pricing for Mistral Large 4?

The docs also list API pricing with two columns: $0.68 and $1.36 per million input tokens, $2.09 and $4.18 per million output tokens (pricing table), and $0.07/$0.14 for cached input.

References

Follow the links in the article for context. The supporting material is collected here for further reading.

  1. Mistral's own model docsdocs.mistral.aiAccessed
  2. Mistral's launch postmistral.aiAccessed
  3. FineMoE paperarxiv.orgAccessed
  4. MoE-Lightningarxiv.orgAccessed
  5. OLED-MoEarxiv.orgAccessed
  6. Kimi K2 technical reportarxiv.orgAccessed
  7. NanoFlow's accelerator tablearxiv.orgAccessed
  8. MosaicQuant's hardware surveyarxiv.orgAccessed
  9. FP8-Flow-MoEarxiv.orgAccessed
  10. QServearxiv.orgAccessed
  11. Tele-FLM retrospectivearxiv.orgAccessed
  12. Mistral Large 2mistral.aiAccessed

Join the discussion

Share a useful perspective or ask a question about this article.

Discussion guidelinesComments privacy