About prices:prices on this page are US street prices in USD, last checked October 2026. They are for reference only. Local prices differ by region and usually include VAT or sales tax, and availability changes quickly, so check the retailer before you buy.
Short answer: not yet, and not on consumer hardware when you can. Mistral Large 4 launched on October 6, 2026 as an API-only preview, with open weights promised by the end of the month. It has 1.05 trillion parameters, about 41% more than GLM-5.2, so plan for roughly 345 GB of combined memory at 2-bit and 525 GB or more at 4-bit. That rules out every single GPU and the 256 GB machines that run GLM-5.2. A 512 GB Mac Studio, a 384 GB multi-GPU workstation, or a rented 8x H200 node are the realistic targets.
Quick Navigation:
What Mistral Actually Announced
Mistral Large 4, nicknamed “Le Chonk” by Mistral itself, is the company’s largest model so far. The launch thread passed 1,800 points on Hacker News within a day, mostly on one question: who can actually run this?
Here is what is confirmed, from Mistral’s announcement and the model card.
Mistral Large 4
Two details are worth flagging. First, the two official sources disagree slightly: the announcement says 1 trillion parameters with 49B active, while the model card says 1.05T total with 52B active. This guide uses the model card’s 1.05T for memory math, since the larger figure is the safer one to plan around. Second, Mistral has not yet named the license. “Open-weight” tells you the weights will be downloadable, not what you are allowed to do with them, so check the terms before building a commercial deployment on it.
Mistral trained the model on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters and serves the preview from the same infrastructure. API pricing is $1.36 per million input tokens and $4.18 per million output tokens.
As with every MoE model, the active parameter count governs speed and the total count governs memory. Only about 50B parameters fire per token, but all 1.05T have to be loaded, because different tokens route to different experts. Our LLM quantization guide explains how parameter count turns into gigabytes.
Estimated Memory by Quantization
No quantized builds of Mistral Large 4 exist yet, because there are no weights to quantize. The numbers below are estimates, scaled from the published sizes of Unsloth’s dynamic GGUF quants for GLM-5.2 (744B parameters) by the parameter ratio of 1.41. Real files will differ by a few percent depending on how the layers are mixed. “Total memory” means VRAM plus system RAM combined.
| Quantization | GLM-5.2 (published) | Mistral Large 4 (estimate) | Smallest Hardware That Fits |
|---|---|---|---|
| Dynamic 1-bit | 223 GB | About 315 GB | 384 GB RAM workstation |
| Dynamic 2-bit | 245 GB | About 345 GB | 512 GB Mac Studio or 4x 96 GB GPUs |
| 3-bit | 290-360 GB | 410-510 GB | 512 GB RAM server or 4x H200 |
| 4-bit | 372-475 GB | 525-670 GB | 8x H200 node |
| 8-bit | 810 GB | About 1,140 GB | Larger than one 8x H200 node |
Key insight: the 256 GB tier is gone. GLM-5.2 squeezed into 256 GB at 2-bit, which made a Mac Studio or a single GPU with 256 GB of RAM viable. A 1.05T model is estimated above 300 GB even at 1-bit, so the entry point moves to 384 GB and realistically 512 GB.
Remember that these figures are for the weights only. A 1M-token context window is expensive in its own right, so leave headroom for the KV cache and expect to run far shorter contexts locally than the API offers.
Hardware That Can Hold It
Path 1: Mac Studio M5 Ultra with 512 GB
The only desktop likely to load a 2-bit build with room left for context. The 512 GB configuration of the M5 Ultra Mac Studio ships from late October 2026, which lines up with Mistral’s weight release. Its 1.2 TB/s of memory bandwidth is the best you can get outside a GPU, but this model has about 25% more active parameters per token than GLM-5.2, so expect low single-digit tokens per second. See our unified memory PC guide for how these machines compare.
Path 2: One GPU + 384 to 512 GB of System RAM
llama.cpp’s MoE offloading keeps attention layers on a single 24 GB or 32 GB GPU and streams expert weights from RAM. That worked for GLM-5.2 on high-end desktop boards. At this size it needs an eight-channel workstation or server platform (Threadripper PRO, Xeon W, or EPYC) with 8x 48 GB or 8x 64 GB DIMMs, and DDR5 is expensive right now. Speed is bound by RAM bandwidth. Our RAM for local LLMs guide covers the offloading setup, and the motherboard guide lists boards that take this much memory.
Path 3: Multi-GPU Workstation (384 GB VRAM)
Four RTX PRO 6000 Blackwell cards at 96 GB each total 384 GB. That is the rig that ran GLM-5.2 at mixed precision with a long context. For Mistral Large 4 the same hardware is estimated to fit only the 2-bit build, with about 40 GB to spare for context. It will be the fastest local option by a wide margin, at a cost of around $50,000. See our multi-GPU setup guide for how memory is split across cards.
Path 4: Rent an 8x H200 Node
Eight H200s give you 1,128 GB of VRAM, enough for a 4-bit build with real context headroom. A 4x H200 node (564 GB) should handle 3-bit and the low end of 4-bit. This is the only practical way to run the model at near-full quality without owning a server rack. Vast.ai usually has the lowest hourly prices on multi-GPU configurations, and RunPod gives new users a $5 credit to test with. Our cloud GPU comparison has current pricing.
What does not make the list: any single consumer GPU, and the 128 GB unified memory boxes such as the DGX Spark or Ryzen AI Max+ mini PCs. They are far too small.
What to Do Before the Weights Land
The weights are a few weeks out, so the useful work right now costs almost nothing.
Test it on the API first. At $4.18 per million output tokens, a few dollars buys enough real usage to tell whether Mistral Large 4 beats what you already run. If GLM-5.2 or DeepSeek V4 Flash does your job as well, you can skip a hardware problem that is 41% bigger.
Do not buy hardware for it yet. Every memory number in this guide is an estimate. Real quant sizes, the license, and the first llama.cpp and vLLM benchmarks will all arrive within days of the weight release. A 512 GB Mac Studio or a 512 GB RAM kit is a large purchase to make on a projection.
If you already own a 256 GB machine, plan on something smaller. Mistral says this model “will also serve as the foundation for a new generation of specialized and optimized Mistral models.” A distilled or pruned variant is the likeliest way this family reaches a 256 GB Mac or a single-GPU rig.
Pro tip: when you do run it locally, quantize the KV cache. With the weights taking 345 GB or more, running llama.cpp with –cache-type-k q4_1 –cache-type-v q4_1 is what decides whether you get a usable context window or none at all.
We will update this guide with measured file sizes and speeds once the weights and the first quantized builds are public.
FAQ
Can you run Mistral Large 4 locally right now?
No. Mistral launched it as an API-only public preview on October 6, 2026 and says the open weights will be released by the end of October. Until then the only way to use it is the Mistral API, priced at $1.36 per million input tokens and $4.18 per million output tokens.
How much VRAM will Mistral Large 4 need?
Mistral’s model card lists 1.05 trillion total parameters, which is about 2.1 TB at FP16. No quantized builds exist yet. Scaling from the published GGUF sizes of the 744B GLM-5.2, expect roughly 345 GB of combined memory at dynamic 2-bit and 525 to 670 GB at 4-bit. These are estimates until real quants are published.
Can Mistral Large 4 run on an RTX 5090 or RTX 4090?
Not on the GPU alone. A 32 GB or 24 GB card holds only a small fraction of a 1.05T-parameter model. It can serve as the GPU half of a MoE offloading setup, but that also needs roughly 384 GB or more of system RAM, which means a workstation or server platform.
Will a 256 GB Mac Studio run Mistral Large 4?
Probably not. The 256 GB tier that fits GLM-5.2 at 2-bit is too small here: even an aggressive 1-bit quant of a 1.05T model is estimated at around 315 GB. The 512 GB M5 Ultra Mac Studio is the only desktop likely to hold a 2-bit build.
What is the cheapest way to try Mistral Large 4?
The API. At $4.18 per million output tokens you can run a lot of real work for a few dollars, which is the sensible way to find out whether the model is worth building or renting hardware for once the weights are public.
Related Reading
GLM-5.2 Hardware Guide
The 744B model these estimates are scaled from
DeepSeek V4 Flash Hardware Guide
A 300B-class MoE model with a much lower memory floor
Unified Memory PCs for Local LLMs
Mac Studio M5 Ultra, DGX Spark, and Ryzen AI Max compared
LLM Quantization Guide
GGUF, GPTQ, AWQ and what the bit levels really cost
Multi-GPU Ollama Setup
How memory splits across several cards
Cloud GPU Providers
Rent H200 nodes by the hour for trillion-parameter models
Building for Big Local Models?
Need a workstation that can hold 512 GB of RAM? Check our Tailored Builds page.
