Short answer: yes, a 125B model now runs on a 12 GB gaming card, and fast. Strata is a free, MIT-licensed inference engine built for one model, Qwen3.8-Flash-Next. Its authors measured 79 tokens per second on an RTX 5070 with 12 GB of VRAM and 64 GB of RAM at the recommended size. The hardware that matters most is system RAM, not VRAM: 48 GB runs the standard sizes and 64 GB runs all of them. You also need about 80 GB of SSD space.
Quick Navigation:
What Strata Is
Strata is an open-source inference engine, written in C++ and built on parts of llama.cpp, that does one thing: run Alibaba’s Qwen3.8-Flash-Next on an ordinary PC. The repository was created on September 24, 2026, reached Hacker News on October 4 with more than 900 points, and had over 16,000 GitHub stars three days later.
The claim sounds like the kind of thing that usually falls apart on inspection. We covered one of those in our AirLLM guide, where a 70B model does run on 4 GB of VRAM but at a speed nobody would use. Strata is different because the speeds are usable, and the reason is the model’s architecture.
Qwen3.8-Flash-Next is a Mixture-of-Experts model with a very fine-grained expert layout. According to the Strata documentation it has 24,576 small experts and uses only 10 of them for each token. Strata splits the work three ways:
- The graphics card runs the core of the model for every token and holds the roughly 4,000 experts that are used most often.
- System RAM holds all 24,576 experts. When a token needs one the card does not have, the CPU computes it.
- The SSD holds a 29 GB lookup table. The model reads a few rows per token, so it never has to be loaded into memory.
On top of that, Strata uses a small draft layer to guess the next few tokens and has the big model verify them in one pass, which the project says makes generation 1.6 to 1.8 times faster with the same output.
One note on the parameter count: Strata and the Hacker News thread describe the model as 125 billion parameters, while the Hugging Face checkpoint metadata lists about 180 billion in total. We use Strata’s figure here because it is the one the project’s memory numbers are built around.
Hardware Requirements
These are the project’s own stated requirements.
| Component | Requirement |
|---|---|
| Graphics card | 12 GB of VRAM or more. NVIDIA RTX 20, 30, 40, or 50 series, or AMD RX 7900 XT / XTX, RX 7800 XT, RX 7700 XT, RX 9060 XT, RX 9070 / 9070 XT, Radeon AI PRO R9700, RX 6800 / 6900 series |
| System RAM | 32 GB minimum, 64 GB to run every size |
| Disk | About 80 GB free, SSD strongly recommended |
| Operating system | Windows 10 or 11, or Linux, with a current GPU driver |
RAM decides which size you can run
This is the part that surprises people who are used to shopping by VRAM. Because every expert lives in system RAM, RAM is the gate. The same model ships in several quantized sizes:
| Size | RAM + VRAM Needed | Fits With | Quality (per the authors) |
|---|---|---|---|
| Coder (IQ1_M) | 29.6 GB loaded | 32 GB RAM | Half the experts removed. Strong at code, weaker elsewhere |
| Q2_0 | 37.6 GB | 48 GB RAM | Good, fastest |
| IQ2_XS | 39.2 GB | 48 GB RAM | Better, the recommended default |
| IQ3_XXS | 47.0 GB | 64 GB RAM | Great |
| IQ3_S | 54.8 GB | 64 GB RAM, little else open | Best. Matches the full model on the published tests |
Key insight: for this workload a 12 GB card with 64 GB of RAM runs every size, while a 24 GB card with 32 GB of RAM is limited to the Coder, Q2_0, and IQ2_XS sizes through Strata’s low-RAM mode. A bigger card makes Strata faster, but outside that mode it does not lower the RAM needed.
If you are deciding what to upgrade, our RAM for local LLMs guide covers kits and capacities, and the RAM comparison page lists current options. The SSD matters too: the first load is much faster from NVMe, and the largest community quants read part of the model from disk while answering. See our NVMe guide.
Measured Speeds
All numbers below are from the Strata project’s own benchmarks, not ours.
| Size | RTX 5070 12 GB, short chat | RTX 5070 12 GB, 128K context | RX 9070 XT 16 GB, short chat |
|---|---|---|---|
| Q2_0 | 94 tokens/s | 76 tokens/s | 60 tokens/s |
| IQ2_XS | 79 tokens/s | 63 tokens/s | 52 tokens/s |
| IQ3_XXS | 62 tokens/s | 49 tokens/s | Not measured |
| IQ3_S | 53 tokens/s | 46 tokens/s | Not measured |
| Coder | 55 tokens/s | 43 tokens/s | 44 tokens/s |
The NVIDIA system was an RTX 5070 with a Ryzen 5 7600 and 64 GB of RAM. The AMD system was an RX 9070 XT with a Ryzen 9 3900X and 47 GB of RAM, running Linux. Prompt reading on a 32K-token prompt ran at 1,620 to 2,650 tokens per second on the RTX 5070.
Two things stand out. The CPU and RAM platform clearly matter: the AMD test used a much older Ryzen 9 3900X on DDR4-era hardware and was slower despite having more VRAM, though the different GPU backend is part of that too. And more VRAM helps: the project estimates an RTX 3090 with 24 GB at roughly 100 to 140 tokens per second, since more experts fit on the card. That is an estimate, not a measurement.
The Catches
Strata is impressive, and it is also a two-week-old project. Know these before you install it.
- It is a quantized model, at 2 to 3 bits. The authors say IQ3_S matches the full model on published tests, but the faster sizes give up some quality. Our quantization guide explains what low-bit quants cost.
- Loading locks up your PC. The documentation warns that the machine can be slow or unresponsive for 1 to 3 minutes while 35 to 55 GB is loaded into RAM.
- One request at a time by default. Parallel requests are possible but slow each answer on a 12 GB card.
- The first long prompt is slow. Strata reads the first message of a chat at about 1 minute per 30,000 tokens. Follow-ups start in seconds.
- The 4-bit options are much slower. The near-full-quality UD-Q4_K_XL quant reads most of the model from the SSD and writes only 7 to 8.5 tokens per second on a 64 GB PC.
- The model has its own license. Strata is MIT. Qwen3.8-Flash-Next ships under the Qwen community license, so read it before commercial use.
- It runs one model. This is not a general replacement for Ollama or llama.cpp. If you want a model that fits entirely in VRAM with no tricks, see our Qwen3.8-27B hardware guide.
Pro tip: if you have 32 GB of RAM today and a 12 GB or 16 GB card, a 64 GB RAM kit is the upgrade that unlocks every Strata size. A new graphics card will not.
What This Means for Buying Hardware
For two years the advice for local LLMs has been simple: buy as much VRAM as you can afford. Strata is the clearest sign yet that fine-grained MoE models change that. When only a tiny share of the weights is used per token, the hot part fits on a mid-range card and the rest can sit in cheap system memory.
This does not make a 24 GB or 32 GB card pointless. Dense models, image and video generation, and fine-tuning still want VRAM, and with RTX 5090 street prices where they are, that money goes a long way elsewhere. It does mean a 12 GB or 16 GB card paired with 64 GB of RAM and a fast NVMe drive is now a serious local LLM machine, at a fraction of the price of the big-VRAM route. The trillion-parameter models are a different story: our Mistral Large 4 guide shows where this approach runs out.
FAQ
What hardware do you need to run Strata?
An NVIDIA RTX 20, 30, 40, or 50 series card or a supported AMD Radeon card with at least 12 GB of VRAM, 32 GB or more of system RAM, about 80 GB of free SSD space, and Windows 10 or 11 or Linux. 64 GB of RAM runs every model size.
How fast is Qwen3.8-Flash-Next on a 12GB GPU with Strata?
The Strata project measured 53 to 94 tokens per second on an RTX 5070 with 12 GB of VRAM, a Ryzen 5 7600, and 64 GB of RAM, depending on the quantization. The recommended IQ2_XS size ran at 79 tokens per second in a short chat and 63 tokens per second at 128K context.
Does Strata need more VRAM or more RAM?
More RAM. All of the model’s experts are held in system RAM, so RAM decides which size fits: 32 GB for the Coder version, 48 GB for Q2_0 and IQ2_XS, and 64 GB for every size. A bigger graphics card makes it faster but does not lower the RAM requirement, except in Strata’s low-RAM mode.
Is the quality the same as the full model?
Not at every size. Strata runs 2-bit and 3-bit quantizations. Its authors say the largest, IQ3_S, matches the full model on the published tests, while the smaller Q2_0 and IQ2_XS trade some quality for speed. The 32 GB Coder version drops half the experts and is weaker outside code.
Does Strata work on AMD graphics cards?
Yes. The project lists the Radeon RX 7900 XT and XTX, RX 7800 XT, RX 7700 XT, RX 9060 XT, RX 9070 and 9070 XT, Radeon AI PRO R9700, and the RX 6800 and 6900 series. It measured 52 tokens per second with IQ2_XS on an RX 9070 XT with 16 GB.
Related Reading
RAM for Local LLMs
The 64 GB upgrade this engine depends on
AirLLM: 70B on a 4GB GPU
The earlier big-model-on-a-small-card trick, and why it is slow
Qwen3.8-27B Hardware Guide
The dense Qwen model that fits on one GPU
LLM Quantization Guide
What 2-bit and 3-bit quants really cost
Mistral Large 4 Hardware Guide
A trillion-parameter model that no gaming PC can hold
Best NVMe SSD for AI
Fast storage for 70 GB model downloads and SSD-backed weights
Speccing a Local LLM Machine?
