Skip to main content
Strata: Run Qwen3.8-Flash-Next (125B) on a 12GB GPU

Image: Strata architecture diagram (Niko1221/Strata on GitHub), MIT

Strata: Run Qwen3.8-Flash-Next (125B) on a 12GB GPU


Short answer: yes, a 125B model now runs on a 12 GB gaming card, and fast. Strata is a free, MIT-licensed inference engine built for one model, Qwen3.8-Flash-Next. Its authors measured 79 tokens per second on an RTX 5070 with 12 GB of VRAM and 64 GB of RAM at the recommended size. The hardware that matters most is system RAM, not VRAM: 48 GB runs the standard sizes and 64 GB runs all of them. You also need about 80 GB of SSD space.

What Strata Is

Strata is an open-source inference engine, written in C++ and built on parts of llama.cpp, that does one thing: run Alibaba’s Qwen3.8-Flash-Next on an ordinary PC. The repository was created on September 24, 2026, reached Hacker News on October 4 with more than 900 points, and had over 16,000 GitHub stars three days later.

The claim sounds like the kind of thing that usually falls apart on inspection. We covered one of those in our AirLLM guide, where a 70B model does run on 4 GB of VRAM but at a speed nobody would use. Strata is different because the speeds are usable, and the reason is the model’s architecture.

Qwen3.8-Flash-Next is a Mixture-of-Experts model with a very fine-grained expert layout. According to the Strata documentation it has 24,576 small experts and uses only 10 of them for each token. Strata splits the work three ways:

  • The graphics card runs the core of the model for every token and holds the roughly 4,000 experts that are used most often.
  • System RAM holds all 24,576 experts. When a token needs one the card does not have, the CPU computes it.
  • The SSD holds a 29 GB lookup table. The model reads a few rows per token, so it never has to be loaded into memory.

On top of that, Strata uses a small draft layer to guess the next few tokens and has the big model verify them in one pass, which the project says makes generation 1.6 to 1.8 times faster with the same output.

One note on the parameter count: Strata and the Hacker News thread describe the model as 125 billion parameters, while the Hugging Face checkpoint metadata lists about 180 billion in total. We use Strata’s figure here because it is the one the project’s memory numbers are built around.

Hardware Requirements

These are the project’s own stated requirements.

ComponentRequirement
Graphics card12 GB of VRAM or more. NVIDIA RTX 20, 30, 40, or 50 series, or AMD RX 7900 XT / XTX, RX 7800 XT, RX 7700 XT, RX 9060 XT, RX 9070 / 9070 XT, Radeon AI PRO R9700, RX 6800 / 6900 series
System RAM32 GB minimum, 64 GB to run every size
DiskAbout 80 GB free, SSD strongly recommended
Operating systemWindows 10 or 11, or Linux, with a current GPU driver

RAM decides which size you can run

This is the part that surprises people who are used to shopping by VRAM. Because every expert lives in system RAM, RAM is the gate. The same model ships in several quantized sizes:

SizeRAM + VRAM NeededFits WithQuality (per the authors)
Coder (IQ1_M)29.6 GB loaded32 GB RAMHalf the experts removed. Strong at code, weaker elsewhere
Q2_037.6 GB48 GB RAMGood, fastest
IQ2_XS39.2 GB48 GB RAMBetter, the recommended default
IQ3_XXS47.0 GB64 GB RAMGreat
IQ3_S54.8 GB64 GB RAM, little else openBest. Matches the full model on the published tests

Key insight: for this workload a 12 GB card with 64 GB of RAM runs every size, while a 24 GB card with 32 GB of RAM is limited to the Coder, Q2_0, and IQ2_XS sizes through Strata’s low-RAM mode. A bigger card makes Strata faster, but outside that mode it does not lower the RAM needed.

If you are deciding what to upgrade, our RAM for local LLMs guide covers kits and capacities, and the RAM comparison page lists current options. The SSD matters too: the first load is much faster from NVMe, and the largest community quants read part of the model from disk while answering. See our NVMe guide.

Measured Speeds

All numbers below are from the Strata project’s own benchmarks, not ours.

SizeRTX 5070 12 GB, short chatRTX 5070 12 GB, 128K contextRX 9070 XT 16 GB, short chat
Q2_094 tokens/s76 tokens/s60 tokens/s
IQ2_XS79 tokens/s63 tokens/s52 tokens/s
IQ3_XXS62 tokens/s49 tokens/sNot measured
IQ3_S53 tokens/s46 tokens/sNot measured
Coder55 tokens/s43 tokens/s44 tokens/s

The NVIDIA system was an RTX 5070 with a Ryzen 5 7600 and 64 GB of RAM. The AMD system was an RX 9070 XT with a Ryzen 9 3900X and 47 GB of RAM, running Linux. Prompt reading on a 32K-token prompt ran at 1,620 to 2,650 tokens per second on the RTX 5070.

Two things stand out. The CPU and RAM platform clearly matter: the AMD test used a much older Ryzen 9 3900X on DDR4-era hardware and was slower despite having more VRAM, though the different GPU backend is part of that too. And more VRAM helps: the project estimates an RTX 3090 with 24 GB at roughly 100 to 140 tokens per second, since more experts fit on the card. That is an estimate, not a measurement.

The Catches

Strata is impressive, and it is also a two-week-old project. Know these before you install it.

  • It is a quantized model, at 2 to 3 bits. The authors say IQ3_S matches the full model on published tests, but the faster sizes give up some quality. Our quantization guide explains what low-bit quants cost.
  • Loading locks up your PC. The documentation warns that the machine can be slow or unresponsive for 1 to 3 minutes while 35 to 55 GB is loaded into RAM.
  • One request at a time by default. Parallel requests are possible but slow each answer on a 12 GB card.
  • The first long prompt is slow. Strata reads the first message of a chat at about 1 minute per 30,000 tokens. Follow-ups start in seconds.
  • The 4-bit options are much slower. The near-full-quality UD-Q4_K_XL quant reads most of the model from the SSD and writes only 7 to 8.5 tokens per second on a 64 GB PC.
  • The model has its own license. Strata is MIT. Qwen3.8-Flash-Next ships under the Qwen community license, so read it before commercial use.
  • It runs one model. This is not a general replacement for Ollama or llama.cpp. If you want a model that fits entirely in VRAM with no tricks, see our Qwen3.8-27B hardware guide.

Pro tip: if you have 32 GB of RAM today and a 12 GB or 16 GB card, a 64 GB RAM kit is the upgrade that unlocks every Strata size. A new graphics card will not.

What This Means for Buying Hardware

For two years the advice for local LLMs has been simple: buy as much VRAM as you can afford. Strata is the clearest sign yet that fine-grained MoE models change that. When only a tiny share of the weights is used per token, the hot part fits on a mid-range card and the rest can sit in cheap system memory.

This does not make a 24 GB or 32 GB card pointless. Dense models, image and video generation, and fine-tuning still want VRAM, and with RTX 5090 street prices where they are, that money goes a long way elsewhere. It does mean a 12 GB or 16 GB card paired with 64 GB of RAM and a fast NVMe drive is now a serious local LLM machine, at a fraction of the price of the big-VRAM route. The trillion-parameter models are a different story: our Mistral Large 4 guide shows where this approach runs out.

FAQ

What hardware do you need to run Strata?

An NVIDIA RTX 20, 30, 40, or 50 series card or a supported AMD Radeon card with at least 12 GB of VRAM, 32 GB or more of system RAM, about 80 GB of free SSD space, and Windows 10 or 11 or Linux. 64 GB of RAM runs every model size.

How fast is Qwen3.8-Flash-Next on a 12GB GPU with Strata?

The Strata project measured 53 to 94 tokens per second on an RTX 5070 with 12 GB of VRAM, a Ryzen 5 7600, and 64 GB of RAM, depending on the quantization. The recommended IQ2_XS size ran at 79 tokens per second in a short chat and 63 tokens per second at 128K context.

Does Strata need more VRAM or more RAM?

More RAM. All of the model’s experts are held in system RAM, so RAM decides which size fits: 32 GB for the Coder version, 48 GB for Q2_0 and IQ2_XS, and 64 GB for every size. A bigger graphics card makes it faster but does not lower the RAM requirement, except in Strata’s low-RAM mode.

Is the quality the same as the full model?

Not at every size. Strata runs 2-bit and 3-bit quantizations. Its authors say the largest, IQ3_S, matches the full model on the published tests, while the smaller Q2_0 and IQ2_XS trade some quality for speed. The 32 GB Coder version drops half the experts and is weaker outside code.

Does Strata work on AMD graphics cards?

Yes. The project lists the Radeon RX 7900 XT and XTX, RX 7800 XT, RX 7700 XT, RX 9060 XT, RX 9070 and 9070 XT, Radeon AI PRO R9700, and the RX 6800 and 6900 series. It measured 52 tokens per second with IQ2_XS on an RX 9070 XT with 16 GB.

Speccing a Local LLM Machine?