Skip to main content
Unified Memory PCs for Local LLMs: DGX Spark, Mac, Ryzen AI

Image: NVIDIA DGX Spark by Daniel Lu (User:dllu), CC BY-SA 4.0, via Wikimedia Commons

ยท Last updated on

Unified Memory PCs for Local LLMs: DGX Spark, Mac, Ryzen AI


About prices:prices on this page are US street prices in USD, last checked October 2026. They are for reference only. Local prices differ by region and usually include VAT or sales tax, and availability changes quickly, so check the retailer before you buy.

Short answer: a unified memory PC lets you load LLMs far bigger than any consumer GPU can hold, but it runs them slower. A 128GB box like the NVIDIA DGX Spark (about $6,950 after an October 2026 price increase) or an AMD Ryzen AI Max+ 395 mini PC (about $2,500 to $3,000) can run 70B models at 8-bit or 100B-class MoE models, which a 32GB RTX 5090 cannot. Their memory bandwidth of 256 to 273 GB/s is about a seventh of the 5090โ€™s, though, so generation speed is lower. The Mac Studio M5 Ultra (up to 512GB at 1.2 TB/s) is the fastest and most expensive option. If your model fits in VRAM, a GPU still wins.

What Unified Memory Means for LLMs

On a normal PC, the GPU has its own VRAM and the CPU has system RAM. A model has to fit in VRAM to run fast, and anything that spills into system RAM crosses the PCIe bus, which is roughly 30 times slower than GDDR7 (see our RAM for local LLMs guide for what offloading costs).

A unified memory system puts the CPU and GPU on one package sharing one large memory pool. The GPU can address most of it directly, so a 128GB machine behaves like a GPU with roughly 100GB or more of usable VRAM. That is why these boxes became popular for local LLMs: with RTX 5090s selling for around $5,000 in late 2026 (see our GPU buying guide), a 128GB unified memory PC is often the cheapest way to get past 32GB.

The Contenders

SystemMax MemoryBandwidthSoftwarePrice (2026)
NVIDIA DGX Spark128 GB273 GB/sFull CUDAabout $6,950 (128GB), from $4,999 (64GB)
NVIDIA RTX Spark PCs128 GBNot announcedCUDA, WindowsFall 2026, not announced
AMD Ryzen AI Max+ 395 mini PCs128 GBabout 256 GB/sROCm, Vulkanabout $2,500 to $3,000
Mac Studio M5 Max128 GB460 GB/sMLX, Metalfrom $2,499 (36GB)
Mac Studio M5 Ultra512 GB1.2 TB/sMLX, Metalfrom $5,499 (96GB)
For reference: RTX 509032 GB1,792 GB/sFull CUDAaround $5,000 street

NVIDIA DGX Spark

The DGX Spark is a 150 by 150 mm desktop box built on the GB10 Grace Blackwell superchip, with 128GB of LPDDR5X unified memory at 273 GB/s. Its main advantage is that it runs the full NVIDIA software stack: CUDA, PyTorch, TensorRT, and anything written for an NVIDIA GPU works without porting. That makes it the only box here that is also reasonable for small fine-tuning jobs, not just inference.

Price has gone the wrong way. It launched at $3,999, and NVIDIA raised the Founders Edition to $4,699 in February 2026, citing memory supply constraints. In September 2026 the NVIDIA store was out of stock and the cheapest Amazon US unit was about $5,000 (Notebookcheck, Toolhalla).

Update, October 7, 2026: prices moved again within days of this guide going live. NVIDIA announced a 64GB DGX Spark starting at $4,999, available from Acer, ASUS, Dell, Gigabyte, HP, and MSI on October 23, and the 128GB model rose to about $6,950 depending on the manufacturer (NVIDIA, ServeTheHome). NVIDIA says the 64GB version supports models up to 100 billion parameters, and that two units linked over the built-in ConnectX-7 port pool to 128GB. At these prices the Ryzen AI Max+ 395 boxes below cost less than half as much for the same 128GB, so the DGX Spark only makes sense if you need CUDA.

NVIDIA RTX Spark

Announced on May 31, 2026, RTX Spark brings the same idea to Windows laptops and compact desktops. It pairs a Blackwell RTX GPU with 6,144 CUDA cores to a 20-core Grace CPU over NVLink-C2C, with up to 128GB of unified memory and 1 petaflop of FP4 AI compute. NVIDIA says it can run 120B-parameter LLMs locally. Systems from ASUS, Dell, HP, Lenovo, Microsoft Surface, and MSI are due in fall 2026, with Acer and GIGABYTE to follow (NVIDIA).

NVIDIA had not announced memory bandwidth or pricing at launch, and those two numbers will decide whether it beats the DGX Spark for LLM work. Wait for real reviews before buying.

AMD Ryzen AI Max+ 395 Mini PCs

AMDโ€™s Ryzen AI Max+ 395 (often called Strix Halo) is the budget route to 128GB. It pairs 16 Zen 5 cores with a large integrated Radeon GPU and LPDDR5X-8000 memory at around 256 GB/s, almost the same bandwidth as the DGX Spark for less than half the price. 128GB models listed in 2026 include the Corsair AI Workstation 300 at $2,499, Framework Desktop at $2,851, and Beelink GTR9 Pro and GMKtec EVO-X2 at about $3,000 (Liliputing, TechPowerUp).

On Windows, AMDโ€™s Variable Graphics Memory lets you dedicate up to 96GB to the GPU. Linux can share memory with the GPU more flexibly. The catch is software: you run models through llama.cppโ€™s Vulkan or ROCm backends, Ollama, or LM Studio, and CUDA-only research code wonโ€™t run. ROCm 10 has improved this, but NVIDIA is still smoother. AMD has also announced a Ryzen AI Halo developer platform on the newer Ryzen AI Max PRO 400 series for later in 2026 (AMD).

Apple Mac Studio M5 Max and M5 Ultra

Apple announced the M5 Mac Studio on August 25, 2026 (Macworld). The M5 Max version offers up to 128GB at 460 GB/s, starting at $2,499. The M5 Ultra offers 96GB, 256GB, or 512GB at 1.2 TB/s, starting at $5,499. A 256GB, 16TB configuration costs $18,299, and the 512GB option ships from late October.

On bandwidth the M5 Ultra is in a different league from the 128GB boxes: 1.2 TB/s is over four times the DGX Spark and about two-thirds of an RTX 5090, with sixteen times the 5090โ€™s memory. It is the only desktop here that can hold a 300B-class model in memory. The trade-off is the software ecosystem: you use Appleโ€™s MLX or llama.cppโ€™s Metal backend, not CUDA.

The Other Extreme: AMD Threadripper Halo Station

At IFA on September 4, 2026, AMD showed the Threadripper Halo Station, a liquid-cooled workstation with a 96-core Threadripper PRO 9995WX, up to 2TB of DDR5, and two Instinct MI350P accelerators with 144GB of HBM3E each, expandable to four for 576GB of GPU memory. AMD pitches it at running trillion-parameter models locally (ServeTheHome). It isnโ€™t unified memory, itโ€™s datacenter accelerators in a tower, and it is due in 2027 with no announced price. It is worth knowing about, but it isnโ€™t a home purchase.

Why Bandwidth Decides Speed

When an LLM generates text, it reads every active weight from memory once per token. That gives a simple upper bound:

Max tokens per second โ‰ˆ memory bandwidth รท size of the active weights

Real-world speed lands below this ceiling because of overhead, but the ratio between machines holds. For a dense 70B model at 4-bit (about 40GB of weights):

SystemBandwidthTheoretical ceiling, 70B at 4-bit
Ryzen AI Max+ 395256 GB/sabout 6 tokens/s
DGX Spark273 GB/sabout 7 tokens/s
Mac Studio M5 Max460 GB/sabout 11 tokens/s
Mac Studio M5 Ultra1.2 TB/sabout 30 tokens/s
2x RTX 3090 (48GB total)936 GB/s eachabout 20 tokens/s (layers split across cards)

Two things change this picture:

  • MoE models are much faster. A mixture-of-experts model only reads its active experts per token. A 100B-class MoE with around 10B active parameters reads only a few GB per token, so even a 256 GB/s box can generate at comfortable reading speed. Most of the big 2026 open models are MoE, which is why these machines work better in practice than the dense-model math suggests.
  • Prompt processing is compute-bound. Reading a long prompt or document (prefill) depends on raw compute, not bandwidth. Here the NVIDIA systems, with Blackwell Tensor Cores, are well ahead of the Ryzen and Apple options. If you paste long documents or run agents with big contexts, that gap matters.

What Fits in 128GB and 512GB

Plan for about 75 to 85% of total memory being usable for the model and its context. The OS and apps need the rest.

  • Any 128GB system: Qwen3.8-27B at full BF16 (54.7GB), 70B dense models at 8-bit, and DeepSeek V4 Flash at its smallest usable quant (about 82.5GB), which otherwise needs four 24GB GPUs.
  • Mac Studio M5 Ultra 512GB: GLM-5.2, which needs 223GB or more even at its smallest quant. That is too tight for the 256GB model once the OS and context take their share. The 256GB version comfortably runs larger quants of everything in the 128GB tier.
  • A single RTX 5090 (32GB): 27B to 32B models at 4 to 8-bit, much faster than any unified memory box, but nothing in the 70B+ class without offloading or a multi-GPU setup.

See the LLM quantization guide for how quantization levels trade size against quality.

Which Should You Buy?

Cheapest Route to 128GB

Pick: Ryzen AI Max+ 395 mini PC with 128GB

About $2,500 to $3,000 for nearly the same bandwidth as a DGX Spark. Best if you run models through Ollama, LM Studio, or llama.cpp and donโ€™t need CUDA.

You Need CUDA

Pick: DGX Spark, or wait for RTX Spark reviews

If you fine-tune, prototype for NVIDIA servers, or rely on CUDA-only libraries, the DGX Spark is the only box here that just works. RTX Spark may be the better Windows option once pricing and bandwidth are public.

Biggest Models, Best Speed

Pick: Mac Studio M5 Ultra, 256GB or 512GB

The only desktop that holds 300B-class models with bandwidth fast enough to use them comfortably. Expensive, and you work in MLX or llama.cpp rather than CUDA.

Your Models Fit in 24 to 48GB

Pick: A GPU, or two used RTX 3090s

For models that fit in VRAM, discrete GPUs are several times faster and also train well. Two used 3090s give you 48GB for roughly the price of one Ryzen AI Max box. See the best GPU for deep learning guide.

Try a Model Before You Buy the Box

Not sure whether a 70B model is worth $3,000 of hardware? Rent a 48GB or 80GB GPU for an afternoon on RunPod (new users get a $5 credit through our link) or Vast.ai, and test the exact model and quant you plan to run.

Referral links: signing up supports TensorRigs at no extra cost to you.

FAQ

Is a unified memory PC better than a GPU for local LLMs?

It is better at fitting big models and worse at running them fast. A 128GB unified memory box can load a 70B model at 8-bit or a 100B-class MoE model that no single consumer GPU can hold, but its memory bandwidth (256 to 460 GB/s on most systems) is a fraction of an RTX 5090โ€™s 1,792 GB/s, so tokens per second are lower. If your model fits in VRAM, a GPU is faster.

How fast is the DGX Spark for local LLMs?

Its 128GB of unified memory runs at 273 GB/s, which caps a dense 70B model at 4-bit (about 40GB) at roughly 7 tokens per second in theory, and less in practice. MoE models that only activate a small slice of their weights per token run much faster. Its strengths are capacity and full CUDA support, not raw generation speed.

How much does a DGX Spark cost in 2026?

The 128GB DGX Spark launched at $3,999 and NVIDIA raised it to $4,699 in February 2026, citing memory supply constraints. In early October 2026 the 128GB model moved to about $6,950 depending on the manufacturer, and NVIDIA announced a 64GB configuration starting at $4,999, available from partners on October 23, 2026.

Can a Mac Studio run large LLMs locally?

Yes. The M5 Ultra Mac Studio offers up to 512GB of unified memory at 1.2 TB/s, the most capacity and bandwidth of any desktop in this class. It starts at $5,499 with 96GB, and the 512GB configuration ships from late October 2026. The trade-off is no CUDA, so you use MLX or llama.cppโ€™s Metal backend instead of the NVIDIA ecosystem.

What is the cheapest way to get 128GB for local LLMs?

An AMD Ryzen AI Max+ 395 mini PC. 128GB models from Corsair, Framework, Beelink, and GMKtec sold for about $2,500 to $3,000 in 2026, with around 256 GB/s of memory bandwidth. On Windows, AMDโ€™s Variable Graphics Memory lets you assign up to 96GB of it to the GPU.

When does NVIDIA RTX Spark come out?

NVIDIA announced RTX Spark on May 31, 2026, pairing a Blackwell RTX GPU with a 20-core Grace CPU and up to 128GB of unified memory. Laptops and compact desktops from ASUS, Dell, HP, Lenovo, Microsoft Surface, and MSI are due in fall 2026. NVIDIA had not announced pricing or memory bandwidth at launch.