DGX Spark 64GB vs Strix Halo 128GB: NVIDIA's $4,999 Spark Has Half the Memory — Should You Buy It?
On October 2 NVIDIA launched a 64GB DGX Spark at $4,999 and moved the 128GB Spark to $6,950. Every 128GB Strix Halo box in our catalog costs less than the new entry Spark. Here's the price-per-GB math, what actually fits in 64GB, and who should still pay for CUDA.
DataHardware Team
Our Top Pick

GMKtec EVO-X2 (Ryzen AI Max+ 395, 128GB)
$3,649Quick answer: buy the 64GB DGX Spark only if CUDA is non-negotiable and your models fit in roughly 30B-class territory — otherwise a 128GB Strix Halo mini PC is the better buy. As of October 2026, NVIDIA's cheapest DGX Spark costs $4,999 for 64GB of unified memory, and the 128GB model is now $6,950. A 128GB AMD Strix Halo box costs less than either: the GMKtec EVO-X2 is $3,649 in stock, and the Framework Desktop lists at $3,449 (currently out of stock). Memory bandwidth is nearly identical — 273 GB/s on GB10 vs 256 GB/s theoretical on the Ryzen AI Max+ 395 — so you're paying $1,350 more for half the memory and buying CUDA plus faster prompt processing, not faster token generation.
That's the whole verdict. The rest of this page is the math behind it, with every price dated and sourced, so you can check whether your situation is the exception.
What NVIDIA changed on October 2
On 2 October 2026 NVIDIA announced a second DGX Spark configuration in its blog post introducing the 64GB model and Sync Cluster Assistant. Same GB10 Grace Blackwell superchip, same 20-core Arm CPU, half the memory. The same day, the 128GB model's price moved to $6,950. The Register ran it under the headline "Nvidia debuts $4,999 DGX Spark with half the RAM and storage, amid memory crunch", and reported that skyrocketing memory prices are to blame — the same DRAM shortage that has repriced every box on this site.
The key facts, each traced to a source:
- 64GB DGX Spark: $4,999, on sale October 23 through Acer, ASUS, Dell, Gigabyte, HP and MSI (NVIDIA). It's partner-only — Hardware Busters notes there is no NVIDIA-branded 64GB unit.
- Storage is halved too, per The Register. NVIDIA hasn't published a single capacity for every OEM build, so check the listing you're buying from.
- Bandwidth is unchanged at 273 GB/s (The Register), which means NVIDIA is using lower-capacity LPDDR5X packages rather than fewer of them.
- 128GB DGX Spark: $6,950 — The Register calls it an increase of nearly 75 percent from this time last year.
| Date | DGX Spark 128GB | DGX Spark 64GB | Source |
|---|---|---|---|
| Launch (Oct 2025) | $3,999 | — | Hardware Busters |
| 27 Feb 2026 | $4,699 | — | NVIDIA price-change notice (memory supply) |
| 2 Oct 2026 | $6,950 | $4,999 (ships Oct 23) | The Register, Hardware Busters, NVIDIA |
Two hikes in eight months, then a cheaper SKU that's cheaper only because it has half the memory. That framing matters: the $4,999 box isn't a price cut. It's the old Spark's price bracket with the capacity that made the Spark interesting removed. If you followed our original DGX Spark vs Strix Halo comparison at the $4,699 price, the gap has roughly doubled since then.
Price per GB of unified memory: the table nobody else is printing
For local LLMs, unified memory capacity decides which models you can load. It's also soldered on every box here, so you can't add more later. Dollars per GB is the fairest way to compare a 64GB box with a 128GB one. Strix Halo prices come from our catalog (re-checked 4 October 2026, each from the vendor's own store unless noted); Spark prices come from the October 2 sources above. The math is simply price ÷ GB.
| Box | Memory | Bandwidth | Price | $/GB | Stack |
|---|---|---|---|---|---|
| Framework Desktop | 128GB | 256 GB/s | $3,449 (out of stock) | $26.95 | ROCm / Vulkan |
| GMKtec EVO-X2 | 128GB | 256 GB/s | $3,649 | $28.51 | ROCm / Vulkan |
| Minisforum MS-S1 Max | 128GB | 256 GB/s | $3,799 | $29.68 | ROCm / Vulkan |
| Beelink GTR9 Pro | 128GB | 256 GB/s | $4,349 (pre-order) | $33.98 | ROCm / Vulkan |
| HP Z2 Mini G1a | 128GB | 256 GB/s | $5,349 – $7,406 | $41.79 (at floor) | ROCm / Vulkan |
| ASUS Ascent GX10 | 128GB | 273 GB/s | $5,999 – $7,999 | $46.87 (at floor) | CUDA |
| NVIDIA DGX Spark 128GB | 128GB | 273 GB/s | $6,950 | $54.30 | CUDA |
| NVIDIA DGX Spark 64GB (OEM) | 64GB | 273 GB/s | $4,999 | $78.11 | CUDA |
Read that bottom row again. Per GB, the new 64GB Spark costs about 2.7× the EVO-X2 ($78.11 vs $28.51) and about 2.9× the Framework Desktop. In absolute dollars, it's $1,350 more than an EVO-X2, $1,200 more than an MS-S1 Max and $650 more than a GTR9 Pro — and every one of those has twice the memory.
One caveat on the GX10 row. Our catalog floor (the 1TB PCIe 4.0 variant, sold by Amazon) was checked on 29 September, before NVIDIA's change. GB10 partner pricing tends to follow NVIDIA's, so treat that row as possibly stale until we re-check it. The DGX Spark vs ASUS Ascent GX10 page tracks the two side by side.
The value pick at 128GB is the GMKtec EVO-X2. At $3,649 for 128GB and 2TB on GMKtec's own store, it's the cheapest 128GB box you can actually order today. The Framework Desktop is cheaper on paper, but it's out of stock. Up to 96GB of the 128GB is GPU-allocatable, which is still 32GB more model memory than the entire 64GB Spark. Our EVO-X2 vs Framework Desktop comparison covers the differences between those two.
What actually fits in 64GB vs 128GB
This is where the 64GB Spark's problem stops being abstract. Below are 4-bit weight footprints using the same math as our VRAM calculator (0.5 bytes per parameter, plus ~20% for runtime and a short context). Long-context KV-cache figures come from our KV-cache guide, which cites SitePoint's numbers.
| Model | Q4 footprint (short context) | DGX Spark 64GB | 128GB Strix Halo (~96GB GPU) | DGX Spark 128GB (~120GB) |
|---|---|---|---|---|
| Gemma 3 27B | ~17GB | Fits, room for long context | Fits | Fits |
| Qwen3 32B | ~20GB | Fits, room for long context | Fits | Fits |
| Llama 3.3 70B | ~42GB | Fits at short context only | Fits; 128K needs quantized KV | Fits, with long-context headroom |
| DeepSeek R1 Distill 70B | ~42GB | Tight — reasoning traces fill context fast | Fits | Fits |
| GPT-OSS 120B | ~63–72GB weights alone | Does not fit | Fits | Fits |
The 70B row is the one that matters. A dense 70B at 4-bit is about 40GB of weights. SitePoint puts the fp16 KV cache for a Llama 70B at 128K context at roughly another 40GB. On 64GB, with the OS also drawing from the same pool, that combination doesn't fit. You can run a 70B on the 64GB Spark, but only with a short context or a quantized KV cache. That rules out the long-document RAG and whole-repo coding agents that people usually buy a 70B-capable box for.
NVIDIA's blog says a single 64GB unit "supports up to 100-billion-parameter models." That's a vendor ceiling, not a recommendation. NVIDIA doesn't state the precision. By our calculator's math, a dense 100B at 4-bit is already ~50GB of weights before any context, so assume that figure means aggressive quantization and short prompts.
GPT-OSS 120B is the cleanest illustration. It's the model most 128GB owners actually run (our GPT-OSS 120B on a mini PC guide explains why), and on a 64GB box it simply doesn't load. It's also the model where the Spark's prompt-processing advantage is best documented. Buy the 64GB Spark and you lose access to the workload that made the Spark's premium easiest to justify.
For the general version of this decision, see how much unified memory you need for local LLMs. In short: if your largest model is 32B or smaller, 64GB is plenty, and on Strix Halo it costs far less than $4,999. Framework lists a 64GB Max+ 395 configuration too, though it is also currently out of stock.
Bandwidth is a wash, so speed on the same model is close
Token generation (decode) is bound by memory bandwidth. The chip re-reads the active weights from memory for every token it writes. On that axis the two platforms are close: 273 GB/s on GB10 against 256 GB/s theoretical on Strix Halo (~215 GB/s measured). That's under 7% apart on paper. Our Strix Halo memory bandwidth deep dive explains where the real-world loss comes from.
Measured results back this up. On GPT-OSS 120B, cited testing puts decode at about 38 tok/s on a DGX Spark vs 34 tok/s on Strix Halo (Micheal Lanham, Medium, 2026), a near-tie. Halving the Spark's memory doesn't change its bandwidth, so the 64GB model should decode any model that fits at the same speed as the 128GB model. You don't gain speed with the cheaper Spark. You just run out of room sooner.
Where the Spark genuinely wins is prompt processing (prefill), which depends on compute rather than bandwidth. The figures behind our prefill and time-to-first-token guide come from hardware-corner.net's DGX Spark benchmarks on GPT-OSS 120B: ~1,723 tok/s on GB10 vs ~340 tok/s on Strix Halo, about 5×. On a 12,000-token prompt, that's roughly 7 seconds of waiting versus roughly 35 seconds before the first word appears. If your workload is long prompts on every turn, that gap is real and it's the best argument for GB10. Just remember those numbers are from a model the 64GB Spark can't load.
| DGX Spark 64GB | DGX Spark 128GB | 128GB Strix Halo (EVO-X2) | |
|---|---|---|---|
| Price | $4,999 (Oct 23, OEM) | $6,950 | $3,649 |
| Unified memory | 64GB | 128GB | 128GB (up to 96GB GPU) |
| Bandwidth | 273 GB/s | 273 GB/s | 256 GB/s theoretical, ~215 real |
| Decode (GPT-OSS 120B) | Can't load | ~38 tok/s | ~34 tok/s |
| Prefill (GPT-OSS 120B) | Can't load | ~1,723 tok/s | ~340 tok/s |
| Software | CUDA, DGX OS | CUDA, DGX OS | ROCm / Vulkan, Windows or Linux |
| Networking | ConnectX-7 200GbE | ConnectX-7 200GbE | 2.5GbE / USB4 |
CUDA vs ROCm: what the extra $1,350 actually buys
This is the section most AMD-friendly coverage skips, so here it is plainly: the CUDA premium is real value for some buyers, not a tax on the uninformed.
What CUDA on a DGX Spark buys you:
- Everything just works. PyTorch, vLLM, TensorRT-LLM, the Hugging Face training stack, and nearly every tutorial written since 2016 assume an NVIDIA GPU. DGX OS ships the toolchain preinstalled.
- A path to the datacenter. Code you prototype on GB10 runs on the same software stack as NVIDIA's server hardware. For teams that will eventually deploy on rented H100/B200 capacity, that matters more than local tok/s.
- Fine-tuning and training tooling. Most LoRA and QLoRA tooling targets CUDA first. Our fine-tuning on a mini PC guide covers what does and doesn't work on Strix Halo today.
- Prefill compute, per the section above.
What ROCm on Strix Halo costs you: rougher edges. Inference is in good shape. llama.cpp, Ollama and LM Studio run well on the Radeon 8060S via ROCm or Vulkan, and that covers what most local-LLM users actually do. Beyond inference, expect more version pinning, more build flags, and the occasional library that hasn't been ported yet.
Who genuinely needs CUDA: ML engineers prototyping for NVIDIA production hardware, anyone whose workflow depends on vLLM or TensorRT-LLM, and researchers running training or fine-tuning jobs. Who doesn't: anyone whose day is spent chatting with, coding with, or running RAG against a local model through llama.cpp-based tools. That's most buyers of a box in this class.
Even if you do need CUDA, ask whether 64GB is enough. CUDA buyers are disproportionately the people running big models and long contexts. If that's you, the honest CUDA pick is the 128GB Spark at $6,950, or a 128GB ASUS Ascent GX10 if its price holds. The 64GB model makes sense mainly for CUDA developers whose models top out around 30B, and who want NVIDIA's software stack more than they want capacity.
Two 64GB Sparks ($9,998) vs one 128GB box
NVIDIA clearly expects people to buy these in pairs. The headline feature alongside the 64GB SKU is Sync Cluster Assistant, which, per NVIDIA's announcement, detects connected units, validates their configuration and sets up the ConnectX-7 network. Two clustered 64GB units pool 128GB and, per NVIDIA, support models up to 200 billion parameters. NVIDIA also says that in its own Qwen 3.8 27B test, two clustered 64GB systems delivered up to 1.7× the performance of one.
Treat that 1.7× as what it is: a vendor's best-case result on one model that already fits on a single unit. It isn't evidence that a pair beats a single 128GB machine on a 70B or 120B model.
The money is less ambiguous:
| Route to 128GB | Cost | vs EVO-X2 |
|---|---|---|
| 1× GMKtec EVO-X2 | $3,649 | — |
| 1× Minisforum MS-S1 Max | $3,799 | +$150 |
| 1× DGX Spark 128GB | $6,950 | +$3,301 |
| 2× DGX Spark 64GB, clustered | $9,998 | +$6,349 |
Two 64GB Sparks cost $3,048 more than one 128GB Spark for the same pooled capacity, plus a 200GbE link, two power bricks and two machines to keep in sync. Our guide to clustering two mini PCs covers the general rule: splitting a model across nodes doesn't add bandwidth per token, because each node still reads its own slice at its own speed. Clustering makes sense once you've outgrown 128GB, for MoE models above roughly 100GB. It's a poor way to reach 128GB in the first place.
If clustering is on your roadmap, the Minisforum MS-S1 Max ($3,799) is the Strix Halo box built for it. It has dual 10GbE, dual USB4 v2 (80Gbps), a PCIe x16 slot and a 2U rack-mount option. ServeTheHome's January 2026 review ran under the headline "Minisforum MS-S1 Max Review – The Best Ryzen AI Max Mini-PC Yet". Two of them cost $7,598 for 256GB pooled. That's $2,400 less than two 64GB Sparks with twice the total memory.
Which to buy: the decision table
| If you are… | Buy | Price | Why |
|---|---|---|---|
| Running 70B dense or 120B MoE models for inference | GMKtec EVO-X2 | $3,649 | Cheapest in-stock 128GB; same decode speed class as GB10 |
| A tinkerer who wants a standard ITX board and can wait for stock | Framework Desktop | $3,449 | Lowest $/GB in the catalog; currently out of stock |
| Planning to cluster, or needing 10GbE / PCIe x16 | Minisforum MS-S1 Max or Beelink GTR9 Pro | $3,799 / $4,349 | Best IO and cooling of the Strix Halo boxes |
| A CUDA developer who needs 70B+ or long context | DGX Spark 128GB or ASUS Ascent GX10 | $6,950 / $5,999 – $7,999 | CUDA, ~5× faster prefill, 200GbE clustering |
| A CUDA developer whose models stay at ~30B or below | DGX Spark 64GB (OEM) | $4,999 | The one buyer this SKU genuinely fits |
| Not blocked today and willing to wait for next-gen AMD | Hold | — | See Gorgon Halo: wait or buy |
Best for most buyers: a 128GB Strix Halo box. The EVO-X2 if you want to order today, the Framework Desktop if you can wait for stock. Best for CUDA work at scale: the 128GB DGX Spark, despite the $6,950 price. Best for clustering: the MS-S1 Max. Hardest to recommend: the 64GB DGX Spark. It's priced above every 128GB alternative and can't run the models that justify buying a box in this class.
A note on timing. NVIDIA's announcement came three weeks before the 64GB unit ships, and the GB10 partners haven't all published final prices. Meanwhile Strix Halo prices have climbed too: the EVO-X2 went from about $1,999 pre-sale to $3,649 today, as tracked in our 2026 price-increase report. We'll update the table above when OEM 64GB listings and the GX10's post-change price are live. For everything else on these two platforms, see the Strix Halo hub and the GB10 / DGX Spark hub. For NVIDIA's Windows-on-Arm sibling, read RTX Spark vs Strix Halo. For the platform matchup at equal memory, see DGX Spark vs EVO-X2.