DGX Spark vs Strix Halo: Which 128GB Box Should Run Your Local LLM?
Both put 128GB of unified memory on your desk for local AI. One costs $3,999 and speaks CUDA; the other costs ~$2,000. Here's the bandwidth-vs-capacity math that actually decides it.
DataHardware
Our Top Pick

GMKtec EVO-X2 (Ryzen AI Max+ 395, 128GB)
$3,399 – $3,499Quick answer: For a single desktop box running 70B-class models locally, the AMD Strix Halo mini PCs (GMKtec EVO-X2, Beelink GTR9 Pro, Framework Desktop) win on price — about $1,900–$2,200 for 128GB of unified memory versus $3,999+ for NVIDIA's DGX Spark — and their memory bandwidth (~215–256 GB/s) is close enough to the Spark's 273 GB/s that token speed lands in the same ballpark. Buy the DGX Spark only if you need CUDA-native tooling or plan to cluster two boxes over 200GbE to reach 405B-class models. And know this up front: neither is a speed play. Both are capacity plays. If you want bandwidth at 128GB, Apple's M3 Ultra (819 GB/s) is roughly 3× either of these.





The two ways to put 128GB on your desk
For two years the only way to load a 70B model locally was to bolt together multiple 24GB GPUs — expensive, power-hungry, and loud. In 2026 there are two clean single-box answers, and they come from opposite directions.
AMD's Strix Halo — officially the Ryzen AI Max+ 395 — fuses a 16-core Zen 5 CPU, a 40-CU Radeon 8060S iGPU, and 128GB of LPDDR5X-8000 unified memory on one package. Up to 96GB of that is allocatable as VRAM. It shows up in a fleet of mini PCs from GMKtec, Beelink, Framework, Minisforum, and HP.
NVIDIA's GB10 "Grace Blackwell" takes the datacenter approach: a 20-core Arm CPU welded to a Blackwell GPU, also with 128GB of coherent unified memory, sold as the DGX Spark and the cheaper ASUS Ascent GX10. Its trump card is a 200GbE ConnectX-7 NIC that lets two boxes cluster into one 256GB pool.





Bandwidth is the number that matters — and it's a wash
Here's the counterintuitive part. Once a model fits in memory, the thing that caps how fast it generates tokens is memory bandwidth, not compute or capacity. And on bandwidth, these two platforms are nearly tied:
- DGX Spark / ASUS GX10 (GB10): 273 GB/s
- Strix Halo boxes (AI Max+ 395): 256 GB/s theoretical, ~215 GB/s real-world
That's a ~15–25% gap on paper and smaller in practice — not the 2–3× difference the Spark's price tag implies. A dense 70B model runs at single-digit tokens/sec on both. For comparison, a discrete GPU moves 800–1,000 GB/s, and Apple's M3 Ultra hits 819 GB/s. So if your mental model was "the $3,999 NVIDIA box must be much faster," correct it now: it isn't, for single-box inference.
What you're actually paying the NVIDIA premium for is the software stack and clustering, not throughput.





Where the DGX Spark earns its price
Two things, and they're real:
1. CUDA-native everything. The Spark runs NVIDIA's DGX OS and the full CUDA toolchain. If your workflow assumes NVIDIA — most fine-tuning scripts, many inference servers, anything that hasn't been ported to ROCm — it just works. On Strix Halo you live in ROCm/Vulkan/llama.cpp land, which is improving fast but still has rough edges for GPU-compute beyond inference.
2. 200GbE clustering. Two Sparks over ConnectX-7 form a 256GB pool that can run 405B-class models — something no single Strix Halo box can touch. If your roadmap is "start with one, scale to a cluster," that NIC is the whole point.
The catch: the headline "1 PFLOP FP4" is the sparse number. Dense compute is roughly half that (~500 TFLOPS, about RTX 5070-class), and street prices have crept above the $3,999 MSRP with thin supply.





Where Strix Halo wins: price and availability
The Strix Halo boxes do the same core job — hold a 70B model in 128GB and generate usable tokens — for roughly half the money, and they're actually in stock. The trade-offs between them are about I/O, cooling, and chassis, not the silicon:
- Framework Desktop — $1,999, the cheapest credible 128GB box; standard mini-ITX board, best Linux/tinkerer story. Direct-only (no Amazon).
- GMKtec EVO-X2 — $1,999–$2,199, the flagship; quiet, dual-M.2 expandable, a practical always-on inference appliance.
- Beelink GTR9 Pro — $1,899–$1,999, the best-connected; dual 10GbE + dual USB4 for pulling models off a fast NAS or wiring boxes together. Watch for driver-dependent 10GbE instability under heavy GPU load.
- Minisforum MS-S1 Max — ~$2,900, enthusiast I/O: dual 10GbE, dual USB4 v2, a PCIe x16 slot, and a 2U-rack option for DIY clusters.
- HP Z2 Mini G1a — $3,300+, the business-grade pick: vPro, ECC, 3-year warranty, true ~2.5L SFF. You pay the HP brand tax.
All of them share the same ceiling: ~215–256 GB/s bandwidth and soldered, non-upgradable memory — you buy the 128GB SKU up front or nothing.





So which one should you buy?
Buy a Strix Halo box if your goal is to run 70B-class models locally for chat, coding, or RAG, you want to spend ~$2,000 instead of ~$4,000, and you're comfortable in the ROCm/llama.cpp world. Start with the Framework Desktop or GMKtec EVO-X2 for value, or the Beelink GTR9 Pro if you want 10GbE networking.
Buy a DGX Spark (or ASUS GX10) if you need CUDA-native tooling for fine-tuning or NVIDIA-only inference servers, or you plan to cluster two boxes to run 405B models. The ASUS Ascent GX10 gets you the identical GB10 platform on cheaper SSD tiers and is the more available of the two.
Buy neither if raw speed at 128GB is the priority — then Apple's M3 Ultra at 819 GB/s is the bandwidth king, at the cost of CUDA and (in 2026) a 96GB-only config during the DRAM shortage.





Bottom line
These aren't really competitors so much as two answers to different questions. "How cheaply can I run a 70B model on my desk?" → Strix Halo, ~$2,000. "How do I run NVIDIA's stack locally and scale to 405B?" → DGX Spark, $3,999+ per box. The one thing both answer the same way: don't expect GPU speed. 128GB at ~256 GB/s is about fitting the model, not racing it.