
NVIDIA DGX Spark (GB10 Grace Blackwell)
$4,699+
NVIDIA's desktop AI supercomputer, and the CUDA-native answer to the Strix Halo boxes. The GB10 Grace Blackwell superchip pairs a 20-core Arm CPU with a Blackwell GPU and 128GB of coherent unified memory, with a 200GbE ConnectX-7 NIC so two units cluster to run 405B-class models. The headline '1 PFLOP FP4' is sparse — dense compute is roughly half (≈RTX 5070-class) — and at 273 GB/s, memory bandwidth is the real ceiling. You're buying CUDA + capacity, not bandwidth.
Affiliate links — We earn a commission on qualifying purchases at no cost to you.
Specifications
| Chip | GB10 Grace Blackwell Superchip |
| CPU | 20-core Arm (10× Cortex-X925 + 10× A725) |
| GPU | Blackwell (5th-gen Tensor Cores), CUDA-native |
| AI Performance | Up to 1 PFLOP FP4 sparse (~500 TFLOPS dense) |
| Unified Memory | 128GB LPDDR5X |
| Memory Bandwidth | 273 GB/s |
| Networking | ConnectX-7 200GbE (2-unit clustering), 10GbE RJ-45 |
| Storage | 4TB NVMe (self-encrypting) |
| OS | NVIDIA DGX OS (Ubuntu-based) |
Pros
- CUDA-native + full NVIDIA/DGX software stack — best dev ergonomics for AI work
- 128GB unified in a 1.2kg box; 2-unit 200GbE stacking reaches 405B-class locally
- Drop-in compatibility with the datacenter toolchain
Cons
- 273 GB/s bandwidth is low for the price — token throughput lags Apple Ultra and GPUs
- Headline '1 PFLOP FP4' is sparse-only; dense compute ~5070-class, not datacenter-class
- NVIDIA raised the official price from $3,999 to $4,699 (2026-02-27) on memory supply; stock is thin
Related Articles
Why Local-AI Mini PC Prices Doubled in 2026 — And Whether to Buy Now or Wait
The GMKtec EVO-X2 went from ~$1,999 to ~$3,399. NVIDIA repriced the DGX Spark from $3,999 to $4,699 on 2026-02-27, citing memory supply. Here's the SKU-by-SKU price damage, why DRAM did it, and why waiting is the wrong call — with TrendForce's own 2027 forecast as the evidence.
Beelink GTR9 Pro Review (2026): The 128GB Strix Halo Box With Dual 10GbE — Real Local-LLM Benchmarks & Who Should Buy
The Beelink GTR9 Pro is a ~$1,899–$1,999 128GB Ryzen AI Max+ 395 mini PC that runs GPT-OSS 120B at ~31 tok/s (~120W) but a dense 70B at only ~5 tok/s. Its real hook is dual 10GbE — plus a real caveat: reported 10GbE instability under heavy GPU load. Benchmarks, thermals, and how it stacks up against the EVO-X2, Framework Desktop, and DGX Spark.
How Much Does It Cost to Run a Local AI Mini PC 24/7? Idle Power, Electricity & True TCO in 2026
A local-AI box left on 24/7 costs roughly $15–$60 a year in electricity — and the spread is almost all idle power, because a personal LLM server sits idle 95%+ of the time. A Strix Halo box idles ~13W (~$20/yr); a DGX Spark idles ~37–40W (~$55/yr). For an always-on box, buy for idle watts, not peak specs.
Why Your Local LLM Feels Slow on a Strix Halo Mini PC: The Prefill (Time-to-First-Token) Problem
On a Ryzen AI Max+ 395 box, token generation ties a DGX Spark — but prompt processing (prefill) is ~5× slower (~340 vs ~1,700 tok/s on gpt-oss 120B). On long prompts, that's the delay that makes RAG and coding agents feel sluggish. Here's who it hits, why, and how to fix it.
Related Products
Disclosure: Some links on this page are affiliate links. We may earn a commission if you make a purchase — at no extra cost to you. This helps support our independent reviews.


