NVIDIA DGX Spark (GB10 Grace Blackwell)
Node KitsFeatured

NVIDIA DGX Spark (GB10 Grace Blackwell)

4/5

$4,699+

NVIDIA's desktop AI supercomputer, and the CUDA-native answer to the Strix Halo boxes. The GB10 Grace Blackwell superchip pairs a 20-core Arm CPU with a Blackwell GPU and 128GB of coherent unified memory, with a 200GbE ConnectX-7 NIC so two units cluster to run 405B-class models. The headline '1 PFLOP FP4' is sparse — dense compute is roughly half (≈RTX 5070-class) — and at 273 GB/s, memory bandwidth is the real ceiling. You're buying CUDA + capacity, not bandwidth.

Affiliate links — We earn a commission on qualifying purchases at no cost to you.

Specifications

ChipGB10 Grace Blackwell Superchip
CPU20-core Arm (10× Cortex-X925 + 10× A725)
GPUBlackwell (5th-gen Tensor Cores), CUDA-native
AI PerformanceUp to 1 PFLOP FP4 sparse (~500 TFLOPS dense)
Unified Memory128GB LPDDR5X
Memory Bandwidth273 GB/s
NetworkingConnectX-7 200GbE (2-unit clustering), 10GbE RJ-45
Storage4TB NVMe (self-encrypting)
OSNVIDIA DGX OS (Ubuntu-based)

Pros

  • CUDA-native + full NVIDIA/DGX software stack — best dev ergonomics for AI work
  • 128GB unified in a 1.2kg box; 2-unit 200GbE stacking reaches 405B-class locally
  • Drop-in compatibility with the datacenter toolchain

Cons

  • 273 GB/s bandwidth is low for the price — token throughput lags Apple Ultra and GPUs
  • Headline '1 PFLOP FP4' is sparse-only; dense compute ~5070-class, not datacenter-class
  • NVIDIA raised the official price from $3,999 to $4,699 (2026-02-27) on memory supply; stock is thin

Related Articles

Guide14 min read

How Much Context Can a 128GB Mini PC Actually Hold? The KV Cache Math Nobody Runs Before Buying

Everyone sizes a unified-memory box against model weights. Almost nobody sizes it against the KV cache — and on a 128K-token agent run, the cache is the number that decides whether the job finishes. Here's the math, box by box, plus the KV-quantization trade that makes long first prompts slower, not faster.

Guide14 min read

Can You Fine-Tune an LLM on a 128GB Mini PC? (And Which Box to Buy in 2026)

LoRA and QLoRA on 20–30B models are practical on 128GB of unified memory; full fine-tunes stop near 12B; dense 70B training doesn't happen on any box in this class. Fine-tuning is the one local-AI workload where the software stack — not memory bandwidth — decides which machine you buy. Here's the capability table, the CUDA tax, and the cloud break-even.

Guide14 min read

Best Mini PC for a Local Coding Agent in 2026 — Why Prefill, Not Tokens/Sec, Decides Your Box

A coding agent re-sends your whole repo context every single turn, which makes it a prefill-bound workload. That flips the buying logic: ~1,700 tok/s prompt processing on a GB10 box vs ~340 tok/s on Strix Halo, while generation is a near-tie. Here's the per-budget verdict, the memory math, and the boxes to skip — with the caveat that the 2026 DRAM spike has narrowed the price gap to about $1,250.

Guide14 min read

Can You Cluster Two Mini PCs to Run Bigger Local LLMs? (2026 Reality Check)

AMD published a four-node Framework Desktop cluster running a trillion-parameter model, and two DGX Sparks pool 256GB over a single cable. Here's when a second box actually helps, when it makes things slower, and which mini PC to buy if clustering is on your roadmap.

Disclosure: Some links on this page are affiliate links. We may earn a commission if you make a purchase — at no extra cost to you. This helps support our independent reviews.