Topic Hub

GB10 / DGX Spark: CUDA-Native Desktop AI

GB10 is NVIDIA's Grace Blackwell superchip: a 20-core Arm CPU fused to a Blackwell GPU with 128GB of coherent unified memory, CUDA-native. The DGX Spark and ASUS Ascent GX10 both ship it, with a 200GbE ConnectX-7 NIC so two units cluster to run 405B-class models. Memory bandwidth is 273 GB/s — a similar ballpark to Strix Halo — so what you're really buying here is CUDA and capacity, not bandwidth.

Top Picks

NVIDIA DGX Spark (GB10 Grace Blackwell)

NVIDIA DGX Spark (GB10 Grace Blackwell)

$6,950

  • Chip: GB10 Grace Blackwell Superchip
  • CPU: 20-core Arm (10× Cortex-X925 + 10× A725)
  • GPU: Blackwell (5th-gen Tensor Cores), CUDA-native
Check Price on Amazon
ASUS Ascent GX10 (GB10 Grace Blackwell)

ASUS Ascent GX10 (GB10 Grace Blackwell)

$5,999 – $7,999

  • Chip: GB10 Grace Blackwell Superchip
  • CPU: 20-core Arm (10× Cortex-X925 + 10× A725)
  • GPU: Blackwell (5th-gen Tensor Cores), CUDA-native
Check Price on Amazon

Comparisons

Models That Run on These Boxes

Llama

Llama 3.3 70B Instruct

Meta's flagship dense chat model — strong general reasoning, coding, and instruction-following that rivals much larger models. The default pick when a single big box has the memory to spare.

Llama

Llama 4 Scout

A mixture-of-experts model — only a fraction of its 109B total parameters are active per token, so it loads like a large model but runs far faster than its size suggests. A natural fit for unified-memory boxes with a long context window.

Qwen3

Qwen3 8B

A capable small chat-and-reasoning model that fits comfortably even on budget boxes — a good everyday assistant when you don't need frontier-level depth.

Qwen3

Qwen3 14B

The mid-size Qwen3 — noticeably stronger reasoning and coding than the 8B while still fitting mid-range hardware at Q4.

Qwen3

Qwen3 32B

Qwen3's large dense model — near-flagship quality that still runs on a single unified-memory box, a strong local alternative to 70B-class models at lower memory cost.

DeepSeek-R1

DeepSeek-R1 Distill Qwen 7B

R1's chain-of-thought reasoning distilled into a 7B Qwen base — the lightest way to get R1-style step-by-step reasoning locally, small enough for budget boxes.

DeepSeek-R1

DeepSeek-R1 Distill Qwen 14B

The 14B R1 distill — a good balance of reasoning depth and hardware cost, running comfortably on mid-range unified-memory boxes.

DeepSeek-R1

DeepSeek-R1 Distill Qwen 32B

The strongest Qwen-based R1 distill — heavy reasoning that still fits a single large-memory box, the sweet spot for local R1-style work.

DeepSeek-R1

DeepSeek-R1 Distill Llama 70B

R1's reasoning distilled onto a Llama 70B base — the highest-quality distill, and the one to reach for when the full 671B R1 won't fit (it never does locally).

Gemma 3

Gemma 3 4B

Google's smallest current Gemma — tiny memory footprint and multimodal, ideal for the most memory-constrained budget boxes and always-on assistants.

Gemma 3

Gemma 3 12B

The mid-size Gemma 3 — solid general-purpose quality with a long context window, fitting mid-range hardware at Q4.

Gemma 3

Gemma 3 27B

The largest Gemma 3 — frontier-adjacent quality for a dense open model, comfortably runnable on a single large-memory box.

GPT-OSS

GPT-OSS 20B

OpenAI's small open-weight MoE — its 21B total parameters load like a mid-size model, but because only a fraction are active per token it runs fast on modest hardware.

GPT-OSS

GPT-OSS 120B

OpenAI's large open-weight MoE — its 117B total parameters fit a single high-memory box, and because only a fraction are active per token it runs far quicker than its total suggests.

Phi

Phi-4

Microsoft's 14B model punches well above its size on math and reasoning — a compact, mid-range-friendly pick when you want strong reasoning without a big memory bill.

Mistral

Mistral Small 3.2 24B

Mistral's 24B dense model — fast, capable, and low-latency for its class, sitting neatly between the mid-size and large boxes.

Related Articles

Comparison

DGX Spark 64GB vs Strix Halo 128GB: NVIDIA's $4,999 Spark Has Half the Memory — Should You Buy It?

On October 2 NVIDIA launched a 64GB DGX Spark at $4,999 and moved the 128GB Spark to $6,950. Every 128GB Strix Halo box in our catalog costs less than the new entry Spark. Here's the price-per-GB math, what actually fits in 64GB, and who should still pay for CUDA.

Read
Comparison

RTX Spark vs Strix Halo: Buy a 128GB Box Now, or Wait for NVIDIA's October Launch?

NVIDIA's RTX Spark (N1X) ships in October with up to 128GB of unified memory — and NVIDIA's own porting guide puts it at 300 GB/s, not the 600 GB/s everyone is quoting. Here's what that actually changes about the box you buy this week.

Read
Guide

How Much Context Can a 128GB Mini PC Actually Hold? The KV Cache Math Nobody Runs Before Buying

Everyone sizes a unified-memory box against model weights. Almost nobody sizes it against the KV cache — and on a 128K-token agent run, the cache is the number that decides whether the job finishes. Here's the math, box by box, plus the KV-quantization trade that makes long first prompts slower, not faster.

Read
Guide

Can You Fine-Tune an LLM on a 128GB Mini PC? (And Which Box to Buy in 2026)

LoRA and QLoRA on 20–30B models are practical on 128GB of unified memory; full fine-tunes stop near 12B; dense 70B training doesn't happen on any box in this class. Fine-tuning is the one local-AI workload where the software stack — not memory bandwidth — decides which machine you buy. Here's the capability table, the CUDA tax, and the cloud break-even.

Read
Guide

Best Mini PC for a Local Coding Agent in 2026 — Why Prefill, Not Tokens/Sec, Decides Your Box

A coding agent re-sends your whole repo context every single turn, which makes it a prefill-bound workload. That flips the buying logic: ~1,700 tok/s prompt processing on a GB10 box vs ~340 tok/s on Strix Halo, while generation is a near-tie. Here's the per-budget verdict, the memory math, and the boxes to skip — with the caveat that the 2026 DRAM spike has narrowed the price gap to about $1,250.

Read
Guide

Can You Cluster Two Mini PCs to Run Bigger Local LLMs? (2026 Reality Check)

AMD published a four-node Framework Desktop cluster running a trillion-parameter model, and two DGX Sparks pool 256GB over a single cable. Here's when a second box actually helps, when it makes things slower, and which mini PC to buy if clustering is on your roadmap.

Read
Economics

How Much Does It Cost to Run a Local AI Mini PC 24/7? Idle Power, Electricity & True TCO in 2026

A local-AI box left on 24/7 costs roughly $15–$60 a year in electricity — and the spread is almost all idle power, because a personal LLM server sits idle 95%+ of the time. A Strix Halo box idles ~13W (~$20/yr); a DGX Spark idles ~37–40W (~$55/yr). For an always-on box, buy for idle watts, not peak specs.

Read
Comparison

DGX Spark vs Strix Halo: Which 128GB Box Should Run Your Local LLM?

Both put 128GB of unified memory on your desk for local AI. One costs $4,699 and speaks CUDA; the other costs ~$3,449. Here's the bandwidth-vs-capacity math that actually decides it.

Read
Guide

The Best Mini PC for Local LLMs in 2026 (By Model Size and Budget)

From a $329 agent host to a 128GB box that runs 70B models, here's the right mini PC for local AI at every tier — matched to the model you actually want to run.

Read

Frequently Asked Questions

What is the GB10 / DGX Spark?

GB10 is NVIDIA's Grace Blackwell superchip — a 20-core Arm CPU paired with a Blackwell GPU and 128GB of coherent unified memory, running the full CUDA stack. Both the NVIDIA DGX Spark and the ASUS Ascent GX10 are built on it.

How is GB10 different from a Strix Halo box?

CUDA. GB10 runs NVIDIA's native software stack, so tools that assume CUDA work out of the box. Memory bandwidth is 273 GB/s — a similar ballpark to Strix Halo — so the differentiator is the software ecosystem and capacity, not raw bandwidth.

DGX Spark or ASUS Ascent GX10?

Both use the same GB10 silicon, 128GB unified memory, 273 GB/s bandwidth, and ConnectX-7 200GbE. The Ascent GX10 ships more widely and offers larger SSD tiers, but at $5,999–$7,999 it no longer undercuts the DGX Spark ($4,699+), which carries NVIDIA's DGX OS and brand.

Can I run 405B-class models on GB10?

Not on a single box, but two units cluster over the 200GbE ConnectX-7 NIC to reach 405B-class models locally.