RTX Spark vs Strix Halo: Buy a 128GB Box Now, or Wait for NVIDIA's October Launch?
NVIDIA's RTX Spark (N1X) ships in October with up to 128GB of unified memory — and NVIDIA's own porting guide puts it at 300 GB/s, not the 600 GB/s everyone is quoting. Here's what that actually changes about the box you buy this week.
DataHardware Team
Our Top Pick

Framework Desktop (Ryzen AI Max+ 395, 128GB)
$3,449Quick answer: if you need 128GB today and you work in llama.cpp or ROCm, buy now — RTX Spark is a bandwidth bump, not a category break. NVIDIA has confirmed an October 2026 launch window and a 128GB unified-memory ceiling for RTX Spark (N1X), and its own Windows on Arm porting guide states the memory interface is 256 bits wide with a bandwidth of 300 GB/s — the widely-repeated 600 GB/s is NVLink-C2C chip-to-chip interconnect from the GB10 platform, not memory bandwidth. NVIDIA has published no retail price; the $2,899 figure circulating is a Morgan Stanley analyst estimate reported secondhand, not a price. Against that, a Framework Desktop is $3,449 and a GMKtec EVO-X2 is $3,649 today, at roughly 215 GB/s real-world. Wait only if you need CUDA on Windows.
What NVIDIA has actually confirmed about RTX Spark (and what it hasn't)
Almost everything currently ranking for "RTX Spark" is a rewrite of the same IFA 2026 press briefing, and the rewrites have drifted. So here is the line between confirmed and unconfirmed, drawn explicitly.
Confirmed by NVIDIA:
- The name and the part. RTX Spark is the retail name; N1X is the SoC. NVIDIA's GPU Identification page enumerates three shipping GPU variants by PCIe device ID: a 6,144-core Blackwell RTX GPU in laptop and desktop forms, and a 5,120-core laptop part.
- The CPU. An Arm v9.2 Grace CPU. The System Overview describes the 20-core part as two clusters of ten, each mixing five Cortex-X925 performance cores with five Cortex-A725 efficiency cores. The CPU-Specific Details page confirms two SKUs: 18-core and 20-core.
- Memory. Up to 128GB of LPDDR5 shared by CPU and iGPU — genuine unified memory, not a carve-out.
- Bandwidth. 300 GB/s across a 256-bit interface. This is first-party and it is the number almost nobody is quoting.
- The OS. Windows on Arm — which is the entire strategic point of the product.
- The partners and the window. October 2026, with Acer, ASUS, Dell, Gigabyte, HP, Lenovo, Microsoft and MSI all announcing devices.
Not confirmed by anyone:
- Retail price. Not for a single SKU, from a single OEM. See the section below on why the $2,899 number you keep seeing is not one.
- Tokens per second. No reviewer has this hardware. Any tok/s figure you find for N1X today is extrapolated or invented — including, we should say plainly, any we might have printed. We haven't, and won't until someone measures it.
- The memory configuration ladder. Reporting from the IFA briefing puts the flagship's range at 24GB to 128GB and the 5,120-core laptop part at 24GB to 32GB, but which OEM ships which tier at which price is unknown.
That second list is the whole reason this decision is hard. You are being asked to defer a purchase against a product whose price and performance are both unpublished.
RTX Spark vs DGX Spark: same family, different product
These get conflated constantly, and the confusion is doing real damage to buying decisions. They are not the same machine.
The NVIDIA DGX Spark — and its cheaper sibling, the ASUS Ascent GX10 — is built on GB10 Grace Blackwell. It runs DGX OS, which is Linux. It ships with a 200GbE ConnectX-7 NIC so two boxes pool into 256GB. It is sold as a developer workstation at $4,699+, and it lives in NVIDIA's datacenter-adjacent channel. We covered that platform against AMD in DGX Spark vs Strix Halo, and the full GB10 hub collects everything we have measured on it; this post is its October sequel.
RTX Spark is N1X, and it is a consumer Windows PC part. It goes into laptops and compact desktops sold by the same OEMs that sell you a ThinkPad. It runs Windows on Arm. It has no ConnectX-7. Nobody is going to hand you DGX OS.
| RTX Spark (N1X) | DGX Spark (GB10) | Strix Halo (AI Max+ 395) | |
|---|---|---|---|
| Status | October 2026 window | Shipping | Shipping |
| CPU | Arm v9.2, 18 or 20-core Grace | 20-core Arm Grace | 16-core Zen 5 (x86) |
| GPU | Blackwell RTX, 5,120 or 6,144 cores | Blackwell, 6,144 cores | Radeon 8060S, 40 CU RDNA 3.5 |
| Max unified memory | 128GB LPDDR5 | 128GB LPDDR5X | 128GB LPDDR5X-8000 |
| Memory bandwidth | 300 GB/s (NVIDIA, 256-bit) | 273 GB/s | 256 GB/s theoretical, ~215 GB/s real |
| OS | Windows on Arm | DGX OS (Linux) | Windows or Linux (x86) |
| Compute API | CUDA | CUDA | ROCm / Vulkan |
| Clustering | None announced | 200GbE ConnectX-7 | Ethernet / USB4 (DIY) |
| Price | Unpublished | $4,699+ | $3,449 and up |
Why this matters for local inference rather than for spec-sheet trivia: the driver and runtime stack you inherit is different. A DGX Spark hands you a Linux box where CUDA, PyTorch and vLLM behave the way every tutorial on the internet assumes. An RTX Spark hands you Windows on Arm, where the question is not "does CUDA exist" — it does — but "has my specific inference stack been compiled for Arm64 Windows yet." Those are very different risks.
The bandwidth question everyone is getting wrong
Here is the correction that no competing page is making, and it is worth stating twice:
The 600 GB/s figure attached to Spark hardware is NVLink-C2C — the chip-to-chip interconnect between the Grace CPU tile and the Blackwell GPU tile inside a single package. It is not memory bandwidth. NVIDIA's own RTX Spark porting guide puts N1X memory bandwidth at 300 GB/s on a 256-bit interface.
Three different numbers get blended in the coverage, and they do three unrelated jobs:
- ~600 GB/s — NVLink-C2C. How fast the CPU die talks to the GPU die inside the package. It never gates token generation, because the weights aren't travelling that path.
- 300 GB/s (N1X) / 273 GB/s (GB10) — memory bandwidth. How fast the chip streams model weights out of LPDDR. This is what sets decode speed.
- ~25 GB/s — box-to-box. What two clustered DGX Sparks get between them, which is why clustering buys capacity and not speed.
Quoting the 600 GB/s figure as memory bandwidth overstates N1X by exactly 2x, and it is the single error that makes RTX Spark look like a generational leap over the boxes you can already buy. It isn't one. At 300 GB/s it is about 10% ahead of the DGX Spark's 273 GB/s and about 17% ahead of Strix Halo's 256 GB/s theoretical — more against the ~215 GB/s these boxes actually achieve, but still the same order of magnitude. We laid out why that ceiling matters more than any other spec in the Strix Halo memory bandwidth guide.
For the scale that would be a leap: the Mac Studio M3 Ultra ran at 819 GB/s, roughly 2.7x N1X. That is what a real bandwidth generation looks like — though Apple retired the M3 Ultra in September 2026 for the M5 Ultra, so that particular box is a secondary-market buy now. The point stands for the line: Apple Silicon is still the answer for anyone whose constraint is speed rather than software.
What 300 GB/s means for tokens/sec on a 70B and on GPT-OSS 120B
We are not going to print a tokens-per-second figure for RTX Spark. The hardware has not shipped, nobody has benchmarked it, and a number invented from a bandwidth spec is a guess wearing a lab coat. What we can do is show you the measured figures for the two platforms it sits between, all of them sourced on our benchmarks reference, and let you reason about the band N1X should land in.
| Hardware | Model | Generation | Prompt processing | Source |
|---|---|---|---|---|
| DGX Spark (GB10) | GPT-OSS 120B MXFP4 | 38.55 tok/s | 1,723 tok/s | hardware-corner.net, llama.cpp CUDA |
| Strix Halo, 128GB | GPT-OSS 120B MXFP4 | 34.13 tok/s | 340 tok/s | hardware-corner.net, same methodology |
| Beelink GTR9 Pro | GPT-OSS 120B | 31.41 tok/s | not stated | ServeTheHome, LM Studio out-of-box |
| Beelink GTR9 Pro | Llama 3.3 70B | ~5 tok/s | not stated | ServeTheHome, approximate |
Read that table carefully, because it contains the most important fact in this post. On generation, the CUDA box beats the AMD box by about 13% on the same model with the same runtime — 38.55 against 34.13 tok/s. That is a bandwidth-shaped gap, exactly as theory predicts. If N1X's 300 GB/s holds, it should sit modestly above the DGX Spark row. Modestly. Not double.
On prompt processing, the gap is a chasm: 1,723 against 340 tok/s, roughly 5x. That is compute-bound rather than bandwidth-bound, and it is where NVIDIA silicon genuinely dominates. If you paste 30,000-token codebases into a model all day, that difference is the one you will feel — we broke it down in the prefill and time-to-first-token analysis. If you mostly chat, you will barely notice it.
One honest gap worth naming: there is no published dense-70B figure for GB10 on comparable methodology. The ~5 tok/s Llama 3.3 70B number above is Strix Halo only. Anyone showing you a dense-70B head-to-head between these platforms is extrapolating. Dense 70B models are brutal on every box in this class — a mixture-of-experts model like GPT-OSS 120B activates a fraction of its weights per token, which is why a 120B model runs seven times faster than a 70B one in that table. Model architecture matters more than your choice of box.
Price: $3,449 on the table today vs a rumor in October
This is where the decision actually gets made, and where the competing coverage is least useful.
What you can pay this week, every figure catalog-verified with a date:
| Box | Price | Verified | Note |
|---|---|---|---|
| Framework Desktop | $3,449 | 2026-08-30 | Cheapest credible 128GB box; direct-only |
| GMKtec EVO-X2 | $3,649 | 2026-08-30 | The flagship Strix Halo appliance |
| Minisforum MS-S1 Max | $3,799 | 2026-09-13 | Listed sold out — not buyable today |
| Beelink GTR9 Pro | $4,349 | 2026-08-30 | Dual 10GbE, best I/O of the group |
| NVIDIA DGX Spark | $4,699+ | 2026-02-27 | CUDA + clustering, today, on Linux |
| HP Z2 Mini G1a | $5,349 – $7,406 | 2026-09-20 | vPro, ECC, 3-year warranty |
| ASUS Ascent GX10 | $6,449 – $7,999 | 2026-09-20 | Street price far above its list price |
Now the rumor. The $2,899 figure attached to RTX Spark traces to a Morgan Stanley note whose channel checks with PC brands at Computex suggested N1X systems would need to price around that level, with lower-tier N1 systems near $1,799. It was reported secondhand by VideoCardz and repeated everywhere since. Three things are wrong with treating it as a price:
- It is an analyst's floor, not a manufacturer's price. The claim is about what OEMs would need to charge, which is a cost argument, not a sticker.
- It predates the current memory market. NVIDIA raised the DGX Spark's official price from $3,999 to $4,699 on 2026-02-27, explicitly citing memory supply. Apple pulled its 256GB and 512GB Mac Studio configurations during the same squeeze. TrendForce has DRAM contract prices up 93–98% quarter-on-quarter. We traced the whole pattern in the 2026 price increase analysis, and a 128GB LPDDR5 consumer launch is precisely the kind of product that gets repriced upward between briefing and shelf.
- $2,899 is the system floor, not the 128GB SKU. Even if it holds, it plausibly describes an entry configuration. The flagship 128GB desktop is a different machine and almost certainly a different number.
Look at the ASUS Ascent GX10 row above for what launch-window pricing does in practice: a $4,699-class platform whose actual in-stock street price is $6,449 – $7,999. That is not a forecast. That is what the last GB10-family launch is costing buyers right now.
CUDA on Windows on Arm: the real reason to care
Strip away the specs and one genuine differentiator survives: RTX Spark puts CUDA on a Windows machine with 128GB of unified memory. No AMD box can offer that, and it isn't close.
If your work depends on CUDA — most fine-tuning scripts, TensorRT-LLM, anything built on custom kernels, the long tail of research code that has never seen ROCm — then Strix Halo isn't a cheaper alternative. It's a different machine that doesn't run your workload. That is a legitimate reason to wait, and it's the only one we'd defend without qualification.
But weigh it against the risk sitting right next to it: Windows on Arm is a different binary target. NVIDIA's own porting documentation describes three application models — native Arm64, x86_64 emulation, and 32-bit x86 emulation — and notes that emulation lets legacy applications run while native Arm applications offer better performance. The existence of a porting guide is itself the tell: NVIDIA is asking developers to do work.
For local AI specifically, the question you need answered before you spend is narrow and practical: does my stack have a Windows-on-Arm64 build on day one? llama.cpp, Ollama, LM Studio, ComfyUI, PyTorch — each one is its own answer, and "runs under emulation" is not the same answer as "runs natively on the GPU." Qualcomm's Snapdragon X launch is the cautionary precedent: capable silicon that spent its first year waiting for software.
The Strix Halo boxes have the opposite profile. ROCm is scrappier than CUDA, but these are x86 machines. Every binary you already run, runs. There is no porting question because there is nothing to port.
Who should wait, and who should buy this week
Wait for RTX Spark if:
- Your workflow is CUDA-dependent and you want Windows rather than Linux. This is the strong case.
- You are not blocked today — you have a working setup and this is an upgrade, not an enabler.
- You want third-party benchmarks before spending $3,000+. Entirely reasonable, and the reason we're not printing a tok/s number for it.
- You want a laptop. No Strix Halo box competes here; that comparison doesn't exist yet.
Buy a 128GB box this week if:
- You are productive in llama.cpp, Ollama, or LM Studio. These are solved on Strix Halo today and unknown on Arm64 Windows.
- You need ~96GB of GPU-allocatable memory this quarter. Use the VRAM calculator to check what your target model actually needs first.
- You would rather pay a known price than gamble on launch-window supply and a rumored one.
- Your models are MoE — GPT-OSS 120B, Qwen3 32B. At 34 tok/s measured on hardware you can order today, the 13% a CUDA box adds is not worth a three-month wait.
The specific picks: Framework Desktop at $3,449 if you want the cheapest credible entry and a tinkerer's platform. GMKtec EVO-X2 at $3,649 if you want a quiet always-on inference appliance — see them side by side. Beelink GTR9 Pro at $4,349 if dual 10GbE matters. And if you need CUDA now rather than in October, the DGX Spark at $4,699+ already exists — compared against the GX10 here, or against the EVO-X2 if you are still weighing CUDA against price.
What waiting actually costs you
Waiting is a position, and positions have costs. Price them honestly:
- Three weeks minimum to first shipments — and "October, depending on OEM" means late October for some.
- Two to six weeks for credible reviews after that. If your rule is "wait for benchmarks," your real timeline is November or December, not October.
- First-wave premium. The ASUS Ascent GX10 is the live example: $6,449 – $7,999 in stock against a $4,699-class list price. The DGX Spark itself went up $700 before general availability.
- Supply. GB10 availability has been thin enough that NVIDIA's own developer forums carry running threads about it, during the same DRAM shortage that pushed Apple to pull its high-memory Mac Studio configs.
- Two to three months of not running anything locally. Whatever you'd have built in that window, you don't build.
There is a symmetric argument on the AMD side too: the Ryzen AI Max 400 "Gorgon Halo" refresh lands in the same Q4 window with 192GB of unified memory. We ran that arithmetic in the Gorgon Halo wait-or-buy analysis and reached the same shape of conclusion: capacity is moving, bandwidth is barely moving, so waiting buys you a bigger tank rather than a faster engine. If you are waiting for RTX Spark and Gorgon Halo and third-party reviews of both, you have talked yourself into 2027.
Bottom line
RTX Spark is a real product with a real date and a genuine advantage — CUDA on Windows with 128GB of unified memory, which nothing else offers. It is not the 2x bandwidth leap the 600 GB/s misquote implies. At NVIDIA's own 300 GB/s, it lands about 10% above the DGX Spark and modestly above Strix Halo, on a platform whose price is unpublished and whose Arm64 software ecosystem is unproven for local inference.
So: CUDA-dependent, unblocked, patient → wait. Everyone else → buy. The Framework Desktop at $3,449 runs GPT-OSS 120B at measured speed, on software that works, at a price with a date attached — and if RTX Spark turns out to be everything the spec sheet suggests, you will have spent three productive months finding out exactly what you need from the next one. Start with the tiered buyer's guide if you're still choosing a tier, or the KV-cache math if you're wondering why 128GB isn't 128GB of usable context.
We will update this post with measured RTX Spark figures the moment a credible third party publishes them. Until then, the absence of a number is the number.