192GB vs 128GB Unified Memory: What the Extra 64GB Actually Buys You
IFA 2026 filled the feeds with 192GB Gorgon Halo mini PCs. Capacity went up 50%; bandwidth went up about 7%. That asymmetry decides the whole purchase — here is the model-by-model list of what only 192GB runs, and why most buyers should still buy 128GB today.
DataHardware Team
Our Top Pick

GMKtec EVO-X2 (Ryzen AI Max+ 395, 128GB)
$3,649Quick answer: below roughly 120GB of resident weights, 192GB buys you headroom and nothing else. AMD's Ryzen AI Max 400 "Gorgon Halo" parts raise the unified-memory ceiling 50%, from 128GB to 192GB, but memory bandwidth rises only about 7% — 256 GB/s to 273 GB/s — on the same Zen 5 cores and the same 40-CU RDNA 3.5 graphics. Because local LLM decode speed is set by bandwidth and not capacity, the extra 64GB lets you load bigger models, not run them faster. The one thing it genuinely unlocks is the 235B–355B sparse MoE class at 4-bit — Qwen3-235B, GLM-class 355B — which needs 118–178GB of weights and simply does not fit in the ~96GB a 128GB box can hand its iGPU. If your workload is a 70B dense model (~35GB at 4-bit) or GPT-OSS 120B (~60GB), a 128GB box shipping today at $3,449–$3,799 does the identical job, months sooner — and the only 192GB price signal on record so far, a “$7…” placeholder on Minisforum’s own store page, points at roughly double that.
What actually changed with Gorgon Halo (and what didn't)
IFA 2026 in Berlin was the launch event, and the coverage was loud. Minisforum announced the MS-S1 Max-P495 — a Ryzen AI Max+ PRO 495 workstation with 192GB of LPDDR5X-8533 unified memory — alongside an AI Agent NAS N5, per Tom's Hardware and HotHardware. Acemagic, GMKtec and Framework all confirmed 495-based boxes reaching the same 192GB ceiling.
What almost none of that coverage said is how little else moved.
| Spec | Ryzen AI Max+ 395 (Strix Halo) | Ryzen AI Max+ PRO 495 (Gorgon Halo) | Change |
|---|---|---|---|
| CPU | 16C/32T Zen 5 | 16C/32T Zen 5 | Same cores, ~+100MHz to 5.2GHz (Tom's Hardware) |
| GPU | Radeon 8060S, 40 CU, RDNA 3.5 | Radeon 8065S, 40 CU, RDNA 3.5 (Wccftech) | Same topology, new badge |
| NPU | 50 TOPS (XDNA 2) | 55 TOPS (XDNA 2) (Minisforum, which rates the whole CPU+GPU+NPU platform at 131 TOPS) | Same NPU generation |
| Max unified memory | 128GB LPDDR5X-8000 | 192GB LPDDR5X-8533 | +50% |
| Memory bandwidth | 256 GB/s theoretical (~215 GB/s real) | 273 GB/s theoretical | +~7% |
| Systems ship | Shipping since 2025 | Minisforum says sales open September 2026 (announced, not verified) | — |
Read the table as a ratio and the story writes itself: capacity up 50%, bandwidth up 7%. Everything else in this post follows from that one asymmetry. If you want the timing question — should I delay a purchase for this? — we answered it separately in Ryzen AI Max 400: wait or buy a 128GB box now?. This post answers the capacity question: what does the extra 64GB actually do?
The trap: a 495 chip does not mean a 192GB machine
This is the single most expensive misreading available right now, and it is very easy to make. "Ryzen AI Max+ PRO 495" is a processor that supports up to 192GB. It does not mean the box in your cart has 192GB in it.
The proof is already on the record. BOSGAME previewed its flagship M5 MAX to Yahoo Finance — a Ryzen AI Max+ PRO 495 machine shipping late September to mid-October 2026 — and the first SKU is 128GB + 2TB at $3,600–$3,800. Same new chip. Same memory ceiling as a box you can buy today. Roughly the same price as today's 128GB boxes.
Buy the memory SKU, not the chip name. A 495 with 128GB is a 128GB machine with a marginally faster CPU. If the listing does not say 192GB in the configuration you are actually adding to the cart, you are not buying the thing the headlines were about.
Because the memory is soldered, this is not recoverable later. There is no SO-DIMM slot, no upgrade kit, no second-year expansion. You choose once.
The models 128GB can't hold — and what they cost you in speed
Here is the honest payoff list. Every weight figure below is derived arithmetic — parameter count × 0.5 bytes for 4-bit quantization — not a measured file size, and real GGUF quants land a few percent either side depending on which tensors are kept at higher precision.
| Model (4-bit) | Weights (derived) | Fits 128GB box (~96GB allocatable)? | Fits 192GB box? |
|---|---|---|---|
| Qwen3 32B dense | ~16GB | Yes, easily | Yes |
| Llama 3.3 70B dense | ~35GB | Yes | Yes |
| GPT-OSS 120B MoE (117B) | ~60GB | Yes | Yes |
| MiniMax ~230B MoE | ~115GB | No | Yes |
| Qwen3-235B-A22B MoE | ~118GB | No | Yes |
| GLM-class 355B-A32B MoE | ~178GB | No | No — exceeds the ~160GB allocatable ceiling |
| Llama 3.1 405B dense | ~203GB | No | No |
| DeepSeek R1 671B MoE | ~336GB | No | No |
That table contains two useful negative results that the launch coverage skipped entirely.
First: the real 128GB ceiling is not 128GB. On Strix Halo, only about 96GB of the 128GB is GPU-allocatable — the rest is reserved for the host OS. So the practical gap between the tiers is not 128 → 192; it is closer to ~96GB usable → ~160GB usable. That is a bigger jump than the headline implies, and it is why the 235B class lands squarely on the wrong side of the line rather than merely being a tight fit.
Second: 192GB does not reach the frontier either. Llama 3.1 405B dense (~203GB at 4-bit) and DeepSeek R1 671B (~336GB) still do not fit. If your reason for wanting 192GB is "so I can finally run R1," the answer is no — that has always needed a 512GB-class machine, and Apple discontinued the 256GB and 512GB Mac Studio configurations during the 2026 DRAM shortage. Nothing announced at IFA changes this.
For the complete picture of what already runs on a 128GB box today, we maintain a dedicated list: the best local LLM models for a 128GB mini PC.
What those models cost you in speed
Fitting is only half the question. Here is the part the press releases omit.
Every tokens-per-second figure in this section is arithmetic derived from published bandwidth numbers, not a benchmark. No 192GB Gorgon Halo system has been independently tested by anyone, including us. Treat every figure as an order-of-magnitude ceiling with a ≈ in front of it, and re-check it against measured results when hardware lands.
Token generation re-reads the weights it needs out of memory once per token emitted. Divide practical bandwidth by bytes-read-per-token and you get the theoretical decode ceiling. On this platform family, real bandwidth runs at roughly 80–85% of theoretical — the relationship behind the 256 GB/s → ~215 GB/s figure we use for the 395 throughout our measured Strix Halo tokens-per-second data. Apply the same ratio to 273 GB/s and Gorgon Halo lands near ≈230 GB/s practical.
| Model at 4-bit | Bytes read per token | ≈ Decode ceiling at ~230 GB/s | Verdict |
|---|---|---|---|
| Qwen3-235B-A22B (sparse, ~22B active) | ~11GB | ≈15–20 tok/s | Genuinely usable |
| GLM-class 355B-A32B (sparse, ~32B active) | ~16GB | ≈10–14 tok/s | Usable, slower |
| Hypothetical dense 200B | ~100GB — the whole model, every token | ≈2 tok/s | Loads. Practically unusable. |
This is the conclusion that reorganizes the whole purchase: 192GB is a mixture-of-experts feature, not a big-dense-model feature. The capacity unlocks a real and valuable class of model — but only the sparse class. Any dense model large enough to require 192GB is, by the same arithmetic, too slow on 273 GB/s to be worth running.
Why 50% more capacity doesn't mean 50% more speed
If you take one idea from this post, take this one, because it governs every buying decision in the category: unified memory capacity determines whether a model loads; memory bandwidth determines how fast it decodes. They are separate specs that happen to be printed next to each other on a box.
The GMKtec EVO-X2 at $3,649 is the clearest illustration in the catalog. Its 128GB of LPDDR5X-8000 will hold a dense 70B model without complaint. Its ~215 GB/s of real bandwidth means that model reads ~35GB per token and lands in single digits of tokens per second. Doubling the memory would not move that number by one token. Only a wider or faster bus would — and Gorgon Halo's bus is the same 256 bits, running 533 MT/s quicker.
Our full breakdown of Strix Halo memory bandwidth walks the arithmetic in detail, but the short version is a division problem: bandwidth ÷ bytes-per-token = tokens per second. Capacity does not appear in the equation at all.
This is also exactly why MoE models have taken over this hardware class. A sparse model activates only a fraction of its parameters per token, so a 235B-total model with 22B active reads roughly 11GB per token instead of 118GB. It occupies memory like a huge model and decodes like a mid-size one. Dense models get no such discount — every parameter is read, every token, forever. If you are new to the capacity-versus-bandwidth distinction, unified memory vs VRAM is the foundational read.
Worth noting where the bandwidth-first machines sit, for contrast: the Mac Studio M3 Ultra runs at 819 GB/s — roughly 3× any Gorgon Halo box — but Apple now sells it new in a 96GB configuration at $3,999 after the 256GB and 512GB tiers were discontinued. It is the fastest unified-memory desktop you can buy and simultaneously lower capacity than a $3,449 Framework Desktop. Capacity and bandwidth genuinely are independent axes, and the market currently forces you to pick one.
The GPU-allocatable trap, now at 192GB
Here is the section the launch-news reprints will not have.
Headline unified-memory capacity is not the number your inference runtime gets to use. On a 128GB Strix Halo box — the Minisforum MS-S1 Max at $3,799 is the direct 395-generation sibling of the newly announced P495 — the specification reads "128GB LPDDR5X-8000 (up to 96GB GPU-allocatable)." That parenthetical is doing enormous work. Roughly 32GB is reserved for the host, and how much of the remainder the iGPU can actually address depends on BIOS UMA settings, your kernel's GTT limits, and which runtime you use.
Extend that to the new tier and the arithmetic holds — and this time it is confirmed rather than inferred. Minisforum’s own MS-S1 MAX-P495 specification lists up to 160GB assignable to graphics out of 192GB total (Guru3D, Tom’s Hardware), exactly the ~32GB host reservation the 128GB boxes use. Which is why the numbers in the fit table above are tighter than they look:
- Qwen3-235B at ~118GB fits in 160GB with about 42GB left for KV cache and overhead — comfortable, and genuinely long context is on the table.
- GLM-class 355B at ~178GB does not fit in 160GB allocatable at all, despite fitting inside a "192GB" box on paper. You would be running partially on the CPU side or at a lower quant — and that is a very different machine from the one the headline sold you.
Anyone planning around a 192GB purchase should read our Strix Halo VRAM allocation guide first and assume the same class of constraint applies. The mechanism does not change with the memory ceiling; only the numbers do. Budget your model against the allocatable figure, never the box-label figure.
Price per usable GB: 128GB today vs 192GB in Q4
Cost-per-GB is the cleanest way to see whether the new tier is actually good value. Catalog prices below are verified retail as of September 2026; the 495 figure is an announced manufacturer number and is labelled as such.
| Box | Price | Memory | $/GB total | $/GB allocatable (~96GB) |
|---|---|---|---|---|
| HP Z2 Mini G1a | $3,300 – $3,734 | 128GB LPDDR5X-8533 ECC | ~$25.78 (low end) | ~$34.38 |
| Framework Desktop | $3,449 | 128GB LPDDR5X-8000 | ~$26.95 | ~$35.93 |
| GMKtec EVO-X2 | $3,649 | 128GB LPDDR5X-8000 | ~$28.51 | ~$38.01 |
| Minisforum MS-S1 Max | $3,799 | 128GB LPDDR5X-8000 | ~$29.68 | ~$39.57 |
| Beelink GTR9 Pro | $4,349 | 128GB LPDDR5X-8000 | ~$33.98 | ~$45.30 |
| NVIDIA DGX Spark | $4,699+ | 128GB LPDDR5X | ~$36.71 | — |
| BOSGAME M5 MAX (495) | $3,600 – $3,800 (announced, not verified) | 128GB + 2TB | ~$28.13 – $29.69 | — |
| Minisforum MS-S1 MAX-P495 (192GB) | Not final — page shows a “$7…” placeholder (Guru3D; ~€7,000 IFA teaser) | 192GB LPDDR5X-8533 | ~$36.46 (at $7,000) | ~$43.75 (160GB allocatable) |
Three things fall out of that table.
One: the new chip is not cheaper per GB. BOSGAME's 495 + 128GB at $3,600–$3,800 lands between the EVO-X2 and the MS-S1 Max on cost per gigabyte. You are paying today's price for a marginally faster CPU and the same memory you can buy this afternoon.
Two: no 192GB price is final, but the one public signal is much higher than the 128GB tier. Minisforum has not published an MSRP — its MS-S1 MAX-P495 store page still reads “Price Reveal Coming Soon.” What that page does show is a placeholder beginning with “$7” (Guru3D), and Minisforum’s own IFA teaser pointed at roughly €7,000. Neither is a confirmed number, and both should be treated as manufacturer signalling rather than a price. But they point at about twice the cost of a 128GB box, not a small premium over one — so do not budget the 192GB tier at $4,000. In the middle of a DRAM shortage, 50% more LPDDR5X per unit is not a rounding error; we tracked that pressure in why local-AI mini PC prices went up in 2026.
Three: the cheapest credible 128GB box is the value anchor. The Framework Desktop at $3,449 holds the lowest verified price per gigabyte in the group at roughly $26.95/GB, on a standard mini-ITX mainboard with open firmware. If your models fit in ~96GB allocatable — and per the fit table, nearly everything short of the 235B class does — that is the number every 192GB box has to beat — and the only 192GB figure anyone has signalled so far is roughly double it.
Which one should you actually buy?
Four decision paths. Pick the one whose first sentence describes you and stop reading the others.
Default — you run 70B dense or 120B-class MoE. Buy 128GB now.
This is roughly nine buyers in ten, and the honest advice is to stop waiting. Llama 3.3 70B at ~35GB and GPT-OSS 120B at ~60GB both sit well inside 96GB allocatable, and a 192GB machine would decode them at the same speed. The GMKtec EVO-X2 at $3,649 is the flagship pick; the Beelink GTR9 Pro at $4,349 costs more but adds dual 10GbE and dual USB4 if you are pulling models off a NAS or clustering two boxes. Cross-shop them directly: EVO-X2 vs GTR9 Pro.
Budget — you want the same capability for the least money.
The Framework Desktop at $3,449, on the strength of the $/GB table above. Standard mini-ITX board, open firmware, the best Linux and tinkerer story in the group. Direct-only — no Amazon. See EVO-X2 vs Framework Desktop for the trade-offs.
You genuinely need more than 120GB resident — and you will have to wait.
If Qwen3-235B or a GLM-class 355B is the actual job, no box in our catalog does it today. The M3 Ultra configurations that once could (256GB and 512GB) were discontinued during the DRAM shortage; Apple sells the Mac Studio M3 Ultra new at 96GB / $3,999, which is less capacity than a Strix Halo box even though it triples the bandwidth at 819 GB/s. That leaves waiting for a 192GB Gorgon Halo SKU — with the caveats stacked in this post: no final price and a placeholder pointing near $7,000, no verified ship date, no independent benchmark, and a 160GB allocatable ceiling that the 355B class still overruns. Track the Strix Halo hub and the Apple Silicon hub for when that changes.
CUDA-locked — your toolchain assumes NVIDIA.
The capacity argument is irrelevant to you, because your ceiling is 128GB either way. The NVIDIA DGX Spark at $4,699+ gives you 128GB at 273 GB/s — the same bandwidth class as Gorgon Halo — plus a CUDA-native stack and ConnectX-7 200GbE two-unit clustering that AMD has no answer to at this tier. It is the most expensive per gigabyte in the table and the only one where the software stack justifies it. See Mac Studio M3 Ultra vs DGX Spark and the GB10 hub.
The bottom line
Gorgon Halo raises the unified-memory ceiling from 128GB to 192GB without meaningfully raising memory bandwidth — 273 GB/s against 256 GB/s, a ~7% step beside a 50% capacity step. The extra 64GB lets you load bigger models, not run them faster, and the models it unlocks are specifically the 235B–355B sparse MoE class. Dense models large enough to need 192GB are too slow on this bandwidth to be worth running, and the genuine frontier models — 405B dense, R1 671B — still do not fit.
So: if you need more than ~120GB resident, 192GB is the only game in town, and you will wait for it at a price Minisforum has so far only hinted at with a “$7…” placeholder — roughly double a 128GB box. If you do not — and most people do not — buy a 128GB box today. For the tier-by-tier version of that decision at 64GB, 96GB and 128GB, start with how much unified memory you actually need for a local LLM. This post is its 192GB sequel, and the answer at the top of the ladder turns out to be the same as the answer in the middle: buy for the model you run, not the model in the headline.