Guide14 min read

Ryzen AI Max 400 "Gorgon Halo": Should You Wait, or Buy a 128GB Box Now?

AMD's next Halo chip raises unified memory 50% — to 192GB, with 160GB GPU-allocatable — but bandwidth only rises about 7%, to 273 GB/s. Divide one by the other and the "300B models locally" headline collapses. Here's the arithmetic, and the honest buy-or-wait verdict.

D

DataHardware Team

Our Top Pick

Beelink GTR9 Pro (Ryzen AI Max+ 395, 128GB)

Beelink GTR9 Pro (Ryzen AI Max+ 395, 128GB)

$1,899 – $1,999
AMD Ryzen AI Max+ 395 (16C/32T, Zen 5)Radeon 8060S (40 CU, RDNA 3.5)50 TOPS (XDNA 2)

Quick answer: for anything up to 120B-class, buy now — Gorgon Halo will not make your models faster. AMD's Ryzen AI Max 400 series raises unified memory 50%, from 128GB to 192GB, with up to 160GB GPU-allocatable (32GB reserved for the CPU) instead of today's ~96GB. But memory bandwidth rises only about 7%, from 256 GB/s to 273 GB/s, on the same Zen 5 CPU, the same 40-CU RDNA 3.5 graphics, and the same XDNA 2 NPU with roughly +100MHz of clock. Because local LLM decode speed is bound by bandwidth and not capacity, the extra 64GB buys you loadability, not speed: a 300B-class sparse MoE model becomes usable at roughly 15–20 tokens/sec, while a dense 300B sits near 1.5 tokens/sec. Systems land Q4 2026, price unannounced. If your workload is a 70B or GPT-OSS 120B, a 128GB box shipping today does the same job, four months sooner.

What AMD actually announced (the specs, without the spin)

The Ryzen AI Max 400 series — codename "Gorgon Halo" — is the successor to the Ryzen AI Max+ 395 "Strix Halo" that powers essentially every 128GB local-AI mini PC on the market. AMD detailed it at Advancing AI 2026, and the coverage that followed (ServeTheHome, Tom's Hardware, HotHardware, Notebookcheck) agrees on the shape of it.

Here is the generation-over-generation table, with the flagship parts side by side:

SpecRyzen AI Max+ 395 (Strix Halo)Ryzen AI Max+ (PRO) 495 (Gorgon Halo)Change
CPU16C/32T Zen 516C/32T Zen 5Same cores, ~+100MHz boost (reported up to 5.2GHz — Tom's Hardware)
GPURadeon 8060S, 40 CU, RDNA 3.540 CU, RDNA 3.5Unchanged topology
NPU50 TOPS (XDNA 2)XDNA 2Same NPU generation
Max unified memory128GB LPDDR5X-8000192GB LPDDR5X-8533+50% capacity
GPU-allocatable~96GBUp to 160GB (32GB reserved for CPU)+~67%
Memory bandwidth256 GB/s theoretical (~215 GB/s real)273 GB/s (AMD figure)+~7%
TDP class~55W class~55W classSame envelope
Systems shipShipping since 2025Q3–Q4 2026 (Oct–Dec per leak-sourced reporting)

The full SKU stack, per HotHardware and Notebookcheck's tables:

  • Ryzen AI Max+ (PRO) 495 — 16C/32T, 40 CU. The flagship, and the only part that matches the 395's compute width.
  • Ryzen AI Max (PRO) 490 / Max+ 492 — 12C/24T, 32 CU.
  • Ryzen AI Max (PRO) 485 / Max+ 488 — 8C/16T, 32 CU, Radeon 8050S graphics.

AMD's own headline claim is that this is the "world's first x86 client processor to run 300B+ LLMs" locally. That claim is true in a narrow, literal sense — and the rest of this post is about how narrow. Counterpoint Research's framing is the useful one for buyers: this is a platform refresh positioned around memory capacity and the ROCm software story, not a new compute architecture.

One structural note worth internalizing before you read the arithmetic: the memory on these parts is soldered. There is no 128GB box you upgrade to 192GB later. The capacity you buy on day one is the capacity you own for the life of the machine — which is precisely why "wait or buy" is a real question here rather than a shrug.

The number nobody is quoting: capacity went up 50%, bandwidth went up 7%

This is the section this post exists for, so the caveat goes first and in bold.

Every tokens-per-second figure in this section is back-of-envelope arithmetic derived from AMD's own published bandwidth number — not a benchmark. No Gorgon Halo silicon has been independently tested by anyone, including us. Treat every figure below as an order-of-magnitude estimate with a ≈ in front of it, and re-check it against real measurements when the hardware lands.

— DataHardware editorial standard

With that said, the arithmetic is not complicated, and it is the thing every press-release paraphrase skips.

Step one: convert the theoretical number to a practical one. On this platform family, real-world memory bandwidth runs at roughly 80–85% of the theoretical peak — that is the relationship behind the 256 GB/s → ~215 GB/s figure this site uses for the 395 throughout our measured Strix Halo tokens-per-second data. Apply the same ratio to 273 GB/s and you get roughly ≈230 GB/s practical on Gorgon Halo.

Step two: size the weights. A 300B-parameter model at FP4 quantization is approximately 150GB of weights. That fits inside 160GB of allocatable memory — with very little room for context, but it fits. AMD's claim survives this step intact.

Step three: divide. Token generation re-reads the weights it needs out of memory for every single token it emits. So:

Model type at 300B totalBytes read per token≈ Estimated decode speed at ~230 GB/sVerdict
Dense 300B (FP4)~150GB — the whole model, every token≈1.5 tok/sLoads. Practically unusable.
Sparse 300B MoE (~22–32B active)~11–16GB per token≈15–20 tok/sGenuinely usable.

Read those two rows against each other and the headline reorganizes itself. On Gorgon Halo, 192GB is a mixture-of-experts feature, not a big-dense-model feature. The capacity increase is real and it unlocks a real class of model — but only the sparse class. If you were reading "runs 300B+ LLMs" and picturing a dense 300B answering you at conversational speed, the bandwidth number says no, and it says no by an order of magnitude.

This is the same trap as NVIDIA's "1 PFLOP FP4" figure on the DGX Spark, which is the sparse number while dense compute is roughly half. The distinction between sparse and dense is the single most load-bearing spec-sheet asterisk in this entire product category, and it shows up on both sides of the AMD/NVIDIA fence.

Now hold the estimate against measured reality on hardware you can buy today. On a Ryzen AI Max+ 395 at ~215 GB/s, ServeTheHome measured GPT-OSS 120B at about 31 tok/s on a Beelink GTR9 Pro at ~120W, and community testing puts Qwen3-30B-A3B in the 70–100 tok/s band. Those are measurements, not estimates. A Gorgon Halo box running those exact same models would read the same weights through a pipe roughly 7% wider. You would not be able to feel the difference in a chat window.

That is the whole verdict compressed into one sentence: the models most people actually run in 2026 are bandwidth-bound, and Gorgon Halo does not meaningfully move bandwidth. If you want the mechanism explained properly rather than asserted, our pillar on unified memory vs VRAM for local AI works through why capacity and speed are separate axes that people persistently collapse into one.

Who should actually wait (be specific, and short)

There is a real population for whom Gorgon Halo is the right machine. It is smaller than the headlines imply, and it is worth naming precisely — because the credibility of "buy now" depends on being honest about who it does not apply to.

Wait if you want a 300B-class sparse MoE model and you accept ~15–20 tok/s. This is the legitimate case, and it is not a small thing: no Strix Halo box shipping today can load that model at all. 96GB of allocatable memory is a hard wall, and 160GB is on the other side of it. If your target is a frontier open-weights MoE and reading speed is acceptable to you, the wait is justified by capability you cannot otherwise buy at this price tier.

Wait if you need more than 96GB allocatable for very long context on a 70B. A dense Llama 3.3 70B at 4-bit is roughly 40–48GB of weights, which leaves the KV cache fighting for what remains of your 96GB ceiling as context grows. If you are running genuinely long contexts — big RAG retrievals, whole repositories — the allocatable headroom on Gorgon is a real, non-cosmetic improvement, even though the decode speed will not change. Our Strix Halo VRAM allocation guide covers how to squeeze the current ceiling first; do that before you conclude you need a new chip.

Wait if you weren't buying until Q4 anyway. Obvious, but worth saying plainly. If your budget or your project timeline lands in Q4 2026 regardless, there is no cost to seeing what launches. Waiting is only expensive when you would otherwise have been running something.

Everyone else: no. And there are two pieces of missing information that should make you especially wary of waiting on price hope:

  • The 192GB SKU price is unannounced — needs verification. AMD has not published one, and neither have the partners. Anyone quoting you a number is guessing. The only adjacent datapoint is that the current-generation Ryzen AI Halo dev platform lists at $3,999, and the second-generation platform (Max+ PRO 495, 192GB, 2TB PCIe 4.0) has no announced price at all.
  • 50% more DRAM in a DRAM shortage does not get cheaper. The memory in these boxes is the single most expensive component and it is the component whose price has been climbing all year — the mechanics of which we documented in why local-AI mini PC prices jumped in 2026. A 192GB box arriving into that market being cheaper than today's 128GB boxes would be an extraordinary outcome, not a default expectation.

Worth drawing the distinction between those two posts explicitly, because they answer the same-sounding question on different axes: that post asks "will the price fall?" — this one asks "will the silicon be better?" You can get a "yes" on one and a "no" on the other, and most readers need both answers before they click buy.

What waiting actually costs you

The cost of waiting is usually treated as zero, which is how four-month delays turn into fourteen-month delays. Price it honestly.

Four months of not owning the box. If you are buying an always-on inference machine as a working tool — the solo-dev and small-studio case — that is a third of a year of not having the thing. Against what? Against a refresh that, by AMD's own bandwidth number, changes nothing measurable for any workload at or below 120B. You are trading real months for a ~7% number you will not perceive.

The counter-argument that is actually good. There is one, and it deserves to be stated at full strength rather than strawmanned: a Gorgon Halo launch may discount existing 128GB Ryzen AI Max+ 395 stock. That is the normal pattern when a successor arrives, and given how far current pricing has moved above launch MSRP — the GMKtec EVO-X2 now sits at $3,399–$3,499 against a pre-sale launch price near $1,999 — there is room for that to matter.

But notice what that argument actually recommends. It is an argument for waiting to buy Strix Halo cheaper, not an argument for buying Gorgon Halo. Those are different decisions with different risk profiles. And the counter-counter-argument is the DRAM shortage again: discounting old stock requires the replacement stock to be plentiful and the component cost to be falling, and in 2026 neither condition has held. The historical pattern this year has been prices rising, not falling.

The honest framing: if you are speculating on a post-launch Strix Halo discount, you are making a market bet, not a technology bet — and you should size it accordingly. If you need the machine, the machine you can order today runs your models at the same speed as the one you would order in December.

The generation that actually matters is Medusa Halo (2027)

Zoom out one generation and the buying cadence gets clearer than any single launch.

Bandwidth is the binding constraint on every box in this class. Not capacity, not TOPS, not CU count — bandwidth. It is why a dense 70B runs at ~5 tok/s on a 395 while a 30B MoE runs at 70–100 tok/s on the same silicon. It is why NVIDIA's GB10 at 273 GB/s does not embarrass a Strix Halo box at 215 GB/s despite costing more than twice as much, a comparison we work through in DGX Spark vs Strix Halo. And it is why Apple's Mac Studio M3 Ultra at 819 GB/s is a genuinely different class of machine for dense models, even though it holds less memory in its current 96GB-only configuration.

Against that backdrop:

  • Gorgon Halo (Q4 2026) adds ~7% bandwidth. That is a refresh.
  • Medusa Halo (2027 at the earliest) is reported to roughly double it. That is a generational break.

So the cadence writes itself, and it is a three-line buying policy:

"Buy now for capacity. Skip Gorgon. Re-evaluate at Medusa. Gorgon Halo is a capacity refresh, not a speed upgrade — the generation that doubles bandwidth is Medusa Halo in 2027."

— DataHardware hardware desk

The one thing that policy is not is a recommendation to wait until 2027. A machine you own and use for eighteen months returns more than a spec sheet you refresh a browser tab on. The point of naming Medusa is to tell you when the next real upgrade is, so that a box bought today has a defined replacement horizon instead of an anxious one.

If you're buying today, which 128GB box

All five Strix Halo boxes we track run the identical Ryzen AI Max+ 395 silicon at the same ~215 GB/s. Peak tokens/sec is effectively identical across them — what differs is I/O, cooling, chassis, and price. Match the box to the buyer, and verify every price at checkout, because this category re-prices weekly and our seeds are date-stamped.

BoxCatalog priceBest forThe caveat
Beelink GTR9 Pro$1,899 – $1,999Best price-per-capacity, plus dual 10GbE for NAS/cluster workReported 10GbE NIC instability under heavy GPU load — driver-dependent
Framework Desktop$1,999Cheapest credible 128GB entry; standard mini-ITX board, best Linux storyDirect-only — not on Amazon
GMKtec EVO-X2$3,399 – $3,499Flagship / best-supported; quiet, dual-M.2, a practical always-on applianceRepriced hard during the DRAM spike; 2.5GbE only
Minisforum MS-S1 Max$2,879 – $3,039Strongest I/O — dual 10GbE, dual USB4 v2, PCIe x16, 2U-rack optionLarger chassis; support reputation weaker than HP
HP Z2 Mini G1a$3,300 – $3,734Business-grade — PRO silicon, vPro, ECC, 3-year warranty, ~2.5L SFFMost expensive per spec; 2.5GbE only

If CUDA is non-negotiable, the GB10 platform is the alternative rather than a competitor: the NVIDIA DGX Spark at $4,699+ or the cheaper ASUS Ascent GX10 at $2,999 – $4,100. Both hold 128GB at 273 GB/s — note that this is the same bandwidth Gorgon Halo will ship with, which tells you something about how little 273 GB/s changes on its own. Head-to-head: DGX Spark vs EVO-X2.

If bandwidth beats capacity for you, look at Apple silicon instead. The Mac Studio M3 Ultra runs at 819 GB/s — roughly 3× any Halo or GB10 box — and at launch its 512GB configuration ran DeepSeek R1 671B at 4-bit entirely in memory. The 2026 caveat is severe and you must price it in: the 256GB and 512GB configurations were pulled during the DRAM shortage, so it sells in 96GB only right now at $3,999. The Mac Studio M4 Max ($1,999 – $5,999) reaches 128GB at 410–546 GB/s and is the more practical Apple pick. Compare: M3 Ultra vs DGX Spark.

Two more comparisons if you are narrowing between Strix Halo boxes specifically: EVO-X2 vs GTR9 Pro and EVO-X2 vs Framework Desktop. And if you are earlier in the funnel than "which 128GB box," start from the best mini PC for local LLMs in 2026, which sizes boxes by model and budget rather than assuming you need this tier at all.

The proof point: you can already do the thing

One last argument, and it is the most persuasive one because it involves no estimates at all.

The workload most people describe when they say "I want to run a big model locally" is a 120B-class MoE. That is already solved, today, on hardware you can order this afternoon. GPT-OSS 120B runs at ~31 tok/s at ~120W on a 395 box per ServeTheHome's measurement, and StorageReview independently ran it on the HP Z2 Mini G1a with no discrete GPU. We wrote the full procedure up in how to run GPT-OSS 120B locally on a mini PC.

Step down a tier and it gets more comfortable, not less: a Qwen3 32B-class model or a 30B MoE sits in the 70–100 tok/s band — faster than you read. Our roundup of the best local LLM models for a 128GB mini PC maps the whole field.

So the question "should I wait for Gorgon Halo?" has a concrete test attached to it, and you can run it in thirty seconds: name the model you actually intend to run. If it is 120B or smaller, a box shipping today runs it at the speed a Gorgon Halo box would, and waiting buys you nothing but four months of not having it. If it is a 300B-class MoE and you have made peace with ~15–20 tok/s, wait — you are the person this chip is for.

Bottom line

  • 192GB and 160GB allocatable are real — a genuine +50% capacity and +~67% allocatable over the Ryzen AI Max+ 395. That part of the headline holds up.
  • 273 GB/s is a ~7% bandwidth gain on the same Zen 5 / RDNA 3.5 40-CU / XDNA 2 silicon with ~+100MHz. That is a refresh, not a new architecture.
  • Capacity ≠ speed. Derived from AMD's own bandwidth figure: dense 300B ≈ 1.5 tok/s, sparse 300B MoE ≈ 15–20 tok/s. Estimates, not benchmarks — no Gorgon silicon has been independently tested.
  • 192GB is a MoE feature. If your model is dense, the extra memory changes what loads, not what runs.
  • For ≤120B workloads, buy now. GPT-OSS 120B at ~31 tok/s and 30B MoE at 70–100 tok/s are measured, today, on ~$1,899–$2,000 hardware.
  • The 192GB price is unannounced — needs verification, and 50% more DRAM in a shortage is unlikely to arrive cheap.
  • The real generational break is Medusa Halo in 2027, which is reported to roughly double bandwidth. Buy now for capacity, skip Gorgon, re-evaluate then.

AMD's "world's first x86 client processor to run 300B+ LLMs" is a defensible marketing claim and a misleading buying signal at the same time. Both things are true, and the gap between them is one division problem: 230 GB/s ÷ 150GB of dense weights ≈ 1.5 tokens per second. Run that division before you postpone a purchase over it.

ryzen-ai-max-400gorgon-halostrix-haloamd-ai-max-395medusa-halounified-memorymemory-bandwidthmoebuying-guide
Beelink GTR9 Pro (Ryzen AI Max+ 395, 128GB)

Beelink GTR9 Pro (Ryzen AI Max+ 395, 128GB)

$1,899 – $1,999

Check Price

More from the blog

Stay ahead in AI hardware

Weekly deals, GPU reviews, and build guides. No spam.

Unsubscribe anytime. We respect your inbox.