AMD Ryzen AI Max+ 395 ("Strix Halo") Explained: The Chip Behind the 128GB Mini PCs
One APU put 128GB of unified memory and a 70B-capable iGPU on your desk for ~$2,000. Here's what Strix Halo actually is, what the numbers mean, and where it falls short.
DataHardware
Our Top Pick

GMKtec EVO-X2 (Ryzen AI Max+ 395, 128GB)
$3,399 – $3,499Quick answer: The AMD Ryzen AI Max+ 395 — codename "Strix Halo" — is a single APU that pairs a 16-core Zen 5 CPU, a 40-CU Radeon 8060S iGPU, and a 50-TOPS NPU with up to 128GB of LPDDR5X-8000 unified memory, of which up to 96GB is assignable as VRAM. That combination is why a ~$2,000 mini PC can now load 70B-class local models that no consumer GPU can hold. The trade-off is bandwidth: at 256 GB/s theoretical (~215 GB/s real) it's about 2× a normal APU but a quarter of a discrete GPU — so it's built to fit big models, not to race them.





What "Strix Halo" actually is
Strix Halo is AMD's bet that, for local AI, a big pool of unified memory beats a fast-but-small GPU. Instead of a CPU plus a separate graphics card with its own VRAM, it puts everything on one package and lets the CPU and GPU share the same 128GB:
- CPU: 16-core / 32-thread Zen 5
- GPU: Radeon 8060S — 40 compute units, RDNA 3.5, roughly RTX 4070-mobile class
- NPU: 50 TOPS (XDNA 2), for low-power on-device AI
- Memory: 128GB LPDDR5X-8000 unified, up to 96GB allocatable to the GPU
The official product name is "Ryzen AI Max+ 395"; "Strix Halo" is the codename everyone actually uses. A "PRO" variant (in the HP Z2 Mini) adds vPro manageability and ECC.





The number that matters: 96GB of allocatable memory
On a normal laptop APU, the iGPU borrows a slice of system RAM and you can't run anything serious on it. Strix Halo changes the math by letting you assign up to 96GB of the 128GB pool directly to the GPU. A 70B model quantized to 4-bit needs roughly 40–48GB plus context — comfortably inside that envelope. That's the entire pitch: capacity a 24GB or 32GB discrete card simply doesn't have, at a fraction of the price of stacking multiple GPUs.





The catch: bandwidth, not capacity
Fitting a model and running it fast are two different problems. Token-generation speed is gated by memory bandwidth, and here Strix Halo is mid-pack:
- Strix Halo (AI Max+ 395): 256 GB/s theoretical, ~215 GB/s real
- NVIDIA GB10 / DGX Spark: 273 GB/s
- Apple M4 Max: up to 546 GB/s · M3 Ultra: 819 GB/s
- Discrete GPU: 800–1,000 GB/s
So a dense 70B model runs at single-digit tokens/sec. Usable for chat, coding, and RAG; not for high-throughput serving. Two more permanent caveats: the memory is soldered (buy the 128GB SKU up front — there's no upgrade path), and you're on ROCm/Vulkan/llama.cpp, not CUDA, so GPU-compute beyond inference is still rougher than on NVIDIA.





Which boxes use it?
The silicon is identical across vendors — what differs is I/O, cooling, chassis, and price:
- Framework Desktop ($1,999) — cheapest credible 128GB box; standard mini-ITX, best tinkerer story. Direct-only.
- GMKtec EVO-X2 ($1,999–$2,199) — the flagship; quiet, dual-M.2, practical always-on appliance.
- Beelink GTR9 Pro ($1,899–$1,999) — dual 10GbE + dual USB4 for NAS/cluster networking.
- Minisforum MS-S1 Max (~$2,900) — enthusiast I/O: dual 10GbE, dual USB4 v2, PCIe x16, 2U-rack option.
- HP Z2 Mini G1a ($3,300+) — business-grade PRO silicon: vPro, ECC, 3-year warranty.





Should you buy into Strix Halo?
Yes if you want to run 70B-class models locally for the lowest realistic price (~$2,000), value capacity over speed, and can live in the ROCm/llama.cpp world. No if you need CUDA tooling (look at NVIDIA's GB10 boxes), if you need maximum token speed at 128GB (Apple M4 Max / M3 Ultra have far more bandwidth), or if your models are small enough (7–13B) that a $400–$600 budget mini PC would do. The chip is a genuine breakthrough for local-model capacity per dollar — just buy it for what it's good at.