Topic Hub

Strix Halo Mini PCs for Local AI

Strix Halo is AMD's Ryzen AI Max+ 395: a 16-core Zen 5 APU with a Radeon 8060S iGPU and 128GB of LPDDR5X unified memory, up to 96GB of it allocatable as VRAM. That's enough to hold 70B-class models no 24–32GB consumer GPU can fit. The trade-off is bandwidth — real-world ~215–273 GB/s depending on the box — so dense 70B runs at single-digit tokens/sec while MoE models fly. Buy these for capacity, not raw speed.

Top Picks

GMKtec EVO-X2 (Ryzen AI Max+ 395, 128GB)

GMKtec EVO-X2 (Ryzen AI Max+ 395, 128GB)

$3,399 – $3,499

  • APU: AMD Ryzen AI Max+ 395 (16C/32T, Zen 5)
  • GPU: Radeon 8060S (40 CU, RDNA 3.5)
  • NPU: 50 TOPS (XDNA 2)
Check Price on Amazon
Beelink GTR9 Pro (Ryzen AI Max+ 395, 128GB)

Beelink GTR9 Pro (Ryzen AI Max+ 395, 128GB)

$1,899 – $1,999

  • APU: AMD Ryzen AI Max+ 395 (16C/32T, Zen 5)
  • GPU: Radeon 8060S (40 CU, RDNA 3.5)
  • NPU: 50 TOPS (XDNA 2)
Check Price on Amazon
Minisforum MS-S1 Max (Ryzen AI Max+ 395, 128GB)

Minisforum MS-S1 Max (Ryzen AI Max+ 395, 128GB)

$2,879 – $3,039

  • APU: AMD Ryzen AI Max+ 395 (16C/32T, Zen 5)
  • GPU: Radeon 8060S (40 CU, RDNA 3.5)
  • NPU: 50 TOPS (XDNA 2)
Check Price on Amazon
HP Z2 Mini G1a (Ryzen AI Max+ PRO 395, 128GB)

HP Z2 Mini G1a (Ryzen AI Max+ PRO 395, 128GB)

$3,300 – $3,734

  • APU: AMD Ryzen AI Max+ PRO 395 (16C/32T, Zen 5, vPro)
  • GPU: Radeon 8060S (40 CU, RDNA 3.5)
  • NPU: 50 TOPS (XDNA 2)
Check Price on Amazon
Framework Desktop (Ryzen AI Max+ 395, 128GB)

Framework Desktop (Ryzen AI Max+ 395, 128GB)

$1,999

  • APU: AMD Ryzen AI Max+ 395 (16C/32T, Zen 5)
  • GPU: Radeon 8060S (40 CU, RDNA 3.5)
  • NPU: 50 TOPS (XDNA 2)
Check Price on Direct

Comparisons

Models That Run on These Boxes

Llama

Llama 3.3 70B Instruct

Meta's flagship dense chat model — strong general reasoning, coding, and instruction-following that rivals much larger models. The default pick when a single big box has the memory to spare.

Llama

Llama 4 Scout

A mixture-of-experts model — only a fraction of its 109B total parameters are active per token, so it loads like a large model but runs far faster than its size suggests. A natural fit for unified-memory boxes with a long context window.

Qwen3

Qwen3 8B

A capable small chat-and-reasoning model that fits comfortably even on budget boxes — a good everyday assistant when you don't need frontier-level depth.

Qwen3

Qwen3 14B

The mid-size Qwen3 — noticeably stronger reasoning and coding than the 8B while still fitting mid-range hardware at Q4.

Qwen3

Qwen3 32B

Qwen3's large dense model — near-flagship quality that still runs on a single unified-memory box, a strong local alternative to 70B-class models at lower memory cost.

DeepSeek-R1

DeepSeek-R1 Distill Qwen 7B

R1's chain-of-thought reasoning distilled into a 7B Qwen base — the lightest way to get R1-style step-by-step reasoning locally, small enough for budget boxes.

DeepSeek-R1

DeepSeek-R1 Distill Qwen 14B

The 14B R1 distill — a good balance of reasoning depth and hardware cost, running comfortably on mid-range unified-memory boxes.

DeepSeek-R1

DeepSeek-R1 Distill Qwen 32B

The strongest Qwen-based R1 distill — heavy reasoning that still fits a single large-memory box, the sweet spot for local R1-style work.

DeepSeek-R1

DeepSeek-R1 Distill Llama 70B

R1's reasoning distilled onto a Llama 70B base — the highest-quality distill, and the one to reach for when the full 671B R1 won't fit (it never does locally).

Gemma 3

Gemma 3 4B

Google's smallest current Gemma — tiny memory footprint and multimodal, ideal for the most memory-constrained budget boxes and always-on assistants.

Gemma 3

Gemma 3 12B

The mid-size Gemma 3 — solid general-purpose quality with a long context window, fitting mid-range hardware at Q4.

Gemma 3

Gemma 3 27B

The largest Gemma 3 — frontier-adjacent quality for a dense open model, comfortably runnable on a single large-memory box.

GPT-OSS

GPT-OSS 20B

OpenAI's small open-weight MoE — its 21B total parameters load like a mid-size model, but because only a fraction are active per token it runs fast on modest hardware.

GPT-OSS

GPT-OSS 120B

OpenAI's large open-weight MoE — its 117B total parameters fit a single high-memory box, and because only a fraction are active per token it runs far quicker than its total suggests.

Phi

Phi-4

Microsoft's 14B model punches well above its size on math and reasoning — a compact, mid-range-friendly pick when you want strong reasoning without a big memory bill.

Mistral

Mistral Small 3.2 24B

Mistral's 24B dense model — fast, capable, and low-latency for its class, sitting neatly between the mid-size and large boxes.

Related Articles

Guide

Can You Cluster Two Mini PCs to Run Bigger Local LLMs? (2026 Reality Check)

AMD published a four-node Framework Desktop cluster running a trillion-parameter model, and two DGX Sparks pool 256GB over a single cable. Here's when a second box actually helps, when it makes things slower, and which mini PC to buy if clustering is on your roadmap.

Read
Comparison

Mac Studio M4 Max vs Strix Halo: Which 128GB Box Actually Runs Your Local LLM Faster?

Tom's Hardware measured the M4 Max at roughly 1.6× a Strix Halo box on tokens/sec. But a 128GB Strix Halo machine starts at $1,999. Here's the bandwidth math, the per-GB math, and the prefill caveat that decides which one you should actually buy.

Read
Guide

Why Local-AI Mini PC Prices Doubled in 2026 — And Whether to Buy Now or Wait

The GMKtec EVO-X2 went from ~$1,999 to ~$3,399. NVIDIA repriced the DGX Spark from $3,999 to $4,699 on 2026-02-27, citing memory supply. Here's the SKU-by-SKU price damage, why DRAM did it, and why waiting is the wrong call — with TrendForce's own 2027 forecast as the evidence.

Read
Guide

Beelink GTR9 Pro Review (2026): The 128GB Strix Halo Box With Dual 10GbE — Real Local-LLM Benchmarks & Who Should Buy

The Beelink GTR9 Pro is a ~$1,899–$1,999 128GB Ryzen AI Max+ 395 mini PC that runs GPT-OSS 120B at ~31 tok/s (~120W) but a dense 70B at only ~5 tok/s. Its real hook is dual 10GbE — plus a real caveat: reported 10GbE instability under heavy GPU load. Benchmarks, thermals, and how it stacks up against the EVO-X2, Framework Desktop, and DGX Spark.

Read
Economics

How Much Does It Cost to Run a Local AI Mini PC 24/7? Idle Power, Electricity & True TCO in 2026

A local-AI box left on 24/7 costs roughly $15–$60 a year in electricity — and the spread is almost all idle power, because a personal LLM server sits idle 95%+ of the time. A Strix Halo box idles ~13W (~$20/yr); a DGX Spark idles ~37–40W (~$55/yr). For an always-on box, buy for idle watts, not peak specs.

Read
Guide

Why Your Local LLM Feels Slow on a Strix Halo Mini PC: The Prefill (Time-to-First-Token) Problem

On a Ryzen AI Max+ 395 box, token generation ties a DGX Spark — but prompt processing (prefill) is ~5× slower (~340 vs ~1,700 tok/s on gpt-oss 120B). On long prompts, that's the delay that makes RAG and coding agents feel sluggish. Here's who it hits, why, and how to fix it.

Read
Guide

Best Local LLM Models to Run on a 128GB Mini PC in 2026 (Matched to Your Box)

You bought (or are eyeing) a 128GB unified-memory box. Here's what to actually load on it: gpt-oss 120B at ~31 tok/s, Qwen3-30B at ~100 tok/s, dense Llama 3.3 70B at ~4–6 tok/s — ranked by job, with the exact SKU for each.

Read
Guide

Unified Memory vs VRAM for Local AI: Why Capacity Now Beats Bandwidth (2026)

VRAM is fast but capped at ~32GB; unified memory trades bandwidth for 128GB+ of capacity. Here's the exact trade-off, why Mixture-of-Experts models flipped the math, and which box to buy for the model you actually run.

Read
Tutorial

How to Unlock the Full 128GB as VRAM on a Ryzen AI Max+ 395 (Strix Halo): BIOS + Linux GTT Guide

You bought a 128GB Strix Halo box and the GPU only sees ~16GB. Here's the fix: Windows caps GPU-allocatable memory at 96GB via the BIOS UMA frame buffer, but Linux with the amdttm GTT kernel params reaches ~110–120GB — and the winning move is counterintuitive.

Read

Frequently Asked Questions

What is Strix Halo and why does it matter for local AI?

Strix Halo is AMD's Ryzen AI Max+ 395 — a 16-core Zen 5 APU paired with a Radeon 8060S iGPU and up to 128GB of LPDDR5X unified memory, up to 96GB of which is allocatable as VRAM. That capacity lets a single small box load 70B-class models that no 24–32GB consumer GPU can hold.

How fast is a Strix Halo box for 70B models?

Memory bandwidth is the ceiling. Real-world bandwidth runs roughly 215–273 GB/s depending on the box, so a dense 70B model runs at single-digit tokens per second. Mixture-of-experts (MoE) models activate only a fraction of their weights per token and run far faster.

Which Strix Halo box should I buy?

The GMKtec EVO-X2 is the flagship. The Beelink GTR9 Pro adds dual 10GbE for clustering. The Minisforum MS-S1 Max has the best IO and cooling. The HP Z2 Mini G1a is the premium, business-grade pick with vPro and ECC. The Framework Desktop is the cheapest credible 128GB box at $1,999.

Can I upgrade the memory later?

No. The 128GB LPDDR5X is soldered on every Strix Halo box, so there is no upgrade path — buy the 128GB configuration up front.