Reference
Local LLM tokens/sec Benchmark Reference
Measured single-stream speeds for running LLMs on mini PCs and unified-memory boxes — Strix Halo, DGX Spark, and two-node clusters. Every number links to the page it was read from. Where no verified benchmark exists, we say so instead of estimating.
Last reviewed 2026-09-22 · Next review 2026-12 · Reviewed quarterly
pp512 / tg128 means a 512-token prompt and 128 generated tokens. Numbers are only comparable within the same model × quantization × backend × run spec — a Vulkan number at 512 tokens and a ROCm number at 8192 are different machines running different code on different work.Qwen3 30B-A3B
| Hardware | Generation | Prompt | Quant | Backend | Run | Confidence | Source |
|---|---|---|---|---|---|---|---|
| Strix Halo (Ryzen AI Max+ 395, 128GB) — box not statedSecond capture of the Level1Techs config above, same author. 72.0 and 75.32 bracket the measured 30B MoE figure — no published run approaches 100. | 75.32 tok/s | — | UD-Q4_K_XL | llama.cpp (Vulkan) | pp512 / tg128 | Community-measured | llm-tracker.info, AMD Strix Halo GPU performance2026-09-22 |
| Framework Desktop (Ryzen AI Max+ 395, 128GB)MoE — ~3B of 30B parameters active per token.Check price$3,449 | 72 tok/s | 604.8 tok/s | UD-Q4_K_XL | llama.cpp (Vulkan) | pp512 / tg128 | Community-measured | Level1Techs forums, Strix Halo LLM benchmark results (lhl)2025-07-22 |
Llama 2 7B
| Hardware | Generation | Prompt | Quant | Backend | Run | Confidence | Source |
|---|---|---|---|---|---|---|---|
| Strix Halo (Ryzen AI Max+ 395, 128GB) — box not stated | 52.73 tok/s | 884.2 tok/s | Q4_0 | llama.cpp (Vulkan + FA) | pp512 / tg128 | Community-measured | llm-tracker.info, AMD Strix Halo GPU performance2026-09-22 |
| Strix Halo (Ryzen AI Max+ 395, 128GB) — box not statedAt 8K context. Generation holds; this is the config that survives long prompts. | 50.97 tok/s | 368.77 tok/s | Q4_0 | llama.cpp (HIP + WMMA + FA) | pp8192 / tg8192 | Community-measured | llm-tracker.info, AMD Strix Halo GPU performance2026-09-22 |
| Strix Halo (Ryzen AI Max+ 395, 128GB) — box not stated | 50.88 tok/s | 343.91 tok/s | Q4_0 | llama.cpp (HIP + WMMA + FA) | pp512 / tg128 | Community-measured | llm-tracker.info, AMD Strix Halo GPU performance2026-09-22 |
| Framework Desktop (Ryzen AI Max+ 395, 128GB)Check price$3,449 | 47.9 tok/s | 906.1 tok/s | Q4_K_M | llama.cpp (HIP/ROCm) | pp512 / tg128 | Community-measured | Level1Techs forums, Strix Halo LLM benchmark results (lhl)2025-07-22 |
| Framework Desktop (Ryzen AI Max+ 395, 128GB)Check price$3,449 | 45.8 tok/s | 998 tok/s | Q4_0 | llama.cpp (Vulkan) | pp512 / tg128 | Community-measured | Level1Techs forums, Strix Halo LLM benchmark results (lhl)2025-07-22 |
| Strix Halo (Ryzen AI Max+ 395, 128GB) — box not statedNo iGPU offload — what the same chip does on CPU alone. | 28.94 tok/s | 294.64 tok/s | Q4_0 | llama.cpp (CPU only) | pp512 / tg128 | Community-measured | llm-tracker.info, AMD Strix Halo GPU performance2026-09-22 |
| Strix Halo (Ryzen AI Max+ 395, 128GB) — box not statedAt 8K context without Flash Attention generation collapses from 52.2 to 7.54 tok/s. Same hardware, same model — a build-flag result, not a silicon one. | 7.54 tok/s | 487.69 tok/s | Q4_0 | llama.cpp (Vulkan, no FA) | pp8192 / tg8192 | Community-measured | llm-tracker.info, AMD Strix Halo GPU performance2026-09-22 |
Shisa V2 8B
| Hardware | Generation | Prompt | Quant | Backend | Run | Confidence | Source |
|---|---|---|---|---|---|---|---|
| Framework Desktop (Ryzen AI Max+ 395, 128GB)Check price$3,449 | 42 tok/s | 878.2 tok/s | i1-Q4_K_M | llama.cpp (HIP/ROCm) | pp512 / tg128 | Community-measured | Level1Techs forums, Strix Halo LLM benchmark results (lhl)2025-07-22 |
gpt-oss 120B
| Hardware | Generation | Prompt | Quant | Backend | Run | Confidence | Source |
|---|---|---|---|---|---|---|---|
| NVIDIA DGX Spark (GB10)Prompt processing is ~5x the Strix Halo row below on the same model; generation is within ~13%.Check price$4,699+ | 38.55 tok/s | 1723.07 tok/s | MXFP4 | llama.cpp (CUDA) | pp2048 / tg32 | Independent-measured | hardware-corner.net, First DGX Spark LLM benchmarks2025-10-15 |
| Strix Halo (Ryzen AI Max+ 395, 128GB) — box not stated | 34.13 tok/s | 339.87 tok/s | MXFP4 | llama.cpp (ROCm/Vulkan) | pp2048 / tg32 | Independent-measured | hardware-corner.net, First DGX Spark LLM benchmarks2025-10-15 |
| Beelink GTR9 Pro (Ryzen AI Max+ 395, 128GB)Reviewer's unoptimized run. The llama.cpp rows above are the tuned comparison.Check price$4,349 | 31.41 tok/s | — | not stated | LM Studio | LM Studio, out-of-the-box | Independent-measured | ServeTheHome, Beelink GTR9 Pro review2026-09-22 |
dots1
| Hardware | Generation | Prompt | Quant | Backend | Run | Confidence | Source |
|---|---|---|---|---|---|---|---|
| Framework Desktop (Ryzen AI Max+ 395, 128GB)Check price$3,449 | 20.6 tok/s | 63.1 tok/s | UD-Q4_K_XL | llama.cpp (Vulkan) | pp512 / tg128 | Community-measured | Level1Techs forums, Strix Halo LLM benchmark results (lhl)2025-07-22 |
Llama 4 Scout
| Hardware | Generation | Prompt | Quant | Backend | Run | Confidence | Source |
|---|---|---|---|---|---|---|---|
| Framework Desktop (Ryzen AI Max+ 395, 128GB)Check price$3,449 | 19.3 tok/s | 264.1 tok/s | UD-Q4_K_XL | llama.cpp (HIP/ROCm) | pp512 / tg128 | Community-measured | Level1Techs forums, Strix Halo LLM benchmark results (lhl)2025-07-22 |
Hunyuan-A13B
| Hardware | Generation | Prompt | Quant | Backend | Run | Confidence | Source |
|---|---|---|---|---|---|---|---|
| Framework Desktop (Ryzen AI Max+ 395, 128GB)Check price$3,449 | 17.1 tok/s | 270.5 tok/s | UD-Q6_K_XL | llama.cpp (Vulkan) | pp512 / tg128 | Community-measured | Level1Techs forums, Strix Halo LLM benchmark results (lhl)2025-07-22 |
Qwen3-235B-A22B
| Hardware | Generation | Prompt | Quant | Backend | Run | Confidence | Source |
|---|---|---|---|---|---|---|---|
| 2x Sapphire Edge AI Max+ 395 (256GB pooled)Model exceeds a single box's ~96GB GPU allocation. Clustering buys capacity, not bandwidth. | 10–15 tok/s | — | Q3_K_S (~101GB) | llama.cpp RPC (two nodes) | not stated | Independent-measured | Starry Hope, Linked Strix Halo mini PCs for 235B LLM inference2026-04-01 |
Mistral Small 3.1
| Hardware | Generation | Prompt | Quant | Backend | Run | Confidence | Source |
|---|---|---|---|---|---|---|---|
| Framework Desktop (Ryzen AI Max+ 395, 128GB)Dense 24B — the architecture penalty in one row against Qwen3 30B-A3B above.Check price$3,449 | 14.3 tok/s | 316.9 tok/s | UD-Q4_K_XL | llama.cpp (HIP/ROCm) | pp512 / tg128 | Community-measured | Level1Techs forums, Strix Halo LLM benchmark results (lhl)2025-07-22 |
Shisa V2 70B
| Hardware | Generation | Prompt | Quant | Backend | Run | Confidence | Source |
|---|---|---|---|---|---|---|---|
| Framework Desktop (Ryzen AI Max+ 395, 128GB)Dense 70B — the bandwidth floor for this class of box.Check price$3,449 | 5 tok/s | 94.7 tok/s | i1-Q4_K_M | llama.cpp (HIP/ROCm) | pp512 / tg128 | Community-measured | Level1Techs forums, Strix Halo LLM benchmark results (lhl)2025-07-22 |
Llama 3.3 70B
| Hardware | Generation | Prompt | Quant | Backend | Run | Confidence | Source |
|---|---|---|---|---|---|---|---|
| Beelink GTR9 Pro (Ryzen AI Max+ 395, 128GB)Source states "around 5 tokens/second" — an approximate figure, not a precise measurement.Check price$4,349 | 5 tok/s | — | not stated | not stated | out-of-the-box | Independent-measured | ServeTheHome, Beelink GTR9 Pro review2026-09-22 |
Known gaps — no number we'd stand behind
These segments have no verified benchmark we can cite. We list them explicitly rather than fill the cell with an estimate — including where that means contradicting numbers in our own older posts.
| Segment | Status | Why |
|---|---|---|
| Budget mini PCs (GMKtec M6 Ultra, GMKtec M8, Beelink SER8, MAGICNUC AS1) | No verified public benchmark | We sell these as 7–13B model hosts and there is no published single-stream tok/s benchmark for any of these exact SKUs. The 7–8B figures in the table above were measured on Strix Halo at ~215 GB/s; these boxes are slower and narrower, so quoting those numbers for them would be an extrapolation, not a measurement. |
| Dense 70B on DGX Spark / GB10 | No verified public benchmark | The DGX Spark comparison above covers gpt-oss 120B only. No published single-stream dense-70B figure exists for GB10 on the same methodology, so the Strix Halo dense-70B number has no like-for-like counterpart. Source |
| Sustained tok/s under thermal load (any Strix Halo chassis) | No verified public benchmark | Every figure on this page is a short burst. Cooling and power delivery differ sharply across the EVO-X2, GTR9 Pro, MS-S1 Max, Framework Desktop and Z2 Mini, but no source publishes a sustained-load tok/s curve — so 'which box stays fastest' has no verified answer, only a plausible one. |
| Dense model at 235B across two nodes | Published figure is a projection, not a measurement | Starry Hope's 3–5 tok/s dense figure is explicitly a projection of what a dense model of that size would do, not a run they performed. It is excluded from the table above for that reason. Source |
Want the story behind these numbers? How fast is Strix Halo, really explains why a 30B MoE model beats a dense 70B on the same box, and the prefill problem covers the prompt-processing column. Need to know what will fit before you worry about speed? Use the VRAM calculator.
Disclosure: Some links on this page are affiliate links. We may earn a commission if you make a purchase — at no extra cost to you. Benchmark source links are never affiliate links.