Reference

Local LLM tokens/sec Benchmark Reference

Measured single-stream speeds for running LLMs on mini PCs and unified-memory boxes — Strix Halo, DGX Spark, and two-node clusters. Every number links to the page it was read from. Where no verified benchmark exists, we say so instead of estimating.

Last reviewed 2026-09-22 · Next review 2026-12 · Reviewed quarterly

How to read this: Generation is text-generation tok/s at batch size 1 — the speed you feel when a model replies. Prompt is prompt-processing (prefill) tok/s, which sets how long you wait before the first token. The Run column is the exact measurement spec the source used: pp512 / tg128 means a 512-token prompt and 128 generated tokens. Numbers are only comparable within the same model × quantization × backend × run spec — a Vulkan number at 512 tokens and a ROCm number at 8192 are different machines running different code on different work.

Qwen3 30B-A3B

HardwareGenerationPromptQuantBackendRunConfidenceSource
Strix Halo (Ryzen AI Max+ 395, 128GB) — box not statedSecond capture of the Level1Techs config above, same author. 72.0 and 75.32 bracket the measured 30B MoE figure — no published run approaches 100.75.32 tok/sUD-Q4_K_XLllama.cpp (Vulkan)pp512 / tg128Community-measuredllm-tracker.info, AMD Strix Halo GPU performance2026-09-22
Framework Desktop (Ryzen AI Max+ 395, 128GB)MoE — ~3B of 30B parameters active per token.Check price$3,44972 tok/s604.8 tok/sUD-Q4_K_XLllama.cpp (Vulkan)pp512 / tg128Community-measuredLevel1Techs forums, Strix Halo LLM benchmark results (lhl)2025-07-22

Llama 2 7B

HardwareGenerationPromptQuantBackendRunConfidenceSource
Strix Halo (Ryzen AI Max+ 395, 128GB) — box not stated52.73 tok/s884.2 tok/sQ4_0llama.cpp (Vulkan + FA)pp512 / tg128Community-measuredllm-tracker.info, AMD Strix Halo GPU performance2026-09-22
Strix Halo (Ryzen AI Max+ 395, 128GB) — box not statedAt 8K context. Generation holds; this is the config that survives long prompts.50.97 tok/s368.77 tok/sQ4_0llama.cpp (HIP + WMMA + FA)pp8192 / tg8192Community-measuredllm-tracker.info, AMD Strix Halo GPU performance2026-09-22
Strix Halo (Ryzen AI Max+ 395, 128GB) — box not stated50.88 tok/s343.91 tok/sQ4_0llama.cpp (HIP + WMMA + FA)pp512 / tg128Community-measuredllm-tracker.info, AMD Strix Halo GPU performance2026-09-22
Framework Desktop (Ryzen AI Max+ 395, 128GB)Check price$3,44947.9 tok/s906.1 tok/sQ4_K_Mllama.cpp (HIP/ROCm)pp512 / tg128Community-measuredLevel1Techs forums, Strix Halo LLM benchmark results (lhl)2025-07-22
Framework Desktop (Ryzen AI Max+ 395, 128GB)Check price$3,44945.8 tok/s998 tok/sQ4_0llama.cpp (Vulkan)pp512 / tg128Community-measuredLevel1Techs forums, Strix Halo LLM benchmark results (lhl)2025-07-22
Strix Halo (Ryzen AI Max+ 395, 128GB) — box not statedNo iGPU offload — what the same chip does on CPU alone.28.94 tok/s294.64 tok/sQ4_0llama.cpp (CPU only)pp512 / tg128Community-measuredllm-tracker.info, AMD Strix Halo GPU performance2026-09-22
Strix Halo (Ryzen AI Max+ 395, 128GB) — box not statedAt 8K context without Flash Attention generation collapses from 52.2 to 7.54 tok/s. Same hardware, same model — a build-flag result, not a silicon one.7.54 tok/s487.69 tok/sQ4_0llama.cpp (Vulkan, no FA)pp8192 / tg8192Community-measuredllm-tracker.info, AMD Strix Halo GPU performance2026-09-22

Shisa V2 8B

HardwareGenerationPromptQuantBackendRunConfidenceSource
Framework Desktop (Ryzen AI Max+ 395, 128GB)Check price$3,44942 tok/s878.2 tok/si1-Q4_K_Mllama.cpp (HIP/ROCm)pp512 / tg128Community-measuredLevel1Techs forums, Strix Halo LLM benchmark results (lhl)2025-07-22

gpt-oss 120B

HardwareGenerationPromptQuantBackendRunConfidenceSource
NVIDIA DGX Spark (GB10)Prompt processing is ~5x the Strix Halo row below on the same model; generation is within ~13%.Check price$4,699+38.55 tok/s1723.07 tok/sMXFP4llama.cpp (CUDA)pp2048 / tg32Independent-measuredhardware-corner.net, First DGX Spark LLM benchmarks2025-10-15
Strix Halo (Ryzen AI Max+ 395, 128GB) — box not stated34.13 tok/s339.87 tok/sMXFP4llama.cpp (ROCm/Vulkan)pp2048 / tg32Independent-measuredhardware-corner.net, First DGX Spark LLM benchmarks2025-10-15
Beelink GTR9 Pro (Ryzen AI Max+ 395, 128GB)Reviewer's unoptimized run. The llama.cpp rows above are the tuned comparison.Check price$4,34931.41 tok/snot statedLM StudioLM Studio, out-of-the-boxIndependent-measuredServeTheHome, Beelink GTR9 Pro review2026-09-22

dots1

HardwareGenerationPromptQuantBackendRunConfidenceSource
Framework Desktop (Ryzen AI Max+ 395, 128GB)Check price$3,44920.6 tok/s63.1 tok/sUD-Q4_K_XLllama.cpp (Vulkan)pp512 / tg128Community-measuredLevel1Techs forums, Strix Halo LLM benchmark results (lhl)2025-07-22

Llama 4 Scout

HardwareGenerationPromptQuantBackendRunConfidenceSource
Framework Desktop (Ryzen AI Max+ 395, 128GB)Check price$3,44919.3 tok/s264.1 tok/sUD-Q4_K_XLllama.cpp (HIP/ROCm)pp512 / tg128Community-measuredLevel1Techs forums, Strix Halo LLM benchmark results (lhl)2025-07-22

Hunyuan-A13B

HardwareGenerationPromptQuantBackendRunConfidenceSource
Framework Desktop (Ryzen AI Max+ 395, 128GB)Check price$3,44917.1 tok/s270.5 tok/sUD-Q6_K_XLllama.cpp (Vulkan)pp512 / tg128Community-measuredLevel1Techs forums, Strix Halo LLM benchmark results (lhl)2025-07-22

Qwen3-235B-A22B

HardwareGenerationPromptQuantBackendRunConfidenceSource
2x Sapphire Edge AI Max+ 395 (256GB pooled)Model exceeds a single box's ~96GB GPU allocation. Clustering buys capacity, not bandwidth.10–15 tok/sQ3_K_S (~101GB)llama.cpp RPC (two nodes)not statedIndependent-measuredStarry Hope, Linked Strix Halo mini PCs for 235B LLM inference2026-04-01

Mistral Small 3.1

HardwareGenerationPromptQuantBackendRunConfidenceSource
Framework Desktop (Ryzen AI Max+ 395, 128GB)Dense 24B — the architecture penalty in one row against Qwen3 30B-A3B above.Check price$3,44914.3 tok/s316.9 tok/sUD-Q4_K_XLllama.cpp (HIP/ROCm)pp512 / tg128Community-measuredLevel1Techs forums, Strix Halo LLM benchmark results (lhl)2025-07-22

Shisa V2 70B

HardwareGenerationPromptQuantBackendRunConfidenceSource
Framework Desktop (Ryzen AI Max+ 395, 128GB)Dense 70B — the bandwidth floor for this class of box.Check price$3,4495 tok/s94.7 tok/si1-Q4_K_Mllama.cpp (HIP/ROCm)pp512 / tg128Community-measuredLevel1Techs forums, Strix Halo LLM benchmark results (lhl)2025-07-22

Llama 3.3 70B

HardwareGenerationPromptQuantBackendRunConfidenceSource
Beelink GTR9 Pro (Ryzen AI Max+ 395, 128GB)Source states "around 5 tokens/second" — an approximate figure, not a precise measurement.Check price$4,3495 tok/snot statednot statedout-of-the-boxIndependent-measuredServeTheHome, Beelink GTR9 Pro review2026-09-22

Known gaps — no number we'd stand behind

These segments have no verified benchmark we can cite. We list them explicitly rather than fill the cell with an estimate — including where that means contradicting numbers in our own older posts.

SegmentStatusWhy
Budget mini PCs (GMKtec M6 Ultra, GMKtec M8, Beelink SER8, MAGICNUC AS1)No verified public benchmarkWe sell these as 7–13B model hosts and there is no published single-stream tok/s benchmark for any of these exact SKUs. The 7–8B figures in the table above were measured on Strix Halo at ~215 GB/s; these boxes are slower and narrower, so quoting those numbers for them would be an extrapolation, not a measurement.
Dense 70B on DGX Spark / GB10No verified public benchmarkThe DGX Spark comparison above covers gpt-oss 120B only. No published single-stream dense-70B figure exists for GB10 on the same methodology, so the Strix Halo dense-70B number has no like-for-like counterpart. Source
Sustained tok/s under thermal load (any Strix Halo chassis)No verified public benchmarkEvery figure on this page is a short burst. Cooling and power delivery differ sharply across the EVO-X2, GTR9 Pro, MS-S1 Max, Framework Desktop and Z2 Mini, but no source publishes a sustained-load tok/s curve — so 'which box stays fastest' has no verified answer, only a plausible one.
Dense model at 235B across two nodesPublished figure is a projection, not a measurementStarry Hope's 3–5 tok/s dense figure is explicitly a projection of what a dense model of that size would do, not a run they performed. It is excluded from the table above for that reason. Source

Want the story behind these numbers? How fast is Strix Halo, really explains why a 30B MoE model beats a dense 70B on the same box, and the prefill problem covers the prompt-processing column. Need to know what will fit before you worry about speed? Use the VRAM calculator.

Disclosure: Some links on this page are affiliate links. We may earn a commission if you make a purchase — at no extra cost to you. Benchmark source links are never affiliate links.