Unified Memory

Unified memory is a single pool of RAM shared by the CPU, GPU, and NPU, so the graphics processor reads model weights directly instead of copying them across a PCIe bus into separate VRAM. It's what lets a box like the GMKtec EVO-X2 expose 128GB of LPDDR5X to the GPU — up to 96GB assignable as VRAM — and load 70B-class models no 24–32GB discrete card can hold. Apple pioneered the approach; the M3 Ultra scales the same idea to 819 GB/s. The trade-off is bandwidth: unified LPDDR5X runs ~215–273 GB/s on the AMD and NVIDIA boxes versus 800–1000 GB/s on a discrete GPU, so you gain capacity but not raw speed.

When buying, unified memory is the reason these mini PCs exist: capacity you can't get on a consumer GPU at any sane price. Decide whether your models fit in memory first, then check bandwidth — capacity gets the model loaded, bandwidth decides how fast it answers.

Related Products

Related Articles

More Terms