Type something to search...
A grey Minisforum mini PC on a dark wooden desk beside a handheld game console, a wireless mouse and a camera control dial

Unified Memory PC or Gaming GPU for Local AI? Who Should Buy Which

A 128GB unified-memory box fits AI models a 32GB RTX 5090 can't, but plays games like an RTX 4060 laptop. Who should buy Strix Halo, DGX Spark, a Mac or a GPU.

Eren Smith10 Sep, 202615 min read
A grey Minisforum mini PC on a dark wooden desk beside a handheld game console, a wireless mouse and a camera control dial

Buy a unified-memory machine if the AI model you want to run is bigger than 32GB. Buy a gaming PC with a discrete GPU if games come first and your models fit in its VRAM. That one number decides most of this.

A 128GB unified-memory box such as AMD's Ryzen AI Max+ 395 "Strix Halo" mini PCs or Nvidia's DGX Spark can load a model that no consumer graphics card can hold. It pays for that with memory bandwidth — the number that sets how fast an AI model writes and how well a game runs. These machines game at roughly the level of an RTX 4060 laptop, on a PC that costs as much as a high-end gaming rig.

Note

Key Takeaways

  • Unified memory buys capacity: 128GB on Strix Halo and DGX Spark, up to 512GB on Apple's new M5 Ultra Mac Studio. The biggest consumer GPU, the RTX 5090, has 32GB.
  • A discrete GPU buys bandwidth: 1,792 GB/s on the RTX 5090 against 256 GB/s on Strix Halo and 273 GB/s on DGX Spark. Speed on models that fit follows that gap.
  • For games, Strix Halo is the only one of these that runs Windows x86 games natively. Its Radeon 8060S averaged 67 FPS in Cyberpunk 2077 at 1080p high in Tom's Hardware's report.
  • DGX Spark runs Linux on Arm and needs an emulation layer for games. Macs can't run kernel-level anti-cheat titles at all.
  • 128GB Strix Halo mini PCs sold for $2,399 to $3,810 in March 2026. The whole category has climbed with memory prices.

What Unified Memory Changes for Local AI

Unified memory lets the GPU use most of the system's RAM as its own, so a 128GB machine can hold a model that would overflow any 32GB graphics card. On a normal PC the GPU only sees its own VRAM. On an APU the CPU and GPU share one pool, and the driver decides how much of it the GPU gets.

AMD's own Ryzen AI Halo box ships with an app that lets you move that CPU and GPU split graphically, according to Tom's Hardware's review. The same idea is covered, for games, in our guide to how much VRAM to allocate to integrated graphics. The difference here is scale: the pool is 128GB rather than 16GB or 32GB.

Capacity decides what loads, bandwidth decides how fast it runs

A large language model has to read its weights from memory for every token it writes. That makes generation speed mostly a memory bandwidth problem, not a compute problem.

You can get a rough ceiling by dividing bandwidth by the size of the weights read per token. A 70-billion-parameter dense model at Q4_K_M, about 4 bits per weight, takes about 43GB: that's the size of Ollama's llama3.3:70b download. Here is the arithmetic, which is a theoretical upper bound and not a benchmark:

MachineMemoryBandwidthCeiling on a 43GB dense model
Ryzen AI Max+ 395 (Strix Halo)128GB256 GB/sabout 6 tokens/s
Ryzen AI Max+ Pro 495 (Gorgon Halo)up to 192GBabout 273 GB/sabout 6 tokens/s
Nvidia DGX Spark (GB10)128GB273 GB/sabout 6 tokens/s
Nvidia RTX Spark (top config)128GBup to 300 GB/sabout 7 tokens/s
Mac Studio M5 Maxup to 128GB614 GB/sabout 14 tokens/s
Mac Studio M5 Ultraup to 512GB1.2 TB/sabout 28 tokens/s
RTX 509032GB1,792 GB/sdoes not fit

Real results land below those numbers. Sometimes well below. The table still explains the market. The four PC platforms sit within about 20% of each other on bandwidth, Apple sits well above them, and the RTX 5090 is several times faster on anything that fits in 32GB.

Why mixture-of-experts models suit unified memory

Mixture-of-experts models only read the experts they activate for each token, not the whole model, so they write far faster than a dense model of the same size. Two of the three models Tom's Hardware ran on the Ryzen AI Halo were MoE designs, Qwen 3.6-35B-A3B and gpt-oss-120B. The third was the small dense Gemma 4 12B.

MoE models are the sweet spot for unified memory: large enough to need the capacity, light enough per token that 256 GB/s feels usable. A large dense model is where these machines slow to single digits. For how model size maps to memory in the first place, see our sibling explainer on how much VRAM local AI models need.

The Machines on Sale or Announced in September 2026

Four unified-memory platforms matter this month: Strix Halo PCs and Nvidia's DGX Spark are on sale, Apple's M5 Mac Studio ships September 22, and Nvidia's RTX Spark arrives in October. They differ more in software and operating system than in memory size.

Strix Halo mini PCs and laptops

The Ryzen AI Max+ 395 pairs a 16-core, 32-thread Zen 5 CPU with a Radeon 8060S GPU and up to 128GB of LPDDR5X-8000 at 256 GB/s. Liliputing's March 2026 roundup listed nine 128GB mini PCs from $2,399 (Bosgame M5) to $3,810 (MINIX ER939-AI). AMD's own Ryzen AI Halo, which ships with Windows 11 Pro or Linux, sells for $3,999 with 128GB and a 2TB SSD. Framework's 128GB Desktop was $3,449 in July.

Gorgon Halo, the Ryzen AI Max 400 refresh

AMD's Ryzen AI Max Pro 400 series, codenamed Gorgon Halo, is a minor update to the same Zen 5 and RDNA 3.5 design. Tom's Hardware lists the flagship Max+ Pro 495 with a 100 MHz higher boost clock, a 40-CU Radeon 8065S, and support for up to 192GB of memory, 160GB of it usable by the GPU. It is a commercial "Pro" part so far. Lenovo's ThinkCentre X Ultra uses it with up to 128GB of LPDDR5X-8533 from November at an expected price of about $3,700, and Framework has previewed a 192GB Desktop with no price or date. Faster 8533 MT/s memory adds about 7% bandwidth over the 395, which doesn't change the picture below.

Nvidia DGX Spark

A 20-core Arm CPU and a Blackwell GPU share 128GB of LPDDR5X at 273 GB/s. Nvidia rates it at up to 1 petaflop of FP4 AI compute and says it can run inference on models up to 200 billion parameters. It runs DGX OS, Nvidia's Ubuntu-based Linux. Tom's Hardware put its price at $4,699 when the $3,999 Ryzen AI Halo launched.

Mac Studio with M5 Max or M5 Ultra

Apple announced both on August 25 and they ship September 22. The M5 Max has up to 128GB at 614 GB/s from $2,499. The M5 Ultra has up to 512GB at 1.2 TB/s from $5,499, and the 512GB configuration follows in late October. Starting prices are for base memory, so a 128GB configuration costs more.

Nvidia RTX Spark, arriving in October

RTX Spark is a Windows on Arm platform for laptops and small desktops, with up to 20 CPU cores, 6,144 CUDA cores, 128GB of LPDDR5X and up to 300 GB/s. Tom's Hardware reports that it launches in October in two configurations, with laptops from Dell, HP, Lenovo, ASUS, MSI and Microsoft. No prices or independent tests exist yet.

How Each One Handles Games

Only Strix Halo plays Windows PC games natively and well, averaging 67 FPS in Cyberpunk 2077 at 1080p high, while DGX Spark needs emulation and a Mac can't run kernel-level anti-cheat games. The GPU isn't the only thing that decides this. The operating system and the CPU instruction set matter as much.

Strix Halo: a real 1080p gaming PC

The Radeon 8060S is the best gaming option in this group, because it runs x86 Windows games on the normal AMD driver with no translation layer. In a test by YouTuber RandomGaminginHD on a Minisforum MS-S1 Max, reported by Tom's Hardware, it averaged these results at 1080p:

GameSettingsAverage FPS
Counter-Strike 2High, MSAA off263.7
Kingdom Come: Deliverance 2High89.8
Battlefield 6High, FSR Native AA86.7
Red Dead Redemption 2Mixed high, ultra textures82.2
Cyberpunk 2077High preset66.8
Cyberpunk 2077RT Ultra, FSR 3 Quality45.4
Borderlands 4Lowest54.8

That is roughly RTX 4060 laptop territory. TechSpot's roundup of early laptop results had the 8060S at 39 FPS in Cyberpunk 2077 against 36 FPS for an RTX 4060 laptop and 37 FPS for an RTX 4070 laptop, while Nvidia's parts were up to 37% faster in GTA V. Plan on 1080p with upscaling, not 1440p ultra. For what ordinary laptop and desktop iGPUs manage by comparison, see our list of games that genuinely run on integrated graphics.

DGX Spark: it can play, with a translation layer in the way

DGX Spark's GPU has the same 6,144 CUDA cores as a desktop RTX 5070, but games still struggle on it. It runs Arm Linux, so Windows games go through Proton plus an x86 emulator, and its 273 GB/s is well under the RTX 5070's 672 GB/s.

A Reddit user's first test, reported by Tom's Hardware, got about 50 FPS in Cyberpunk 2077 at 1080p medium using Box64. An Nvidia representative replied that switching to FEX and Proton 10.2-2 beta produced more than 175 FPS at 1080p high with Ultra ray tracing, but that figure uses DLSS 4 Multi Frame Generation. Generated frames raise the counter without improving responsiveness the same way, as our breakdown of frame generation and input lag explains. The same Redditor called stability "actually pretty good" but still hit occasional crashes.

Mac Studio: fast hardware, missing games

The M5 Max and M5 Ultra have far more bandwidth than any PC APU, but macOS limits what you can play. Native Mac ports are a small slice of the PC library, and the rest runs through Apple's Game Porting Toolkit or CodeWeavers' CrossOver.

CrossOver 26, released in early 2026, got several anti-cheat-protected games running, including Helldivers 2, according to WINE for Mac's coverage. Games whose anti-cheat runs at the Windows kernel level, such as Valorant's Vanguard, still don't run on macOS. Riot lists Valorant for PC, Xbox Series X|S and PlayStation 5 only. If your group plays competitive shooters, a Mac isn't your gaming machine — however much memory it has.

RTX Spark: promising, unproven

Nvidia claims RTX Spark is good for "100 FPS 1440p gaming", with DLSS 4.5 and Multi Frame Generation likely doing much of that work. It also runs Windows on Arm, so most games will go through Microsoft's Prism x86 emulation. Treat the gaming claim as a vendor claim until independent reviews arrive after the October launch.

Why a Gaming GPU Still Wins for Most People

A discrete graphics card is faster at every AI model that fits in its VRAM and at every game, which covers most of what gamers actually run. The RTX 5090 has 1,792 GB/s of bandwidth, about seven times Strix Halo's 256 GB/s. Nvidia's $1,999 launch MSRP for the card is also below the $2,399 to $4,699 list prices of the 128GB boxes covered here, though street prices for any of them can differ.

Macro photograph of a GPU die on its package ringed by the graphics card's GDDR memory chips

The catch is the 32GB ceiling. Nothing gets around it. A 16GB card such as the RTX 5070 Ti handles models up to about 20B parameters at 4-bit comfortably. The RTX 5090 stretches that into the 30B range. Past about 32GB of weights, the rest of the model spills into system RAM, and speed falls off hard. The slow part is the RAM itself, about 96 GB/s for dual-channel DDR5-6000, not the PCIe slot.

The gaming PC also brings things unified-memory boxes can't. You can upgrade the GPU in two years. You get CUDA on Windows with no emulation. And none of the game compatibility problems above apply. Our best budget GPU for 2026 comparison covers the cheaper entry points, and how much VRAM you need for gaming covers the gaming side of the capacity question.

The trade is power. Nvidia rates the RTX 5090 alone at 575W of total graphics power and lists 1,000W of required system power. DGX Spark ships with a 240W brick, and AMD's Ryzen AI Halo has a 120W chip and a 240W adapter. If a box will run an AI agent around the clock, that difference shows up on the bill, as our guide to the cost of running a gaming PC works through.

Who Should Buy Which

Choose by the largest model you actually intend to run and by whether you need Windows games, not by the AI TOPS figure on the box. Most gamers curious about local AI will be happier with a GPU. A minority with a specific large model in mind won't be.

If this describes youBuyWhy
Games come first, and your models are 30B parameters or smaller at 4-bitA gaming PC with a 16GB to 32GB GPUFastest at every model that fits, fastest image generation, CUDA on Windows, and anti-cheat games run
You want 70B-class or large MoE models and Windows games at 1080p on one machineA Strix Halo mini PC or laptopThe only 128GB option that games natively, though ROCm is less mature than CUDA
AI development is the job and gaming barely mattersDGX SparkNvidia's CUDA stack plus 200 Gbps networking that links up to four units, which Nvidia says handles models up to 700 billion parameters
You need more bandwidth than any PC APU, or more than 128GBMac Studio M5 Max or M5 Ultra614 GB/s or 1.2 TB/s and up to 512GB, with games left to another device
You want a thin 128GB laptopNothing yetRTX Spark's game compatibility is unproven until reviews land after the October launch

Our position is blunter than most buying guides: if you can't name the model that needs more than 32GB, you don't need unified memory. That's the whole test.

This table is only as good as September 2026, though. Nobody outside Nvidia has published RTX Spark game tests, and AMD hasn't said whether consumer Gorgon Halo chips are coming, so two rows could look different by the end of the year.

If you are weighing an NPU laptop instead, note that the NPU is not what runs these large models. Our sibling piece on whether an NPU AI PC helps gaming covers what that chip actually does.

The Costs the Spec Sheet Leaves Out

A unified-memory box also costs you in slow long-prompt responses, less mature software, fan noise and memory you can never upgrade. Each shows up in reviews of shipping hardware, not just on paper.

Start with latency. Tom's Hardware found the Ryzen AI Halo's generation speed acceptable but slower than a Dell Pro Max GB10. The bigger gap was time to first token, which "can quickly rise to non-interactive levels with long contexts." At the extremes, the reviewer waited two to four minutes for a response to start. That matters for coding assistants, where the context grows all session.

Software is the second gap. DGX Spark ships with Nvidia's own stack and playbooks. AMD answered with preloaded ROCm, Lemonade and its own playbooks on the Ryzen AI Halo, though the review found the preinstalled vLLM failed to load Qwen 3.6-35B-A3B and some paths weren't documented.

Then noise. The same review describes the AI Halo's twin blowers as audible at idle, with "a notable high-pitched whine" under load, and found it louder than the Dell GB10 box during image generation. Our look at whether mini PC power supplies are good enough covers the external-brick side of that design.

Finally, memory prices. Framework's 128GB Desktop cost $2,851 in March 2026 and $3,449 by July. The GEEKOM A9 Mega went from a planned $2,099 to $3,199. Unified memory is soldered LPDDR5X, so you are buying all of it up front and can never add more later. That also makes the 32GB-versus-64GB question from our is 32GB of RAM overkill guide irrelevant here: you pick your ceiling at checkout.

Frequently Asked Questions

Can a Strix Halo mini PC replace a gaming PC?

For 1080p gaming, largely yes, because the Radeon 8060S performs close to an RTX 4060 laptop GPU. It averaged 67 FPS in Cyberpunk 2077 at 1080p high and 87 FPS in Battlefield 6 in Tom's Hardware's report. It won't match a desktop RTX 5070 or better at 1440p, and you can't upgrade the GPU later.

Is a unified-memory PC faster than an RTX 5090 for local AI?

Only when the model doesn't fit in 32GB. For anything that fits, the RTX 5090's 1,792 GB/s beats Strix Halo's 256 GB/s and DGX Spark's 273 GB/s by a wide margin. The unified-memory box wins by being able to load 70B-class dense or large mixture-of-experts models at all.

Can the Nvidia DGX Spark play games?

Technically, yes, but it isn't built for it. It runs Arm Linux, so games need Proton and an x86 emulator such as FEX. Early tests got about 50 FPS in Cyberpunk 2077 at 1080p medium, and Nvidia's own 175 FPS figure relies on DLSS 4 Multi Frame Generation.

Is a Mac Studio good for gaming?

Its hardware is strong, with 614 GB/s on the M5 Max, but game support is the limit. Native Mac ports are few, compatibility layers such as CrossOver run many Windows games, and titles with kernel-level anti-cheat such as Valorant do not run on macOS at all.

How much unified memory do I need for local AI?

Match it to your largest model plus room for context. A 70B dense model at Q4_K_M needs about 43GB for weights alone, so 64GB is tight and 128GB is comfortable. Nvidia rates its 128GB DGX Spark for inference on models up to 200 billion parameters. Past that you're looking at Gorgon Halo's 192GB or the M5 Ultra's 512GB. If your models stay at 30B or below, a 16GB to 32GB graphics card is the faster and cheaper option.

The Bottom Line

A unified-memory box is a capacity purchase: it loads AI models no graphics card can hold, at a fraction of a high-end GPU's speed. If you don't have a specific model that needs more than 32GB, you are paying $2,400 to $4,700 for memory you won't use — and giving up a lot of gaming performance to do it.

For a gamer with a real large-model need, Strix Halo is the only option here that is also a competent Windows gaming PC. DGX Spark suits AI developers who happen to game occasionally. The M5 Mac Studio wins on bandwidth and loses on games. Everyone else should put the money into a discrete GPU with as much VRAM as the budget allows. For how local AI is starting to show up inside games, see our coverage of AI NPCs in games.

Hardware photography courtesy of the respective manufacturers and publications, used for editorial coverage.

Sources

Share