Buy a unified-memory machine if the AI model you want to run is bigger than 32GB. Buy a gaming PC with a discrete GPU if games come first and your models fit in its VRAM. That one number decides most of this.
A 128GB unified-memory box such as AMD's Ryzen AI Max+ 395 "Strix Halo" mini PCs or Nvidia's DGX Spark can load a model that no consumer graphics card can hold. It pays for that with memory bandwidth — the number that sets how fast an AI model writes and how well a game runs. These machines game at roughly the level of an RTX 4060 laptop, on a PC that costs as much as a high-end gaming rig.
Note
Key Takeaways
- Unified memory buys capacity: 128GB on Strix Halo and DGX Spark, up to 512GB on Apple's new M5 Ultra Mac Studio. The biggest consumer GPU, the RTX 5090, has 32GB.
- A discrete GPU buys bandwidth: 1,792 GB/s on the RTX 5090 against 256 GB/s on Strix Halo and 273 GB/s on DGX Spark. Speed on models that fit follows that gap.
- For games, Strix Halo is the only one of these that runs Windows x86 games natively. Its Radeon 8060S averaged 67 FPS in Cyberpunk 2077 at 1080p high in Tom's Hardware's report.
- DGX Spark runs Linux on Arm and needs an emulation layer for games. Macs can't run kernel-level anti-cheat titles at all.
- 128GB Strix Halo mini PCs sold for $2,399 to $3,810 in March 2026. The whole category has climbed with memory prices.
What Unified Memory Changes for Local AI
Unified memory lets the GPU use most of the system's RAM as its own, so a 128GB machine can hold a model that would overflow any 32GB graphics card. On a normal PC the GPU only sees its own VRAM. On an APU the CPU and GPU share one pool, and the driver decides how much of it the GPU gets.
AMD's own Ryzen AI Halo box ships with an app that lets you move that CPU and GPU split graphically, according to Tom's Hardware's review. The same idea is covered, for games, in our guide to how much VRAM to allocate to integrated graphics. The difference here is scale: the pool is 128GB rather than 16GB or 32GB.
Capacity decides what loads, bandwidth decides how fast it runs
A large language model has to read its weights from memory for every token it writes. That makes generation speed mostly a memory bandwidth problem, not a compute problem.
You can get a rough ceiling by dividing bandwidth by the size of the weights read per token. A 70-billion-parameter dense model at Q4_K_M, about 4 bits per weight, takes about 43GB: that's the size of Ollama's llama3.3:70b download. Here is the arithmetic, which is a theoretical upper bound and not a benchmark:
| Machine | Memory | Bandwidth | Ceiling on a 43GB dense model |
|---|---|---|---|
| Ryzen AI Max+ 395 (Strix Halo) | 128GB | 256 GB/s | about 6 tokens/s |
| Ryzen AI Max+ Pro 495 (Gorgon Halo) | up to 192GB | about 273 GB/s | about 6 tokens/s |
| Nvidia DGX Spark (GB10) | 128GB | 273 GB/s | about 6 tokens/s |
| Nvidia RTX Spark (top config) | 128GB | up to 300 GB/s | about 7 tokens/s |
| Mac Studio M5 Max | up to 128GB | 614 GB/s | about 14 tokens/s |
| Mac Studio M5 Ultra | up to 512GB | 1.2 TB/s | about 28 tokens/s |
| RTX 5090 | 32GB | 1,792 GB/s | does not fit |
Real results land below those numbers. Sometimes well below. The table still explains the market. The four PC platforms sit within about 20% of each other on bandwidth, Apple sits well above them, and the RTX 5090 is several times faster on anything that fits in 32GB.
Why mixture-of-experts models suit unified memory
Mixture-of-experts models only read the experts they activate for each token, not the whole model, so they write far faster than a dense model of the same size. Two of the three models Tom's Hardware ran on the Ryzen AI Halo were MoE designs, Qwen 3.6-35B-A3B and gpt-oss-120B. The third was the small dense Gemma 4 12B.
MoE models are the sweet spot for unified memory: large enough to need the capacity, light enough per token that 256 GB/s feels usable. A large dense model is where these machines slow to single digits. For how model size maps to memory in the first place, see our sibling explainer on how much VRAM local AI models need.
The Machines on Sale or Announced in September 2026
Four unified-memory platforms matter this month: Strix Halo PCs and Nvidia's DGX Spark are on sale, Apple's M5 Mac Studio ships September 22, and Nvidia's RTX Spark arrives in October. They differ more in software and operating system than in memory size.
Strix Halo mini PCs and laptops
The Ryzen AI Max+ 395 pairs a 16-core, 32-thread Zen 5 CPU with a Radeon 8060S GPU and up to 128GB of LPDDR5X-8000 at 256 GB/s. Liliputing's March 2026 roundup listed nine 128GB mini PCs from $2,399 (Bosgame M5) to $3,810 (MINIX ER939-AI). AMD's own Ryzen AI Halo, which ships with Windows 11 Pro or Linux, sells for $3,999 with 128GB and a 2TB SSD. Framework's 128GB Desktop was $3,449 in July.
Gorgon Halo, the Ryzen AI Max 400 refresh
AMD's Ryzen AI Max Pro 400 series, codenamed Gorgon Halo, is a minor update to the same Zen 5 and RDNA 3.5 design. Tom's Hardware lists the flagship Max+ Pro 495 with a 100 MHz higher boost clock, a 40-CU Radeon 8065S, and support for up to 192GB of memory, 160GB of it usable by the GPU. It is a commercial "Pro" part so far. Lenovo's ThinkCentre X Ultra uses it with up to 128GB of LPDDR5X-8533 from November at an expected price of about $3,700, and Framework has previewed a 192GB Desktop with no price or date. Faster 8533 MT/s memory adds about 7% bandwidth over the 395, which doesn't change the picture below.
Nvidia DGX Spark
A 20-core Arm CPU and a Blackwell GPU share 128GB of LPDDR5X at 273 GB/s. Nvidia rates it at up to 1 petaflop of FP4 AI compute and says it can run inference on models up to 200 billion parameters. It runs DGX OS, Nvidia's Ubuntu-based Linux. Tom's Hardware put its price at $4,699 when the $3,999 Ryzen AI Halo launched.
Mac Studio with M5 Max or M5 Ultra
Apple announced both on August 25 and they ship September 22. The M5 Max has up to 128GB at 614 GB/s from $2,499. The M5 Ultra has up to 512GB at 1.2 TB/s from $5,499, and the 512GB configuration follows in late October. Starting prices are for base memory, so a 128GB configuration costs more.
Nvidia RTX Spark, arriving in October
RTX Spark is a Windows on Arm platform for laptops and small desktops, with up to 20 CPU cores, 6,144 CUDA cores, 128GB of LPDDR5X and up to 300 GB/s. Tom's Hardware reports that it launches in October in two configurations, with laptops from Dell, HP, Lenovo, ASUS, MSI and Microsoft. No prices or independent tests exist yet.
How Each One Handles Games
Only Strix Halo plays Windows PC games natively and well, averaging 67 FPS in Cyberpunk 2077 at 1080p high, while DGX Spark needs emulation and a Mac can't run kernel-level anti-cheat games. The GPU isn't the only thing that decides this. The operating system and the CPU instruction set matter as much.
Strix Halo: a real 1080p gaming PC
The Radeon 8060S is the best gaming option in this group, because it runs x86 Windows games on the normal AMD driver with no translation layer. In a test by YouTuber RandomGaminginHD on a Minisforum MS-S1 Max, reported by Tom's Hardware, it averaged these results at 1080p:
| Game | Settings | Average FPS |
|---|---|---|
| Counter-Strike 2 | High, MSAA off | 263.7 |
| Kingdom Come: Deliverance 2 | High | 89.8 |
| Battlefield 6 | High, FSR Native AA | 86.7 |
| Red Dead Redemption 2 | Mixed high, ultra textures | 82.2 |
| Cyberpunk 2077 | High preset | 66.8 |
| Cyberpunk 2077 | RT Ultra, FSR 3 Quality | 45.4 |
| Borderlands 4 | Lowest | 54.8 |
That is roughly RTX 4060 laptop territory. TechSpot's roundup of early laptop results had the 8060S at 39 FPS in Cyberpunk 2077 against 36 FPS for an RTX 4060 laptop and 37 FPS for an RTX 4070 laptop, while Nvidia's parts were up to 37% faster in GTA V. Plan on 1080p with upscaling, not 1440p ultra. For what ordinary laptop and desktop iGPUs manage by comparison, see our list of games that genuinely run on integrated graphics.
DGX Spark: it can play, with a translation layer in the way
DGX Spark's GPU has the same 6,144 CUDA cores as a desktop RTX 5070, but games still struggle on it. It runs Arm Linux, so Windows games go through Proton plus an x86 emulator, and its 273 GB/s is well under the RTX 5070's 672 GB/s.
A Reddit user's first test, reported by Tom's Hardware, got about 50 FPS in Cyberpunk 2077 at 1080p medium using Box64. An Nvidia representative replied that switching to FEX and Proton 10.2-2 beta produced more than 175 FPS at 1080p high with Ultra ray tracing, but that figure uses DLSS 4 Multi Frame Generation. Generated frames raise the counter without improving responsiveness the same way, as our breakdown of frame generation and input lag explains. The same Redditor called stability "actually pretty good" but still hit occasional crashes.
Mac Studio: fast hardware, missing games
The M5 Max and M5 Ultra have far more bandwidth than any PC APU, but macOS limits what you can play. Native Mac ports are a small slice of the PC library, and the rest runs through Apple's Game Porting Toolkit or CodeWeavers' CrossOver.
CrossOver 26, released in early 2026, got several anti-cheat-protected games running, including Helldivers 2, according to WINE for Mac's coverage. Games whose anti-cheat runs at the Windows kernel level, such as Valorant's Vanguard, still don't run on macOS. Riot lists Valorant for PC, Xbox Series X|S and PlayStation 5 only. If your group plays competitive shooters, a Mac isn't your gaming machine — however much memory it has.
RTX Spark: promising, unproven
Nvidia claims RTX Spark is good for "100 FPS 1440p gaming", with DLSS 4.5 and Multi Frame Generation likely doing much of that work. It also runs Windows on Arm, so most games will go through Microsoft's Prism x86 emulation. Treat the gaming claim as a vendor claim until independent reviews arrive after the October launch.
Why a Gaming GPU Still Wins for Most People
A discrete graphics card is faster at every AI model that fits in its VRAM and at every game, which covers most of what gamers actually run. The RTX 5090 has 1,792 GB/s of bandwidth, about seven times Strix Halo's 256 GB/s. Nvidia's $1,999 launch MSRP for the card is also below the $2,399 to $4,699 list prices of the 128GB boxes covered here, though street prices for any of them can differ.

The catch is the 32GB ceiling. Nothing gets around it. A 16GB card such as the RTX 5070 Ti handles models up to about 20B parameters at 4-bit comfortably. The RTX 5090 stretches that into the 30B range. Past about 32GB of weights, the rest of the model spills into system RAM, and speed falls off hard. The slow part is the RAM itself, about 96 GB/s for dual-channel DDR5-6000, not the PCIe slot.
The gaming PC also brings things unified-memory boxes can't. You can upgrade the GPU in two years. You get CUDA on Windows with no emulation. And none of the game compatibility problems above apply. Our best budget GPU for 2026 comparison covers the cheaper entry points, and how much VRAM you need for gaming covers the gaming side of the capacity question.
The trade is power. Nvidia rates the RTX 5090 alone at 575W of total graphics power and lists 1,000W of required system power. DGX Spark ships with a 240W brick, and AMD's Ryzen AI Halo has a 120W chip and a 240W adapter. If a box will run an AI agent around the clock, that difference shows up on the bill, as our guide to the cost of running a gaming PC works through.
Who Should Buy Which
Choose by the largest model you actually intend to run and by whether you need Windows games, not by the AI TOPS figure on the box. Most gamers curious about local AI will be happier with a GPU. A minority with a specific large model in mind won't be.
| If this describes you | Buy | Why |
|---|---|---|
| Games come first, and your models are 30B parameters or smaller at 4-bit | A gaming PC with a 16GB to 32GB GPU | Fastest at every model that fits, fastest image generation, CUDA on Windows, and anti-cheat games run |
| You want 70B-class or large MoE models and Windows games at 1080p on one machine | A Strix Halo mini PC or laptop | The only 128GB option that games natively, though ROCm is less mature than CUDA |
| AI development is the job and gaming barely matters | DGX Spark | Nvidia's CUDA stack plus 200 Gbps networking that links up to four units, which Nvidia says handles models up to 700 billion parameters |
| You need more bandwidth than any PC APU, or more than 128GB | Mac Studio M5 Max or M5 Ultra | 614 GB/s or 1.2 TB/s and up to 512GB, with games left to another device |
| You want a thin 128GB laptop | Nothing yet | RTX Spark's game compatibility is unproven until reviews land after the October launch |
Our position is blunter than most buying guides: if you can't name the model that needs more than 32GB, you don't need unified memory. That's the whole test.
This table is only as good as September 2026, though. Nobody outside Nvidia has published RTX Spark game tests, and AMD hasn't said whether consumer Gorgon Halo chips are coming, so two rows could look different by the end of the year.
If you are weighing an NPU laptop instead, note that the NPU is not what runs these large models. Our sibling piece on whether an NPU AI PC helps gaming covers what that chip actually does.
The Costs the Spec Sheet Leaves Out
A unified-memory box also costs you in slow long-prompt responses, less mature software, fan noise and memory you can never upgrade. Each shows up in reviews of shipping hardware, not just on paper.
Start with latency. Tom's Hardware found the Ryzen AI Halo's generation speed acceptable but slower than a Dell Pro Max GB10. The bigger gap was time to first token, which "can quickly rise to non-interactive levels with long contexts." At the extremes, the reviewer waited two to four minutes for a response to start. That matters for coding assistants, where the context grows all session.
Software is the second gap. DGX Spark ships with Nvidia's own stack and playbooks. AMD answered with preloaded ROCm, Lemonade and its own playbooks on the Ryzen AI Halo, though the review found the preinstalled vLLM failed to load Qwen 3.6-35B-A3B and some paths weren't documented.
Then noise. The same review describes the AI Halo's twin blowers as audible at idle, with "a notable high-pitched whine" under load, and found it louder than the Dell GB10 box during image generation. Our look at whether mini PC power supplies are good enough covers the external-brick side of that design.
Finally, memory prices. Framework's 128GB Desktop cost $2,851 in March 2026 and $3,449 by July. The GEEKOM A9 Mega went from a planned $2,099 to $3,199. Unified memory is soldered LPDDR5X, so you are buying all of it up front and can never add more later. That also makes the 32GB-versus-64GB question from our is 32GB of RAM overkill guide irrelevant here: you pick your ceiling at checkout.
Frequently Asked Questions
Can a Strix Halo mini PC replace a gaming PC?
For 1080p gaming, largely yes, because the Radeon 8060S performs close to an RTX 4060 laptop GPU. It averaged 67 FPS in Cyberpunk 2077 at 1080p high and 87 FPS in Battlefield 6 in Tom's Hardware's report. It won't match a desktop RTX 5070 or better at 1440p, and you can't upgrade the GPU later.
Is a unified-memory PC faster than an RTX 5090 for local AI?
Only when the model doesn't fit in 32GB. For anything that fits, the RTX 5090's 1,792 GB/s beats Strix Halo's 256 GB/s and DGX Spark's 273 GB/s by a wide margin. The unified-memory box wins by being able to load 70B-class dense or large mixture-of-experts models at all.
Can the Nvidia DGX Spark play games?
Technically, yes, but it isn't built for it. It runs Arm Linux, so games need Proton and an x86 emulator such as FEX. Early tests got about 50 FPS in Cyberpunk 2077 at 1080p medium, and Nvidia's own 175 FPS figure relies on DLSS 4 Multi Frame Generation.
Is a Mac Studio good for gaming?
Its hardware is strong, with 614 GB/s on the M5 Max, but game support is the limit. Native Mac ports are few, compatibility layers such as CrossOver run many Windows games, and titles with kernel-level anti-cheat such as Valorant do not run on macOS at all.
How much unified memory do I need for local AI?
Match it to your largest model plus room for context. A 70B dense model at Q4_K_M needs about 43GB for weights alone, so 64GB is tight and 128GB is comfortable. Nvidia rates its 128GB DGX Spark for inference on models up to 200 billion parameters. Past that you're looking at Gorgon Halo's 192GB or the M5 Ultra's 512GB. If your models stay at 30B or below, a 16GB to 32GB graphics card is the faster and cheaper option.
The Bottom Line
A unified-memory box is a capacity purchase: it loads AI models no graphics card can hold, at a fraction of a high-end GPU's speed. If you don't have a specific model that needs more than 32GB, you are paying $2,400 to $4,700 for memory you won't use — and giving up a lot of gaming performance to do it.
For a gamer with a real large-model need, Strix Halo is the only option here that is also a competent Windows gaming PC. DGX Spark suits AI developers who happen to game occasionally. The M5 Mac Studio wins on bandwidth and loses on games. Everyone else should put the money into a discrete GPU with as much VRAM as the budget allows. For how local AI is starting to show up inside games, see our coverage of AI NPCs in games.
Hardware photography courtesy of the respective manufacturers and publications, used for editorial coverage.
Sources
- Apple Newsroom, Apple introduces new Mac Studio with M5 Max and M5 Ultra, retrieved 2026-09-10, https://www.apple.com/newsroom/2026/08/apple-introduces-new-mac-studio-with-m5-max-and-m5-ultra/
- NVIDIA, DGX Spark: Personal AI Supercomputer, retrieved 2026-09-10, https://www.nvidia.com/en-us/products/workstations/dgx-spark/
- NVIDIA, GeForce RTX 5090 specifications, retrieved 2026-09-10, https://www.nvidia.com/en-us/geforce/graphics-cards/50-series/rtx-5090/
- NVIDIA Newsroom, NVIDIA Blackwell GeForce RTX 50 Series Opens New World of AI Computer Graphics, retrieved 2026-09-10, https://nvidianews.nvidia.com/news/nvidia-blackwell-geforce-rtx-50-series-opens-new-world-of-ai-computer-graphics
- Ollama, llama3.3 model tags, retrieved 2026-09-10, https://ollama.com/library/llama3.3/tags
- Tom's Hardware, AMD Ryzen AI Halo review: AMD builds a DGX Spark of its own, retrieved 2026-09-10, https://www.tomshardware.com/pc-components/gpus/embargo-mon-july-6-8am-pt-1100-edt-amd-ryzen-ai-halo-review
- Tom's Hardware, AMD challenges Nvidia's DGX Spark with $3,999 Ryzen AI Halo with Windows 11 support, retrieved 2026-09-10, https://www.tomshardware.com/desktops/mini-pcs/amd-challenges-nvidias-dgx-spark-with-usd3-999-ryzen-ai-halo-with-windows-11-support-strix-halo-desktop-undercuts-nvidia-by-usd700-packs-128gb-of-unified-memory
- Tom's Hardware, Strix Halo Radeon 8060S iGPU benchmarked in games, delivers butter-smooth 1080p performance, retrieved 2026-09-10, https://www.tomshardware.com/pc-components/gpus/strix-halo-radeon-8060s-benchmarked-in-games-delivers-butter-smooth-1080p-performance-ryzen-ai-max-395-apu-is-a-pretty-solid-gaming-offering
- Tom's Hardware, Nvidia's $3,999 mini AI supercomputer tested in gaming: DGX Spark tested at 1080p on medium settings in Cyberpunk 2077, retrieved 2026-09-10, https://www.tomshardware.com/video-games/pc-gaming/as-expected-nvidias-usd3-999-mini-ai-supercomputer-is-terrible-for-gaming-dgx-spark-struggles-to-hit-50-fps-at-1080p-on-medium-settings-in-cyberpunk-2077
- Tom's Hardware, Nvidia unveils RTX Spark Superchip for laptops and desktop PCs at Computex 2026, retrieved 2026-09-10, https://www.tomshardware.com/laptops/nvidia-unveils-rtx-spark-superchip-at-computex-2026-new-platform-promises-to-turn-windows-into-an-agentic-ai-os-with-arm-cpu-blackwell-gpu-and-128gb-unified-memory
- Tom's Hardware, Nvidia's RTX Spark N1X launches in October for laptops and desktops, retrieved 2026-09-10, https://www.tomshardware.com/laptops/nvidias-rtx-spark-n1x-launches-in-october-for-laptops-and-desktops-18-or-20-cpu-cores-paired-with-5-120-or-6-144-cuda-cores-up-to-128gb-of-unified-memory
- Tom's Hardware, AMD Ryzen AI Max 400 'Gorgon Halo' packs up to 192GB of unified memory, retrieved 2026-09-10, https://www.tomshardware.com/pc-components/cpus/amd-ryzen-ai-max-400-gorgon-halo-packs-up-to-192gb-of-unified-memory-refreshed-apu-uses-zen-5-and-rdna-3-5-and-can-clock-up-to-5-2-ghz
- Tom's Hardware, Lenovo ThinkCentre X Ultra packs Gorgon Halo, retrieved 2026-09-10, https://www.tomshardware.com/laptops/lenovo-thinkcentre-x-ultra-packs-gorgon-halo-amd-ryzen-ai-max-pro-495-shows-up-in-mini-workstation
- TechSpot, AMD Ryzen Strix Halo APU is matching the RTX 4070 laptop GPU in gaming benchmarks, retrieved 2026-09-10, https://www.techspot.com/news/106835-amd-ryzen-strix-halo-laptop-apus-match-rtx.html
- Liliputing, More Ryzen AI Max+ 395 mini PCs with 128GB are now available... if you can afford one, retrieved 2026-09-10, https://liliputing.com/more-ryzen-ai-max-395-mini-pcs-with-128gb-are-now-available-if-you-can-afford-one/
- Notebookcheck, Framework launches world's first mini-ITX desktop PC with Ryzen AI Max+ Pro 495 and 192 GB RAM, retrieved 2026-09-10, https://www.notebookcheck.net/Framework-launches-world-s-first-mini-ITX-desktop-PC-with-Ryzen-AI-Max-Pro-495-and-192-GB-RAM.1349336.0.html
- Wikipedia, GeForce RTX 50 series, retrieved 2026-09-10, https://en.wikipedia.org/wiki/GeForce_RTX_50_series
- WINE for Mac, CrossOver 26 Breaks the Anti-Cheat Barrier, retrieved 2026-09-10, https://wineformac.org/news/blog-crossover-26-anti-cheat-2026.html
- Riot Games, VALORANT platform selection, retrieved 2026-09-10, https://playvalorant.com/en-us/platform-selection/
- Riot Games, Vanguard x VALORANT, retrieved 2026-09-10, https://playvalorant.com/en-us/news/game-updates/vanguard-x-valorant/



