AMD Ryzen AI Halo: 128GB of Unified Memory for US$700 Less Than the Spark

The AMD Ryzen AI Halo is a CA$5,604 (US$3,999) mini PC with 128GB of unified memory, and it exists for exactly one reason: NVIDIA’s DGX Spark proved people will pay workstation money for a tiny box that runs huge language models at home. AMD looked at that, looked at the Spark’s price climbing to CA$6,585 (US$4,699), and slid its own version across the table for CA$981 (US$700) less. With Windows.

Currency note: Canadian dollar equivalents use the Bank of Canada daily average of 1 USD = 1.4014 CAD on July 17, 2026. Converted amounts are approximate; the original US price is shown in brackets.

I’ve been paying close attention to this category, partly because I already covered the NVIDIA Spark hardware on this site, and partly for a selfish reason: the laptop that runs my local AI models has 8GB of VRAM. Eight. Every interesting model released in the past year laughs at that number. Unified memory machines are the first real answer for people like me, so let’s take the AMD Ryzen AI Halo apart properly, including the parts AMD would rather leave out of the headline.

Why 128GB of Unified Memory Is the Whole Story

On a normal PC, your GPU has its own small pool of fast memory, and a model either fits in it or spills into system RAM and slows to a crawl. Unified memory tears that wall down. The CPU and GPU share one big pool, so the GPU can address nearly all of the 128GB. Models that would need a stack of consumer graphics cards suddenly fit in a box the size of a sandwich.

That’s the pitch, and it’s genuinely true. AMD says the Ryzen AI Halo handles models up to around 200 billion parameters locally. Fitting a model and running it quickly are two different things, though, and that gap is where this whole product category gets interesting. Hold that thought.

The Spec Sheet That Matters

Inside the 150 x 150 x 45mm chassis (yes, it’s basically a stack of coasters that weighs just over a kilogram) sits the Ryzen AI Max+ 395: 16 Zen 5 cores, a Radeon 8060S integrated GPU with 40 RDNA 3.5 compute units, and a 50 TOPS XDNA 2 NPU. The 128GB of LPDDR5X-8000 memory feeds all of it at 256 GB/s. You get a 2TB NVMe drive included at the CA$5,604 (US$3,999) price, plus 10GbE Ethernet, Wi-Fi 7, HDMI 2.1b, and three USB-C data ports. The whole thing sips 120W through a fourth USB-C port. Full details are in AMD’s announcement coverage on TechPowerUp.

The detail I like most: it ships in two variants with identical hardware, one with Windows 11 Pro and one with a Linux developer image. More on why that matters in a minute.

AMD Ryzen AI Halo vs DGX Spark: The Real Numbers

I don’t have either machine on my desk, so I’ll lean on people who do. StorageReview’s testing put both boxes through vLLM serving benchmarks, and the results are lopsided in a way the price gap doesn’t capture. Running GPT-OSS 120B at batch size 64, the Halo produced 222 tokens per second against the Spark’s 701. On Llama 3.1 8B it was 407 versus 1,330. Switch to FP4 quantization, where NVIDIA’s Blackwell silicon has dedicated hardware, and the gap blows out to more than 13x.

AMD Ryzen AI Halo vs NVIDIA DGX Spark comparison chart: price, unified memory, bandwidth, benchmarks, and operating systems
Where each box wins. Benchmark rows are StorageReview’s vLLM measurements.

Before you close the tab: those are server-style numbers, measured with dozens of simultaneous requests. That’s the workload the Spark was built for, and it’s not how most home users run models. If it’s just you chatting with a model, generation speed is mostly limited by memory bandwidth, and there the two boxes are neighbors: 256 GB/s for the Halo, 273 GB/s for the Spark. For single-user local AI, the practical experience is far closer than the benchmark chart suggests.

A rough rule of thumb for this class of machine: divide memory bandwidth by model size. A 70B model quantized to 4-bit weighs about 40GB, so 256 GB/s gives you a ceiling in the neighborhood of 6 tokens per second. That’s usable reading speed, not magic. The same math applies to the Spark, which is exactly why neither of these boxes replaces a rack of real GPUs. They replace not being able to run the model at all.

The Catches Nobody Puts in the Headline

First, that bandwidth number is the ceiling on everything. 256 GB/s is enormous for an integrated GPU and unimpressive next to a discrete card; a desktop RTX card moves memory several times faster. Big models fit, but they stroll rather than sprint.

Second, software. NVIDIA’s CUDA ecosystem has a fifteen-year head start, and most AI tooling treats it as the default. AMD’s ROCm support has improved a lot, and the popular runtimes (llama.cpp, LM Studio, Ollama) run fine on this hardware, but you will eventually hit a research project or a quantization format that assumes NVIDIA. Buying the AMD Ryzen AI Halo means occasionally being second in line.

Third, the storage is half as fast as the Spark’s. StorageReview measured about 6,900 MB/s sequential reads against roughly 13,400 MB/s, which you’ll feel every time a 60GB model loads into memory. And fourth, there’s no answer to the Spark’s 200G cluster fabric. Two Sparks can be linked to run even bigger models; the Halo’s single 10GbE port ends that conversation. If any of these four points made you wince, that’s useful information about which machine you need.

The Windows Thing Actually Matters

Here’s the argument I find most persuasive, via Tom’s Hardware: this is the only machine in the category that’s a normal computer. The DGX Spark runs NVIDIA’s DGX OS, a specialized Linux. A Mac Studio runs macOS. The AMD Ryzen AI Halo is a regular x86 PC that happens to have 128GB of unified memory, so the Windows variant runs your existing software, your games (the Radeon 8060S is the strongest integrated GPU AMD has ever shipped), and your AI stack on one box.

For a developer who lives in Linux anyway, that’s a shrug. For someone who wants one desktop that does everything and also runs a serious home AI agent, it’s arguably the whole reason to pick AMD here. When this box eventually stops being your AI machine, it’s still a killer general-purpose PC. A used DGX Spark is a paperweight with a great story.

Should You Wait for the 192GB Version?

AMD has already said a Ryzen AI Max+ 495 variant with 192GB is coming around Q3 2026, with support for 300B+ parameter models. Normally I’d say waiting for the next version is a trap you can fall into forever. In this case the calendar makes it a fair question, because Q3 is close.

My take: if 128GB covers the models you actually want to run today, buy the current one, because memory prices are rising and the Spark’s own price hike shows this category doesn’t get cheaper while you wait. If your dream model needs more than 128GB, wait for the 495. Buying hardware for the workload you have, not the one you imagine, has never once failed me.

Who Should Skip This

Skip it if your models fit in 16GB or 24GB, because a used RTX 4090 or a mid-range gaming desktop will generate tokens several times faster for less money. Skip it if you’re serving models to a team or building anything multi-user; StorageReview’s concurrency numbers show that’s the Spark’s home turf, and the price difference won’t matter to a business. Skip it if you’re already comfortable in the Apple ecosystem and bandwidth is your bottleneck, since a Mac Studio out-muscles both of these on that front. And obviously skip it if local AI is a curiosity rather than a habit. CA$5,604 (US$3,999) is a lot of money for a hobby you might have until October. If you’re not sure which camp you’re in, my Spark hardware comparison covers the NVIDIA side of the fence.

FAQ

Can the AMD Ryzen AI Halo really run 200B parameter models?

Yes, with quantization, and that’s AMD’s own claim rather than an independent one. A 200B model compressed to 4-bit fits in roughly 110GB, which squeaks under the 128GB ceiling. It will run. It won’t be quick, because at that size the 256 GB/s memory bus becomes the speed limit.

Is it faster than the DGX Spark?

No. In StorageReview’s vLLM testing the Spark was 2x to 13x faster depending on model and quantization, with the biggest gaps in FP4 workloads where Blackwell has dedicated hardware. The Halo’s case is price, Windows support, and being a usable general-purpose PC, not raw inference speed.

Can you game on it?

The Windows 11 variant is a normal x86 PC, and the Radeon 8060S with 40 compute units is the most capable integrated GPU AMD has shipped. It’s no desktop RTX card, but it’s a real gaming-capable chip, which is more than either the Spark or a Mac can say about your Steam library.

What ports and storage does it come with?

A 2TB NVMe SSD is included at CA$5,604 (US$3,999), alongside 10GbE Ethernet, Wi-Fi 7, Bluetooth 5.4, HDMI 2.1b, and three USB-C data ports. Power comes in over a separate USB-C port at 120W, so there’s no external brick drama.


The Short Version

The AMD Ryzen AI Halo is the value pick in a category that didn’t have one. It’s CA$981 (US$700) cheaper than the DGX Spark, it includes 2TB of storage, and it’s the only option that doubles as an everyday Windows machine. The Spark remains meaningfully faster for serious serving work, the CUDA ecosystem is still the safer bet for bleeding-edge tooling, and a 192GB version is only a quarter away. But if you want 128GB of unified memory on your desk without leaving Windows or spending five grand, the AMD Ryzen AI Halo is the box I’d put my own money on.