Beelink GTR9 Pro Review: 128GB, Dual 10GbE, and the Real LLM Numbers

The Beelink GTR9 Pro is the first Strix Halo mini PC I’d actually tell someone to buy for local AI work, and that’s mostly because it stops pretending. It’s a roughly one-litre box with AMD’s Ryzen AI Max+ 395, 128GB of soldered LPDDR5X, dual 10-gigabit networking, and a price that lands between CA$2,523 (US$1,800) and CA$2,801 (US$1,999) depending on the day. On paper it runs a 120-billion-parameter model. In practice, whether that’s exciting or a slideshow depends entirely on what kind of 120-billion-parameter model you feed it.

Currency note: Canadian dollar equivalents use the Bank of Canada daily average of 1 USD = 1.4014 CAD on July 17, 2026. Converted amounts are approximate; the original US price is shown in brackets.

I haven’t had one on my own desk, so everything below is measured by the reviewers who did: ServeTheHome ran it as a full workstation, Notebookcheck tore it down, and a stack of hobbyists posted raw token counts. The numbers agree with each other, and they tell a more useful story than the box copy does. Here’s what the hardware is, what it’s genuinely good at, and the three things that would make me wait a firmware update or two before buying.

Table of Contents

What the Beelink GTR9 Pro actually is

Strip the marketing and the Beelink GTR9 Pro is a delivery vehicle for one chip: AMD’s Ryzen AI Max+ 395, codenamed Strix Halo. That’s 16 Zen 5 cores, 32 threads, and a Radeon 8060S integrated GPU with 40 RDNA 3.5 compute units, all sharing a single 128GB pool of LPDDR5X-8000. The rest of the box is built to keep that chip fed and cool: a 230W internal power supply, a 2TB NVMe SSD, and a vapor-chamber dual-fan cooler that Notebookcheck’s teardown found fills most of the chassis.

AMD Ryzen AI Max+ 395 processor, the chip inside the Beelink GTR9 Pro
Image: AMD newsroom. The Ryzen AI Max+ 395 is the whole reason this box exists.

What separates it from the pile of near-identical Strix Halo boxes is the connectivity. ServeTheHome measured dual 10GBase-T ports on an Intel E610 controller, two USB4 ports at 40Gbps, HDMI 2.1, and DisplayPort 2.1 capable of driving an 8K display. For a machine you’re likely to park on a shelf and hit over the network, that dual 10-gigabit link is the difference between a toy and a real inference server on your LAN. And if it’s going to live in a closet as an always-on box, configure auto power on after a power outage so it comes back on its own.

The 128GB unified memory trick

The reason anyone cares about a Beelink GTR9 Pro instead of a normal desktop is the memory model. Like Apple Silicon, Strix Halo gives the CPU, GPU, and NPU one shared pool. You can hand up to 96GB of that 128GB to the GPU on Windows, or a more generous 110GB on Linux. A 70-billion-parameter model that would need two CA$2,803 (US$2,000) graphics cards to hold in VRAM fits here in one box that draws about 120W under load.

The catch is bandwidth, and it’s a big one. That memory tops out around 256 GB/s on paper, and runaihome’s testing measured roughly 215 GB/s in real mixed workloads. That sounds fast until you learn an RTX 5090 moves 1,792 GB/s. Token generation is bandwidth-bound: the chip has to read the model’s weights for every single token it produces. So the GTR9 Pro can hold enormous models that a gaming GPU can’t, but it reads them roughly eight times slower. That trade defines everything about how it performs.

The token numbers that actually matter

Here’s where the box copy gets sneaky. Yes, the GTR9 Pro runs a 120-billion-parameter model. Notebookcheck clocked GPT-OSS 120B at around 31 tokens per second while the box sat at 120W, and a tuned Linux setup pushes that past 55 tokens per second per runaihome. That’s genuinely usable, real-time chat speed from a puck on your desk.

But GPT-OSS 120B is a Mixture-of-Experts model. It has 120 billion weights total and activates only a small fraction of them per token, so the memory system only has to read a few billion parameters each pass. Feed the same box a dense 70-billion-parameter model like Llama 3.3, where every weight fires on every token, and Hardware Corner’s numbers drop to 5 to 9 tokens per second. That’s slower than you can comfortably read. Same hardware, same quantization, a 10x difference driven entirely by the model’s architecture.

Beelink GTR9 Pro tokens per second chart comparing MoE and dense LLM speeds
Chart: thedeskbrief.com

The practical takeaway: this machine loves sparse MoE models and small dense ones. A 30-billion-parameter MoE like Qwen3-30B-A3B screams along near 100 tokens per second. If your workflow lives on today’s open MoE releases, the GTR9 Pro is a bargain. If you’re wedded to a big dense 70B, buy it knowing you’re getting a patient assistant, not a fast one.

The catches: heat, drivers, and a firmware to-do list

No product this new ships clean, and the Beelink GTR9 Pro is no exception. The most concerning report comes from ServeTheHome’s testing, where the 10-gigabit NICs crashed under sustained load. For a box whose whole pitch is being a networked inference server, a network stack that falls over when you actually push it is not a small footnote. It looks like a driver and firmware problem rather than dead silicon, so it may well be patched, but I’d confirm it’s fixed before I trusted this thing with anything unattended.

Noise is the second question mark. Notebookcheck’s controlled teardown measured a reasonable 36 to 41 dBA under a 120W LLM load, which is quiet for the wattage. But at least one hands-on reviewer reported fan noise hitting an unpleasant 52 dBA, which suggests the cooling behavior depends heavily on ambient temperature and how hard you’re pushing it. And ServeTheHome flagged general BIOS and firmware rough edges that need cleaning up. None of this is a dealbreaker on a CA$2,523 (US$1,800) machine, but it’s the reason I’d treat an early unit as an enthusiast purchase, not an appliance.

Beelink GTR9 Pro vs the alternatives

Compared to AMD’s own Ryzen AI Halo developer box, which uses the identical chip but sells for CA$5,604 (US$3,999) through a single US retailer, the GTR9 Pro is the same silicon at less than half the money with better networking. That’s an easy call unless you specifically need AMD’s supported dev-kit software stack. Against Lenovo’s Yoga AI mini PC, which pairs a slower Panther Lake chip with narrower memory bandwidth, the Beelink wins on raw inference throughput by a wide margin.

The real competition is a used discrete GPU. A single 24GB card will destroy this box on any model that fits in 24GB of VRAM, running several times faster. The GTR9 Pro only pulls ahead when the model is too big for consumer VRAM, which is exactly the 70B-and-up territory where its huge memory pool becomes the only affordable option. If you’re weighing a box like this against NVIDIA’s RTX Spark approach, the deciding question is always the same: how big are your models, and are they sparse or dense?

Who Should Skip This

Skip the GTR9 Pro if your models fit in a 16GB or 24GB graphics card, because a GPU you might already own will run them faster and cheaper. Skip it if you mainly run big dense models and expect snappy replies; 7 tokens per second on a dense 70B will test your patience daily. Skip it if you need a plug-and-forget appliance today, because the NIC and firmware reports mean an early unit still wants an owner who reads release notes.

Buy it if you want the cheapest way to keep large MoE models resident in memory, run them privately, and reach the box over a fast home network. That’s a real and growing use case, and nothing else at this price does it in this footprint.

Beelink GTR9 Pro FAQ

How much does the Beelink GTR9 Pro cost?

Street pricing for the 128GB, 2TB configuration sits between CA$2,523 (US$1,800) and CA$2,801 (US$1,999), which undercuts AMD’s own Ryzen AI Halo dev box using the same chip by roughly CA$2,803 (US$2,000).

Can it really run a 120B model?

Yes, and reasonably fast, but only because GPT-OSS 120B is a Mixture-of-Experts model that activates a fraction of its weights per token. A dense model of the same size would run far slower on the same hardware.

Is the Beelink GTR9 Pro good for gaming?

The Radeon 8060S handles 1080p gaming well, but that’s not why this box exists. If gaming is your priority, a desktop with a discrete GPU at this price is a smarter buy. This is an inference machine that can game, not the reverse.

Windows or Linux for local AI on this box?

Linux, if you’re comfortable with it. You can allocate 110GB to the GPU versus 96GB on Windows, and the tuned Linux token rates on GPT-OSS 120B ran noticeably higher in testing. If you’d rather not tinker, running a home AI agent on Windows still works fine.

The Short Version

The Beelink GTR9 Pro is the most sensible Strix Halo box shipping right now: 128GB of unified memory, dual 10GbE, and vapor-chamber cooling for between CA$2,523 (US$1,800) and CA$2,801 (US$1,999). It’s the cheapest honest path to running large Mixture-of-Experts models locally and privately. Just go in clear-eyed. Sparse MoE models fly at 31 to 100 tokens per second, dense 70B models crawl at 7, and early units carry NIC-driver and firmware rough edges that reward a buyer who’s happy to babysit a fix or two. Match your model type to the hardware and it’s a genuine bargain. Buy it expecting a fast dense-70B machine and you’ll be disappointed.


Sources: ServeTheHome’s GTR9 Pro review, Notebookcheck’s teardown, and runaihome’s Strix Halo LLM benchmarks.