devtake.dev

AMD's Helios rack packs 72 GPUs and 31TB of memory to take on Nvidia

AMD launched the Instinct MI455X and its Helios rack-scale system, tying 72 GPUs and 31TB of HBM4 into one machine to challenge Nvidia's data-center grip.

Hiro Tanaka · · 4 min read · 4 sources
AMD launch slide reading 'Launching AMD Instinct MI455X GPUs' with a data-center photo and specs: CDNA 5, 2nm process node, HBM4, FP4/FP6/FP8, 320B transistors
Image via Phoronix · Source

AMD just launched Helios, a rack-scale AI system built to fight Nvidia where it’s strongest. The rack ties 72 of AMD’s new Instinct MI455X accelerators into what behaves like one giant machine, and it’s the company’s most direct shot yet at the GB200-class systems that run today’s biggest AI models.

Rack-scale is where the fight is now. Nvidia’s lead comes from systems like the NVL72, where 72 GPUs act as one shared pool of memory and compute, more than from any single chip. AMD has had competitive accelerators for a while. What it lacked was that rack. Helios is the answer, and AMD says customers including Microsoft and Anthropic are already lined up.

What we know

Start with the chip. The Instinct MI455X is AMD’s first CDNA 5 accelerator, a 320-billion-transistor part built on TSMC’s 2nm process. Its headline feature is memory: 432GB of HBM4 per GPU, which AMD says is roughly 50% more than the previous MI355X and more than Nvidia’s competing Rubin-class parts carry. Bandwidth climbs too, to 23.3TB/s per GPU. For AI work, how much memory sits on one accelerator decides how large a model you can hold without splitting it, so that gap matters.

The confirmed specs:

  • Each Helios rack packs 72 MI455X accelerators for 31TB of HBM4 and a claimed 2.9 exaflops of FP4 compute, per AMD’s launch numbers.
  • The GPUs talk over UALink-over-Ethernet, an open fabric AMD built with Broadcom’s Tomahawk switches, with HPE as lead integration partner, reports Phoronix.
  • AMD’s own next-gen EPYC Venice server CPUs run the rack, for roughly 4,600 CPU cores and 18,000 GPU compute units per cabinet (Phoronix).
  • Against Nvidia’s Vera Rubin NVL72, Helios carries 31TB of HBM4 to Nvidia’s 20.7TB, about 50% more (StorageReview).
  • AMD says the racks are in production, with shipments this quarter ramping into Q4 and the first half of 2027, and names Microsoft, Meta, Oracle, OpenAI, and Anthropic as customers (TechCrunch).

AMD isn’t shy about the stakes. “By 2030, the AI accelerator market is going to reach about $1.4 trillion,” CEO Lisa Su said at the launch, adding that it will “approach the size of the entire semiconductor market today,” according to TechCrunch. Helios is AMD’s bid for a real slice of that.

What we don’t know

The specs are AMD’s, and vendor benchmarks flatter the vendor. AMD claims Helios beats the NVL72 on tokens per dollar for inference, but independent numbers on real models aren’t out yet. Nvidia has a long habit of cherry-picking its own comparisons, and there’s no reason to assume AMD won’t do the same.

Pricing is the second unknown. Reports peg Helios at a premium over Nvidia’s second-generation Rubin, and AMD is betting customers pay it for the extra memory. Whether that math holds outside the handful of named hyperscalers is still open.

The real elephant is software. CUDA, Nvidia’s programming stack, is 18 years of tooling, libraries, and muscle memory that AMD’s ROCm still has to match. On the r/hardware launch thread, the recurring take was that the CUDA moat is finally eroding, partly because AI coding tools now make porting kernels far cheaper than before. One commenter’s read: OpenAI won’t blink at assigning 20 engineers to port a CUDA kernel if it lifts throughput. That’s the bet AMD is making too.

Phoronix and StorageReview published the detailed spec breakdowns; TechCrunch covered the launch and the customer list. AMD’s separate press release with Anthropic put hard numbers on at least one deployment: up to 2 gigawatts of Instinct GPUs, with the first gigawatt starting in the first half of 2027.

What this means for you

If you rent GPU time, this is the first credible reason to expect Nvidia’s pricing to feel pressure. A second vendor with a rack that matches or beats the NVL72 on memory hands hyperscalers real bargaining power, and that eventually shows up in what you pay for inference. Don’t expect it tomorrow. Helios volume doesn’t ramp until late 2026 into 2027, and the software gap means the early wins go to shops big enough to port their own kernels, like OpenAI and Anthropic. For most developers, the signal is simpler. If you’ve written off AMD because “it’s all CUDA anyway,” that assumption is worth re-testing. ROCm plus AI-assisted porting is a different proposition than it was two years ago. Benchmark before you commit, but AMD is back in the conversation.

Share this article

Quick reference

HBM4
High-bandwidth memory, stacks of DRAM bonded together for huge throughput. It's the memory AI accelerators like Nvidia's GPUs depend on, and the supply now competes with phones and PCs.
UALink
Ultra Accelerator Link, an open standard that wires many accelerators together so a rack of GPUs acts as one shared memory pool. It's AMD's answer to Nvidia's proprietary NVLink.

Sources

Mentioned in this article