Skip to main content
Tech News

AMD Helios claims 15% more AI compute than NVIDIA’s Vera Rubin NVL72

The AI data center wars have officially gone rack-scale, and AMD just fired its biggest shot. AMD Helios is a complete rack-scale AI system combining 72 Instinct MI455X GPUs, sixth-generation EPYC Venice CPUs, Pensando networking and ROCm software into a single deployable unit, and AMD says it delivers 15 percent more peak FP4 compute than NVIDIA’s flagship Vera Rubin NVL72.

Key Takeaways

  • AMD Helios is a rack-scale AI system with 72 Instinct MI455X GPUs and EPYC Venice CPUs.
  • It delivers 2.9 ExaFlops of FP4 compute, a claimed 15 percent lead over Vera Rubin NVL72.
  • 31TB of HBM4 memory is 50 percent more than the NVIDIA rack, with 1.7 PB/s of memory bandwidth.
  • AMD is pitching 30 percent more tokens per dollar as the economic argument.

What Helios actually is

Helios is not a GPU launch; it is a system launch, and that distinction is the story. Modern AI infrastructure is bought by the rack, not the chip, because networking, memory bandwidth and power delivery matter as much as raw compute once you are wiring thousands of accelerators together. Helios packages AMD’s best of everything into one building block: MI455X accelerators for compute, Venice server CPUs for orchestration, Pensando for the networking fabric and ROCm as the open software layer. Scale from one rack to a gigawatt cluster by adding more of the same unit.

The headline specs are aggressive: 2.9 ExaFlops of FP4 compute per rack, 31TB of HBM4 memory, 43 TB/s of scale-out bandwidth and 1.7 PB/s of aggregate memory bandwidth. Each figure is aimed at a specific NVIDIA comparison, and AMD is not being subtle about it.

The Vera Rubin comparison, with salt

Vendor-versus-vendor benchmarks deserve their usual seasoning. The 15 percent FP4 compute lead is peak theoretical throughput on one precision format, the one AMD chose to highlight. Real-world training and inference performance depends on software maturity, interconnect efficiency and workload shape, and NVIDIA’s software ecosystem remains its deepest moat. That said, the memory story is harder to spin: 50 percent more HBM4 per rack is a physical advantage, and memory capacity increasingly decides which models fit and how efficiently they run.

Tokens per dollar is the metric that matters

AMD’s sharpest claim is economic: 30 percent more tokens per dollar. For the companies spending hundreds of billions on AI infrastructure, that is the number that moves purchase orders. Capital expenditure at this scale is an accounting exercise, and a credible efficiency edge, even a contested one, forces every procurement team to run the comparison. Whether or not Helios wins, NVIDIA now has to defend on price-performance, which it has not seriously had to do in years.

The software moat is the real battlefield

Ask anyone deploying AI infrastructure what actually decides the vendor, and the answer is rarely the silicon. It is the software. NVIDIA’s CUDA ecosystem has a decade of libraries, tooling, trained engineers and production-proven stacks behind it, and switching costs are measured in engineering years, not benchmark deltas. AMD’s answer is ROCm, and to its credit the platform has matured from an afterthought into something genuinely deployable, with the major frameworks running well. But “runs well” and “runs identically to the stack your team already knows” are different claims, and the second one is what wins nine-figure contracts.

This is why the Helios pitch leads with system economics rather than chip speeds. AMD is not trying to out-CUDA CUDA; it is trying to make the total cost argument strong enough that buyers justify the migration. The 30 percent tokens-per-dollar claim is aimed precisely at that calculation.

What to watch next

The proof points that will settle this are not spec sheets but deployments: which frontier labs and cloud providers commit to Helios at scale, what real-world training throughput looks like against NVIDIA’s racks, and whether ROCm’s software cadence keeps pace. Watch the design-win announcements over the next two quarters. They will tell you more than any benchmark.

Why the rack wars matter beyond the data center

Competition at this tier cascades downward. A genuine two-horse race in AI infrastructure accelerates the technology, pressures pricing and ultimately shapes what AI costs for everyone downstream, from startups renting clusters to consumers buying the devices AI services run on. The spending context is staggering, as we detailed in our report on data centers heading for 20 percent of US power, and every efficiency gain at the rack level compounds across that entire buildout.

It also validates AMD’s long game. Years of EPYC share gains funded the accelerator program, the Instinct line matured, and now the company is competing for the most lucrative infrastructure market in tech history with a complete system, not a challenger chip. On the desktop side we track the same rivalry in pieces like our Nova Lake leak analysis, but the rack is where the money now lives.

There is also an ecosystem argument running alongside the performance one. AMD has positioned ROCm and its rack architecture as the open alternative, and a meaningful slice of the market actively wants a credible second source for AI compute, if only for negotiating leverage and supply resilience. Every percentage point of share AMD takes makes the entire AI buildout less dependent on a single vendor’s roadmap, pricing and allocation decisions. Even buyers who never deploy Helios benefit from its existence at the negotiating table.

The timing matters as much as the technology. AI infrastructure orders are being placed now for capacity that comes online over the next several years, which means the vendor decisions being made this year lock in market share deep into the decade. AMD does not need to win the majority of those decisions for Helios to matter; winning a meaningful slice establishes the company as a permanent second pole in AI compute, with all the roadmap influence and pricing pressure that follows from it.

Full Helios specifications are published on AMD’s official site.

The bottom line

The Bottom Line

Helios is AMD’s most credible challenge yet to NVIDIA’s AI infrastructure dominance: more FP4 compute, 50 percent more memory, and a tokens-per-dollar argument aimed straight at procurement spreadsheets. The claims need independent validation, but the era of uncontested rack-scale supremacy is over, and that alone is worth the announcement.

Does NVIDIA finally have a real rack-scale rival? Tell the tech desk.

Share this article
Alex Mercer

Alex Mercer

Alex Mercer is the managing editor of Tech News Xplore. A hardware journalist for over a decade, he has benchmarked more CPUs, GPUs and handhelds than he can count, and leads the publication's testing methodology and editorial standards.

Join the discussion

Keep it civil and on topic. Comments are moderated according to our comment policy in the Terms of Use.