AMD Helios Takes AI Infrastructure Fight to Rack Scale

AMD’s first production rack-scale AI system combines MI455X GPUs, EPYC CPUs, Pensando networking and ROCm software as the company pushes deeper into the infrastructure surrounding AI compute.

Key Highlights

  • Helios combines 72 liquid-cooled AMD Instinct MI455X GPUs with EPYC CPUs and Pensando networking, delivering high-density AI compute for training and inference.
  • The platform emphasizes rack-scale integration using open standards, enabling flexible sourcing and reducing reliance on proprietary ecosystems.
  • AMD's new EPYC 9006 processors support complex AI workloads, reflecting the industry's need for substantial CPU capacity alongside accelerators.
  • Deployment plans include major cloud providers like Microsoft Azure and commitments from OpenAI, indicating strong industry validation.
  • The evolution of AI infrastructure now extends beyond silicon to encompass data center power, cooling, and network architecture, shaping future AI campus designs.

AMD is escalating its challenge to Nvidia with Helios, a rack-scale AI system that puts the company squarely into the race to define how the next generation of AI factories are built.

Unveiled in production form at AMD’s Advancing AI 2026 event in San Francisco, Helios combines 72 Instinct MI455X GPUs with sixth-generation EPYC “Venice” CPUs, Pensando networking and AMD’s ROCm software stack.

The significance goes beyond another generation of faster accelerators.

Like Nvidia’s Vera Rubin platform, Helios treats the rack as an integrated compute system in which GPUs, CPUs, memory, networking, power delivery and cooling increasingly have to be engineered together. For data center operators, that means the competitive battle between the two chip companies is moving directly into infrastructure design.

AMD said Helios is now in production, with deployments beginning during the second half of 2026.

The Rack Becomes the System

Helios is built around AMD’s Instinct MI455X, a liquid-cooled accelerator based on the company’s CDNA 5 architecture and equipped with HBM4 memory.

A complete Helios rack delivers 72 GPUs along with EPYC host CPUs and Pensando networking for front-end, scale-up and scale-out traffic. AMD is positioning the platform for both large-scale training and increasingly important inference workloads.

AMD says Helios can deliver up to 30% more inference tokens per dollar than a competing system. The company also claims the MI455X provides more peak AI compute and substantially greater memory capacity than Nvidia’s Rubin GPU.

Those numbers are AMD benchmarks rather than independent comparisons. But the larger architecture may matter more than the percentages.

AI infrastructure is rapidly moving beyond the model of servers being installed as largely independent pieces of IT equipment. Accelerators have to exchange enormous volumes of data with each other while CPUs orchestrate workloads and networking connects increasingly large clusters across rows, halls and campuses.

AMD’s answer is to make Helios the repeatable building block.

The architecture uses open standards including the Open Compute Project’s Open Rack Wide design, UALink and Ethernet-based networking. That gives AMD a potentially important point of differentiation: rack-scale integration without requiring customers to adopt every element of a single proprietary ecosystem.

For operators and hyperscalers trying to preserve multiple sourcing options, that could become as important as raw accelerator performance.

A Data Center Infrastructure Contest

Helios also illustrates how closely the server roadmap is becoming tied to the physical data center.

The system is direct-liquid cooled, while its Open Rack Wide form factor departs from the conventional 19-inch rack that defined enterprise computing for decades.

That distinction matters as AI campuses move toward much denser infrastructure.

Nvidia has already laid out a roadmap moving from today's high-density Blackwell systems toward Vera Rubin and ultimately rack architectures operating at several hundred kilowatts. Those systems are driving the industry toward liquid cooling, higher-voltage power distribution and tighter coordination between IT equipment and the electrical and mechanical infrastructure supporting it.

AMD is now entering that same conversation with its own rack architecture.

That infrastructure requirement is already coming into focus. Schneider Electric’s reference design for Helios is built around 246-kW liquid-cooled racks, with modular AI clusters scaling to as much as 10.4 MW. The design integrates medium- and low-voltage electrical distribution, cooling and physical infrastructure around the AMD platform — a useful measure of how quickly the competition between AI architectures is extending beyond silicon and directly into data center power and thermal design.

The result is a competitive dynamic that increasingly reaches beyond GPUs. A hyperscaler selecting an AI platform is also making decisions about network architecture, cooling topology, rack dimensions, electrical distribution and ultimately how efficiently thousands of racks can be deployed across a multi-hundred-megawatt or gigawatt-scale campus.

CPUs Return to the AI Conversation

AMD also used Advancing AI to launch its sixth-generation EPYC 9006 family, code-named Venice.

The new lineup includes processors aimed at conventional enterprise and cloud workloads alongside high-density AI host systems, with configurations reaching 256 cores.

That CPU story is becoming more relevant as the industry moves toward agentic AI.

GPUs remain the primary engines performing AI computation, but increasingly complex inference environments require substantial CPU capacity around those accelerators. Agents retrieve information, call tools, manage state, interact with applications and coordinate workflows before and after GPU execution.

In other words, the growth of AI does not eliminate conventional compute. At large enough scale, it creates more of it.

AMD sees that dynamic as an opportunity to defend and expand EPYC's position just as Nvidia pushes its Arm-based Vera CPU into the data center.

It also introduces another architectural question for operators: whether CPUs and GPUs increasingly arrive as tightly paired components of an AI platform rather than separate procurement decisions.

Helios Gets Hyperscale Validation

AMD’s latest financial results add some commercial weight to that momentum. The company reported $6.7 billion in Data Center revenue for the second quarter, up 107% year over year, driven by strong demand for EPYC processors and Instinct GPUs. CEO Lisa Su said AMD entered the second half of 2026 with EPYC demand accelerating, Instinct deployments scaling and Helios beginning to ramp — an important distinction as AMD moves from announcing its rack-scale architecture to proving it can deploy it at volume.

The bigger hurdle for AMD has historically been less about producing competitive silicon than establishing an ecosystem capable of challenging Nvidia's enormous installed base and CUDA software advantage.

Helios is beginning to assemble some meaningful customers.

Microsoft plans to deploy the platform in Azure for frontier-model inference, Azure AI services and customer workloads.

Anthropic has committed to deploy up to 2 GW of AMD Instinct MI450-series GPUs using Helios systems, with the first gigawatt expected to begin deployment in the first half of 2027.

Meta is validating Helios and sixth-generation EPYC systems for future deployments, while Cerebras is working with AMD on an inference architecture pairing Helios with its wafer-scale systems.

And the biggest validation predates the July launch. OpenAI and AMD have already laid out a multi-generation agreement covering as much as 6 GW of AMD GPU capacity, beginning with MI450-series infrastructure.

Taken together, those commitments move Helios beyond the realm of an alternative architecture looking for customers.

AMD now has to prove that those commitments can translate into repeatable deployment at scale.

A Second AI Factory Architecture

Nvidia still occupies the commanding position in AI infrastructure, and its advantage extends well beyond GPU performance. CUDA, NVLink, networking, DGX systems and a broad partner ecosystem have allowed Nvidia to increasingly define the AI factory as an integrated system.

AMD's response with Helios is revealing.

Rather than trying to win the market one GPU at a time, AMD is assembling its own full-stack proposition around Instinct accelerators, EPYC CPUs, Pensando networking, ROCm software and open rack-scale hardware.

And the cadence is accelerating. AMD plans another Instinct generation and Helios platform in 2027, followed by MI600-generation hardware in 2028.

For data center developers and operators, that gives the industry something it has been seeking as AI infrastructure scales into hundreds of megawatts and gigawatts: another serious architectural path.

The next phase of the GPU battle will therefore be fought well outside the GPU.

It will play out across racks, networks, power systems, cooling plants and entire AI campuses — where the ability to turn silicon roadmaps into deployable capacity may ultimately matter as much as whose chip wins the benchmark.

 

At Data Center Frontier, we talk the industry talk and walk the industry walk. In that spirit, DCF Staff members may occasionally use AI tools to assist with content. 

Keep pace with the fast-moving world of data centers and cloud computing by connecting with Data Center Frontier on LinkedIn, following us on X/Twitter and Facebook, as well as on BlueSky, and signing up for our weekly newsletters using the form below.

About the Author

Matt Vincent

Matt Vincent is Editor in Chief of Data Center Frontier, where he leads editorial strategy and coverage focused on the infrastructure powering cloud computing, artificial intelligence, and the digital economy. A veteran B2B technology journalist with more than two decades of experience, Vincent specializes in the intersection of data centers, power, cooling, and emerging AI-era infrastructure. Since assuming the EIC role in 2023, he has helped guide Data Center Frontier’s coverage of the industry’s transition into the gigawatt-scale AI era, with a focus on hyperscale development, behind-the-meter power strategies, liquid cooling architectures, and the evolving energy demands of high-density compute, while working closely with the Digital Infrastructure Group at Endeavor Business Media to expand the brand’s analytical and multimedia footprint. Vincent also hosts The Data Center Frontier Show podcast, where he interviews industry leaders across hyperscale, colocation, utilities, and the data center supply chain to examine the technologies and business models reshaping digital infrastructure. Since its inception he serves as Head of Content for the Data Center Frontier Trends Summit. Before becoming Editor in Chief, he served in multiple senior editorial roles across Endeavor Business Media’s digital infrastructure portfolio, with coverage spanning data centers and hyperscale infrastructure, structured cabling and networking, telecom and datacom, IP physical security, and wireless and Pro AV markets. He began his career in 2005 within PennWell’s Advanced Technology Division and later held senior editorial positions supporting brands such as Cabling Installation & Maintenance, Lightwave Online, Broadband Technology Report, and Smart Buildings Technology. Vincent is a frequent moderator, interviewer, and keynote speaker at industry events including the HPC Forum, where he delivers forward-looking analysis on how AI and high-performance computing are reshaping digital infrastructure. He graduated with honors from Indiana University Bloomington with a B.A. in English Literature and Creative Writing and lives in southern New Hampshire with his family, remaining an active musician in his spare time.

You can connect with Matt via LinkedIn or email.

You can connect with Matt via LinkedIn or email.

Sign up for our eNewsletters
Get the latest news and updates
Aree_S/Shutterstock.com
Source: Aree_S/Shutterstock.com
Sponsored
Justin Loritz of Rehlko says it's time to shift away from calendar-driven servicing and reactive repairs in favor of more targeted, data-informed maintenance.
SNEHIT PHOTO/Shutterstock.com
Source: SNEHIT PHOTO/Shutterstock.com
Sponsored
A paused data center project is not a project on hold — it's a project in transition. Bill Tierney of The ProLift Rigging Company explains what it means for rigging and installation...