top of page

AMD Helios Unites 72 GPUs, but Nvidia Sets the Test

Aug 12
14 min read

AMD has launched Helios as one 72-GPU system, not a loose collection of accelerator servers connected inside a cabinet. The AMD ServeTheHome architecture story matters because the company now controls nearly every major layer of its AI rack.

Helios combines Instinct MI455X accelerators, EPYC Venice processors, Pensando networking, ROCm software, and liquid-cooled power infrastructure. Broadcom supplies the merchant Ethernet switching silicon that connects the GPUs through one scale-up fabric.

That combination creates the real conflict. AMD is no longer challenging Nvidia with an accelerator card alone. It is challenging Nvidia’s rack-scale model while rejecting the closed networking approach that helped make Nvidia’s systems difficult to copy.

Helios looks credible on paper. However, peak specifications do not establish delivered application performance, software maturity, supply volume, or operating economics. Those unresolved questions will determine whether openness becomes a purchasing advantage or merely an architectural preference.

What AMD ServeTheHome Coverage Reveals Inside Helios

Helios changes AMD’s unit of competition from one GPU to an integrated AI rack.

The system contains 72 Instinct MI455X accelerators across 18 liquid-cooled compute trays. Each tray carries four GPUs and one socket for a sixth-generation EPYC 9006 processor, code-named Venice.

AMD assigns 432 GB of HBM4 memory to each MI455X. Across the rack, that produces roughly 31 TB of high-bandwidth memory available within the scale-up domain.

High-bandwidth memory, or HBM, sits close to the GPU and supplies data at much higher rates than conventional server memory. That capacity matters for large models, long contexts, inference caches, and workloads that would otherwise require more partitioning.

AMD lists up to 23.3 TB/s of memory bandwidth for each MI455X in its CDNA 5 architecture documentation. The rack aggregates those devices into a memory system with more than one petabyte per second of theoretical bandwidth.

The physical layout is as important as the accelerator count. Helios uses an Open Compute Project Open Rack Wide frame, which is wider than a conventional rack. The design provides space for dense compute trays, cabling, power equipment, and liquid-cooling hardware.

The rack includes six scale-up switch trays. Each tray contains two Broadcom Tomahawk 6 switch chips, giving Helios 12 switch chips across the system.

Each Tomahawk 6 device offers 102.4 Tbps of switching capacity. Its Ethernet lanes form the internal fabric that lets every accelerator communicate with every other accelerator through a single switch hop.

That fabric carries UALink over Ethernet, often shortened to UALoE. UALink defines a scale-up connection for accelerators, while Ethernet supplies the underlying transport and merchant switching hardware.

AMD says each MI455X receives 3.6 TB/s of bidirectional scale-up bandwidth. Across 72 devices, the aggregate figure reaches about 260 TB/s.

This is the mechanism behind AMD’s claim that Helios behaves like one 72-GPU system. Traditional clusters often divide memory ownership among separate servers, then pass data through multiple networking stages.

Helios reduces those boundaries inside the rack. It still contains distinct processors and memory devices, but its single-hop fabric gives software a more tightly connected accelerator domain.

AMD’s MI455X specifications also show how much networking moved onto the GPU package. Each accelerator module includes two enhanced I/O dies and 36 bidirectional UALoE links.

That change makes networking part of the accelerator architecture. It is no longer an accessory selected after the compute platform has been designed.

Helios also uses Pensando hardware for communication beyond the scale-up domain. Each GPU can connect to three 800 Gbps Vulcano network interface cards, providing up to 2.4 Tbps of scale-out bandwidth per accelerator.

Scale-out networking connects multiple racks into a larger deployment. A Pensando Salina data processing unit handles front-end traffic, including management, storage access, and application requests.

The rack therefore contains two distinct network layers. Broadcom silicon joins the accelerators inside Helios, while AMD Pensando devices connect Helios to storage, services, and other racks.

Power delivery completes the system. A rear 50-volt direct-current bus bar supplies the rack, and liquid cooling removes heat from its dense components.

Reporting from AMD’s Advancing AI event placed the rack’s load between 225 kW and 245 kW. That requirement makes Helios unsuitable for data halls without high-density power and liquid-cooling support.

The rack can also weigh around 5,000 pounds, according to published platform details. Buyers must treat deployment as a facility project, not a routine server refresh.

ServeTheHome’s architecture focus is therefore justified. The central product is the integration itself, including compute, memory, networking, power, cooling, mechanics, and software.

Why AMD Had to Build the Entire Rack Now

Nvidia forced every serious accelerator vendor to compete at system scale.

Modern AI workloads spend substantial time moving data among accelerators. Faster arithmetic helps only when models, activations, and intermediate results can reach the compute engines without creating long stalls.

That reality favors systems designed as coordinated units. Nvidia established this model with its NVL72 platforms, which combine 72 GPUs, CPUs, NVLink switching, networking, cooling, and software.

AMD previously sold competitive accelerators that system vendors assembled into servers and clusters. That model gave customers choice, but it left more integration work outside AMD’s direct control.

The company could improve a GPU while still losing at the system level. Network topology, collective communication, cooling limits, software tuning, and server design could erase an advantage measured on the accelerator.

Helios addresses that weakness by providing a complete reference architecture. AMD specifies the compute tray, scale-up topology, scale-out network, rack format, power delivery, cooling approach, and associated software.

The timing also reflects the arrival of CDNA 5, AMD’s latest dedicated data-center compute architecture. MI455X uses eight compute chiplets manufactured with a 2 nm process and places them above two 3 nm fabric-and-cache dies.

Two additional I/O dies manage external communication. Twelve HBM4 stacks surround the logic, while advanced packaging connects the components into one accelerator module.

This chiplet design lets AMD optimize different functions with different manufacturing processes. Compute density, memory interfaces, cache, and networking do not all need the same silicon characteristics.

MI455X contains 320 billion transistors and 256 work group processors. A work group processor organizes execution resources that process groups of AI and scientific-computing operations.

AMD targets low-precision AI calculations with formats such as MXFP4, MXFP6, MXFP8, and FP8. These formats represent model values with fewer bits, increasing throughput when an application can preserve acceptable accuracy.

The accelerator reaches a listed peak of 40.3 petaflops for MXFP4 operations. Across the rack, AMD advertises up to 2.9 exaflops of peak FP4 performance and 1.4 exaflops at FP8.

Those figures are theoretical peaks, not measured results for a complete model. However, they explain why rack integration arrived alongside MI455X.

A single accelerator now moves enough data and consumes enough power that its surrounding system determines whether applications can use the available arithmetic. AMD needed Helios to expose CDNA 5’s capabilities at meaningful scale.

The company also acquired ZT Systems in 2025, gaining engineering experience in hyperscale rack design. AMD later separated the manufacturing business, limiting direct competition with established server partners.

That transaction gave AMD deeper system expertise without requiring it to become the sole Helios supplier. HPE, Supermicro, cloud providers, and other partners can build products from the reference design.

Microsoft has announced plans to deploy Helios in its data centers. HPE previously committed to adopting the architecture, giving AMD routes into both hyperscale and enterprise infrastructure.

AMD said at CES that Helios would serve as a blueprint for much larger AI systems. Its rack-scale preview connected the design to training, inference, and future multi-rack deployments.

The pressure therefore extends beyond Nvidia. Server manufacturers must decide how much AMD engineering to adopt, while cloud providers must decide whether an open reference platform reduces integration risk.

Chip suppliers also face a new purchasing pattern. Customers increasingly evaluate accelerators as parts of complete systems, not as interchangeable cards with isolated benchmark scores.

Open Ethernet Is AMD’s Main Challenge to Nvidia

The primary contest is AMD’s open Ethernet rack against Nvidia’s vertically controlled NVLink platform.

Helios resembles Nvidia’s Vera Rubin NVL72 in several important ways. Both place 72 accelerators across 18 liquid-cooled compute trays, and both use dedicated switch trays for high-bandwidth communication.

The difference lies in control of the scale-up network. Nvidia builds NVLink and NVLink switching into its platform, giving the company direct authority over the protocol, silicon, topology, and software integration.

AMD uses an open accelerator link carried through Ethernet. Broadcom provides the Tomahawk 6 switch silicon, while UALink defines how accelerators communicate across that fabric.

This choice allows AMD to use merchant networking technology rather than a proprietary switch architecture. System vendors and hyperscalers can work with familiar Ethernet tools, suppliers, and operational practices.

Openness does not mean every component can be exchanged without engineering work. Scale-up networks have tight latency, congestion, reliability, synchronization, and software requirements.

However, published interfaces give partners more room to customize the rack. A cloud provider can adjust networking, management, or deployment choices without depending on one vendor for every layer.

Broadcom’s role makes that claim more concrete. Tomahawk 6 supplies 512 lanes at 200 Gbps, with enough capacity to feed the 72 GPUs at AMD’s stated scale-up rate.

The Helios architecture report shows why this partnership matters. AMD’s hardware portfolio is broad, but it does not need to manufacture every component to control the system design.

This creates a coalition-based alternative to Nvidia. AMD contributes the GPUs, CPUs, Pensando network devices, software, and platform engineering. Broadcom contributes the central scale-up switching silicon.

HPE and other manufacturers can then turn that blueprint into systems. Cloud operators can deploy those systems while retaining more influence over networking and software choices.

Nvidia offers the opposite proposition. Its tighter integration reduces the number of external variables and gives customers a platform optimized by one vendor.

That control can simplify performance tuning. Nvidia can coordinate GPU behavior, switch silicon, communication libraries, drivers, networking, and application frameworks through a single roadmap.

AMD is betting that open standards can reach comparable efficiency without requiring one company to own every link. That proposition must survive actual workloads, not only topology diagrams.

The two systems also differ outside the rack. Helios assigns three 800 Gbps Vulcano interfaces to each accelerator, reaching 2.4 Tbps of scale-out bandwidth per GPU.

Published Vera Rubin configurations pair each GPU with one 1.6 Tbps ConnectX-9 interface. AMD therefore claims 50 percent more scale-out bandwidth per accelerator.

That comparison favors workloads spanning several racks, assuming software can use the links efficiently. Large training jobs and distributed inference services depend on this layer when one rack cannot hold the entire workload.

AMD also claims more memory capacity and bandwidth. Helios offers 31 TB of HBM4 across the rack, while AMD says that represents 50 percent more capacity than the competing Nvidia platform.

Memory capacity can reduce model partitioning and leave more space for key-value caches. A key-value cache stores attention data so an inference system can generate later tokens without recomputing earlier context.

That advantage is particularly relevant to long-context inference and models serving many concurrent requests. Yet capacity alone does not determine latency or throughput.

Software scheduling, kernel quality, communication efficiency, and workload shape still control how much useful performance reaches customers. Nvidia’s CUDA environment remains the established reference point for many AI teams.

ROCm has improved across recent Instinct generations, and major frameworks now support AMD hardware. Helios still places greater responsibility on AMD to make 72 accelerators operate predictably as one platform.

The open-versus-controlled contest is therefore not philosophical. It is a measurable question about whether a partner-based architecture can match an integrated platform’s delivered performance and reliability.

The Specifications Do Not Settle Performance

AMD’s strongest numbers remain vendor claims until independent rack-level testing confirms them.

AMD claims Helios provides 15 percent more peak FP4 performance per accelerator than the leading competing system. It also projects up to 30 percent better token economics.

Those claims require careful framing. Peak FP4 arithmetic describes the fastest supported low-precision operation under ideal conditions, not the sustained speed of a deployed model.

A token-per-dollar comparison requires even more assumptions. Hardware utilization, electricity, cooling, software licensing, staffing, model accuracy, batch size, and system availability all affect the result.

AMD has not disclosed public pricing in a form that supports a neutral ownership comparison. Buyers will also negotiate hardware, support, networking, and deployment agreements at different scales.

Rack-level figures complicate the performance story. AMD lists 2.9 exaflops of peak FP4 performance for Helios, while Nvidia has published a higher 3.6-exaflops rack figure for Vera Rubin NVL72.

The definitions behind those numbers may differ. Nvidia’s result can incorporate compression behavior suited to some inference tasks, while AMD emphasizes raw supported precision rates.

The Register’s rack comparison noted this distinction. Some workloads can benefit from Nvidia’s adaptive compression, while others need calculations that align more closely with AMD’s uncompressed figure.

Neither comparison represents a universal winner. Training, fine-tuning, dense inference, mixture-of-experts models, and long-context serving stress hardware differently.

A mixture-of-experts model activates only selected groups of parameters for each token. It can reduce computation, but it also creates demanding communication patterns among GPUs.

Helios may perform well when memory capacity or scale-out bandwidth limits an application. Nvidia may retain an advantage when its software stack extracts more work from lower nominal resources.

The same caution applies to AMD’s single-hop memory language. Helios provides a tightly connected HBM domain, but it does not turn 72 physical memory pools into one conventional uniform-memory device.

Software must still understand data placement, accelerator ownership, synchronization, and communication costs. A remote HBM access across a switch does not behave like a local access inside one GPU package.

Latency also matters alongside bandwidth. AMD publishes impressive aggregate throughput, but detailed application results will show how the fabric behaves under contention and irregular traffic.

Reliability creates another challenge. A 72-GPU rack combines accelerators, processors, switches, network interfaces, cooling connections, power components, cables, and software into one operating domain.

Failures become more expensive when workloads treat the rack as one system. Operators need fault isolation, telemetry, checkpointing, serviceability, and predictable recovery procedures.

The double-width rack presents facility constraints as well. Its 225 kW to 245 kW power range exceeds the capacity of many existing enterprise data halls.

Liquid cooling is mandatory at this density. Buyers need suitable coolant distribution, power conversion, floor loading, maintenance access, and trained operations teams.

These requirements do not weaken Helios relative to every competitor. Nvidia’s rack-scale systems create similar infrastructure demands.

They do limit the addressable market, however. Helios initially belongs in hyperscale data centers, specialized AI facilities, national laboratories, and sites designed for dense liquid-cooled equipment.

Software remains the largest uncertain variable. ROCm must support the rack’s topology while matching the usability and performance developers expect from mature Nvidia deployments.

Customers will need stable collective communication, optimized kernels, framework integration, observability, orchestration, and rapid support for new model architectures.

AMD controls more of that path than before. Owning the processor, accelerator, network interfaces, and reference system gives its engineers more opportunities to eliminate cross-vendor problems.

Yet control does not instantly create maturity. The first production deployments will expose issues that presentation benchmarks and reference designs cannot anticipate.

Helios Pressures Buyers as Much as Nvidia

Helios gives infrastructure buyers a second rack-scale path, but it also makes their evaluation work more demanding.

A credible alternative can improve negotiating leverage. Hyperscalers no longer need to compare an integrated Nvidia rack with an assortment of independently assembled AMD servers.

They can compare two 72-GPU rack architectures with similar physical ambitions. Both arrive as coordinated platforms for training and inference at data-center scale.

That makes procurement comparisons more meaningful. Buyers can examine memory per rack, scale-up bandwidth, scale-out bandwidth, power, cooling, software readiness, serviceability, and workload results.

The AMD ServeTheHome deep dive also highlights the importance of supply-chain choice. Broadcom switching and an open rack standard create more room for manufacturers to differentiate their implementations.

HPE can combine Helios with its own system engineering and Juniper networking expertise. Supermicro can target customers that already operate its liquid-cooled platforms.

Cloud providers can expose MI455X capacity through managed services, reducing the need for smaller customers to install the physical racks themselves.

AMD gains reach from this model, but it also depends on partners to execute consistently. Poor integration by one manufacturer could damage perceptions of the broader platform.

Nvidia’s controlled design limits that variation. Customers receive fewer architectural choices, but they also benefit from a more uniform target for application tuning and support.

Enterprise buyers should therefore avoid treating openness as an automatic cost reduction. Customization creates value only when the organization has the engineering capacity to use it.

A company running standard models through a managed cloud service may care more about delivered tokens and availability than the underlying switch supplier.

A hyperscaler building custom networking and scheduling software may assign much more value to published interfaces and merchant silicon.

Researchers working with models that exceed one server’s memory can benefit from Helios’s 31 TB HBM domain. The same capacity can support larger inference caches and more concurrent requests.

Developers should care because hardware diversity affects software portability. A second viable rack architecture gives framework projects and model vendors stronger reasons to optimize outside CUDA.

That process will not happen automatically. Application teams must validate kernels, numerical behavior, communication libraries, container images, monitoring tools, and deployment workflows.

The broader industry also benefits from a standards contest. UALink and Ultra Ethernet now have a high-profile system where their performance can be measured under demanding AI workloads.

Success would encourage more switch vendors, accelerator designers, and system builders to participate. Weak results would strengthen the case for proprietary integration.

AMD is also placing pressure on itself. Annual accelerator releases require the rack design, networking, software, and manufacturing chain to advance on the same schedule.

A delayed component can hold back the entire system. The platform is only ready when accelerators, CPUs, switches, network interfaces, cooling, firmware, drivers, and frameworks work together.

The company says Helios has entered production, with shipments expected by the end of the third quarter of 2026. That schedule puts it near Nvidia’s Vera Rubin rollout rather than a full generation behind.

Timing narrows AMD’s traditional disadvantage. It also removes excuses if software or supply fails to meet customer expectations.

The production announcement reported planned deployments from major partners. The next test is whether those commitments become accessible production capacity.

A rack installed for evaluation is not the same as a fleet serving revenue-generating workloads. Buyers should request repeatable results over weeks, not short demonstrations under controlled conditions.

Three Signals Will Decide Whether Helios Works

Shipment volume, independent workload results, and multi-rack reliability will decide whether Helios changes the market.

The first signal is production delivery. AMD expects Helios shipments by the end of the third quarter, so customers should watch which systems arrive before September closes.

Named installations matter, but quantities matter more. Multiple operating fleets would show that AMD and its partners can obtain HBM4, package MI455X devices, assemble racks, and commission liquid-cooled sites.

Delays would weaken AMD’s timing argument. Nvidia’s installed base and production experience become more valuable with every quarter that Helios remains scarce.

The second signal is independent application benchmarking. Reviewers need complete rack results for current language models, mixture-of-experts workloads, fine-tuning jobs, and long-context inference.

Useful tests should report latency, throughput, power, utilization, accuracy, and failure behavior. Peak arithmetic alone cannot show whether 72 GPUs remain busy during a real workload.

Benchmarks should also compare software effort. A platform that requires weeks of custom kernel work may deliver attractive hardware results while increasing project risk.

The most informative comparisons will use matched models, batch sizes, numerical formats, and service-level targets. Otherwise, vendors can select settings that favor their architecture.

Results should cover memory-intensive workloads as well as compute-intensive ones. Helios’s capacity and bandwidth claims become stronger if applications can use both without excessive communication overhead.

The third signal is reliable multi-rack scaling. One Helios rack tests UALink over Ethernet, while larger deployments also test Pensando Vulcano interfaces and the surrounding Ultra Ethernet network.

AMD must show that performance remains predictable as jobs cross rack boundaries. Congestion, collective operations, job scheduling, and component failures become harder at that scale.

Large customers will watch how quickly a failed accelerator or network link can be isolated. They will also measure whether a job can continue, restart, or recover from a recent checkpoint.

Strong multi-rack results would support AMD’s open-network thesis. They would show that merchant switching, open specifications, and partner engineering can deliver a coordinated AI system.

Weak scaling would favor Nvidia’s tighter integration. It would suggest that control over the entire fabric still provides an operational advantage that specifications cannot capture.

These signals should shape how readers interpret future AMD ServeTheHome reports. New diagrams and peak figures will be useful, but production evidence now carries more weight than architectural intent.

Helios is not just another Instinct server. It is AMD’s attempt to combine its accumulated data-center assets into one competitive product.

The MI455X supplies low-precision compute and 432 GB of HBM4. Venice supplies host processing, Pensando provides scale-out networking, and ROCm connects the system to AI frameworks.

Broadcom’s Tomahawk 6 silicon supplies the central link among the 72 accelerators. Open Rack Wide hardware gives manufacturers a common physical foundation.

That combination makes Helios AMD’s most complete challenge to Nvidia’s AI infrastructure strategy. It also exposes AMD to a more demanding standard.

Customers will no longer judge the company only by accelerator specifications. They will judge rack availability, application performance, software quality, facility requirements, reliability, and support as one package.

The next question is practical: will independent operators reproduce AMD’s claims after Helios reaches production data centers? Follow the first large deployments, compare complete workload results, and watch whether open Ethernet remains efficient beyond one rack.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page