Amazon EKS MoE Reinforcement Learning Gets 40% More Throughput, but the Benchmark Has Limits
Amazon says Amazon EKS MoE reinforcement learning produced 40% more aggregate rollout throughput after its engineers enabled DeepEP over Elastic Fabric Adapter. The test covered 48 P5en instances, with 16 assigned to policy training and 32 generating inference rollouts. That is a substantial result for an expensive stage of large-model post-training.
The important change is not simply another faster GPU configuration. Amazon has adapted DeepEP’s expert-communication path to use libfabric, giving DeepEP v2 native support for AWS networking. That integration targets the irregular traffic pattern created when a Mixture-of-Experts model routes tokens among experts on different GPUs.
The result also needs careful framing. AWS disclosed a relative throughput gain for one super-sparse MoE workload, not a universal benchmark for every model, cluster, or reinforcement learning framework. Independent systems research has documented the difficulty of moving expert-parallel communication across different GPU and network architectures.
What Amazon Changed in Its MoE Reinforcement Learning Stack
Amazon’s reported gain came from changing how routed tokens travel between experts, not from adding more machines to the measured cluster.
The AWS architecture divides a reinforcement learning job into several worker groups. GPU nodes handle policy training, reward-model inference, and rollout generation. CPU nodes run environments and preprocessing, while memory-optimized nodes hold experience buffers and checkpoint caches.
Amazon EKS acts as the orchestration layer. It places containers, manages separate node groups, coordinates failures, and lets operators scale each part of the workload independently. Amazon S3 stores datasets, checkpoints, final weights, and other durable artifacts outside the latency-sensitive execution path.
That separation matters because reinforcement learning is not one uniform computation. During rollout generation, an inference worker uses the current policy to interact with an environment and produce candidate responses or trajectories. A reward system evaluates those samples, and the policy-training workers consume the resulting experience.
The updated weights then return to the rollout fleet, beginning another policy iteration. This loop appears in Reinforcement Learning from Human Feedback, or RLHF, which uses preference signals to improve a model. It also appears in Group Relative Policy Optimization, or GRPO, which evaluates outputs relative to other samples in a group.
Each stage stresses infrastructure differently. Rollout generation resembles distributed inference and can often divide work among independent workers. Policy training requires tighter synchronization because participating GPUs must complete coordinated operations before the next step can proceed.
The architecture therefore separates 32 inference instances from 16 training instances in the reported test. Across those 48 P5en systems, Amazon says DeepEP over EFA raised aggregate rollout throughput by 40% compared with the configuration without DeepEP.
P5en instances use NVIDIA H200 GPUs and high-bandwidth AWS networking. A fully populated 48-instance deployment represents hundreds of accelerators, although AWS reports its larger architecture can extend toward roughly one thousand accelerators. The published percentage describes the 48-instance comparison, not every possible cluster size.
The model itself is described only as a super-sparse MoE model. A Mixture-of-Experts model contains multiple specialized feed-forward blocks but activates only a subset for each token. Sparsity reduces per-token computation, yet it creates a demanding routing problem when experts reside on different GPUs.
Standard collective operations work well when every rank exchanges predictable blocks of similar size. MoE routing is different. Tokens choose experts dynamically, so traffic can be sparse, uneven, and composed of many small transfers.
That difference explains why infrastructure matters so much. More theoretical model capacity does not automatically deliver more useful tokens per second. If expert dispatch and result collection overwhelm the network, expensive GPUs wait for activations instead of processing them.
Why Amazon EKS MoE Reinforcement Learning Hits a Network Wall
Sparse computation saves arithmetic, but expert parallelism can return that cost as communication delay.
Expert parallelism distributes a model’s experts across GPUs. When a router selects remote experts, the system must dispatch each token’s activation to the correct device. After the expert processes it, a combine operation returns the output to its original execution path.
Those exchanges happen repeatedly throughout the model. Their destinations depend on routing decisions made at runtime, and different experts can receive different numbers of tokens. The network must therefore handle many fine-grained transfers without allowing a few busy destinations to stall every participant.
DeepEP was created for this pattern. The DeepEP project provides specialized dispatch and combine kernels for expert-parallel workloads. It uses NVLink for communication inside a server and an RDMA-capable transport between servers.
Remote Direct Memory Access, or RDMA, lets one machine transfer data directly to another machine’s memory with less CPU involvement. That shorter path can reduce software overhead and make high-speed network hardware more useful.
Elastic Fabric Adapter, or EFA, is AWS’s low-latency network interface for tightly coupled computing. The EFA documentation describes an operating-system bypass path built on AWS Scalable Reliable Datagram. EKS can expose EFA devices to pods running distributed machine learning applications.
Inside each P5en instance, NVLink and NVSwitch carry GPU traffic across the local accelerator fabric. For cross-instance transfers, EFA becomes the relevant path. Amazon’s integration uses libfabric, an interface that lets applications access different high-performance network providers through a common API.
Amazon says its engineers contributed features that moved DeepEP communication primitives from a CUDA-specific RDMA backend to libfabric. With that work, DeepEP v2 can send inter-node data over EFA while retaining specialized kernels for expert dispatch and combine.
The distinction between specialized expert operations and dense collectives is central to the result. NCCL remains useful for regular operations such as all-reduce, all-gather, and reduce-scatter. DeepEP targets the sparse all-to-all exchanges around MoE layers.
Recent research reflects this division. The authors of NCCL EP describe separate low-latency and high-throughput modes for expert communication. Their high-throughput design aggregates data within NVLink domains before transmitting it across inter-node RDMA connections.
That hierarchy reduces the amount of fine-grained traffic crossing the slower boundary between machines. It also recognizes that a cluster is not one uniform network. Communication within a server has different bandwidth and latency characteristics from communication across servers.
The AWS implementation follows the same broad principle. Local traffic stays on NVLink, while libfabric carries cross-node DeepEP traffic over EFA. This topology-aware path replaces a generic treatment of every token transfer.
The resulting 40% increase refers to aggregate rollout output, not merely a communication microbenchmark. That end-to-end measure is valuable because a faster kernel does not always accelerate the entire reinforcement learning loop. The gain suggests expert communication was important enough to affect completed rollout work.
However, rollout throughput is still only one layer of the system. Policy iteration time also depends on environment execution, reward evaluation, sample buffering, checkpoint publication, training computation, and weight synchronization. Optimizing one stage can expose a bottleneck elsewhere.
The Real Contest Is Specialized Routing Versus Generic Collectives
The primary contest is between communication designed for dynamic expert routing and collective operations designed for regular data movement.
Generic collectives are attractive because they are mature, broadly supported, and easier to integrate. They work across many training frameworks and hardware configurations. Operators can also test them with familiar tools and reason about their synchronization behavior.
MoE traffic violates several assumptions that make those collectives efficient. Every token can select a different set of experts. Some experts become temporarily popular, message sizes stay small, and the system performs dispatch and combine operations at every MoE layer.
A conventional implementation can package this traffic into all-to-all operations. That approach remains functional, but synchronization and message-handling overhead grow as expert parallelism spans more nodes. More GPUs then create more communication relationships rather than proportionally more useful computation.
DeepEP attacks this problem with kernels built around the semantics of expert routing. The dispatch kernel sends token activations to selected experts. The combine kernel returns processed activations, while avoiding work that a general collective might perform for unused destinations.
The design also seeks to overlap communication with computation. If a GPU can continue useful matrix operations while transfers progress, some network time disappears from the critical path. That overlap becomes harder when communication requires repeated CPU coordination or strict global synchronization.
Amazon’s libfabric migration matters because the original optimization was closely associated with NVIDIA GPUs and InfiniBand-style networking. A communication library that performs well on one fabric does not automatically retain its behavior on another. Ordering guarantees, message initiation, and device interfaces vary.
The integration therefore represents more than changing a network address. DeepEP’s assumptions must map onto EFA’s transport semantics, and the implementation must preserve correct token delivery. It must also avoid introducing enough software overhead to erase the benefits of specialized routing.
Amazon states that supported P5 and P6 systems can use GPUDirect RDMA with EFA. GPUDirect RDMA allows network transfers to read from and write to GPU memory without staging every payload through ordinary host memory. The operating system remains outside the primary data path.
This design places pressure on generic MoE deployments that rely only on standard collectives. Infrastructure teams using large expert-parallel models now have evidence that a specialized path can improve one production-relevant reinforcement learning workload.
The result also pressures framework maintainers. DeepEP support must reach serving engines, reinforcement learning systems, container images, schedulers, and observability tools. A fast transport that requires a fragile custom build can lose its advantage during deployment or recovery.
NCCL 2.31 adds another part of the picture. AWS says that release includes newer EFA optimizations for dense collective communication. A realistic MoE training stack therefore uses different mechanisms for different traffic classes rather than declaring one universal winner.
DeepEP handles irregular expert dispatch and combine. NCCL continues handling dense synchronization around attention layers, tensor parallelism, data parallelism, and optimizer state. EFA carries both classes across machines through paths optimized for their respective patterns.
This division is the larger architectural lesson. MoE scaling depends on identifying communication by shape and purpose. Treating every transfer as interchangeable leaves performance available on the table.
What the 40% DeepEP Throughput Claim Does Not Establish
The benchmark supports a specific architecture decision, but it does not establish a universal 40% gain for DeepEP over EFA.
Amazon identifies the instance allocation, the relative improvement, and the model’s broad sparsity profile. It does not publish the model’s parameter count, expert count, routing distribution, sequence lengths, batch sizes, or complete baseline configuration.
Those details directly affect expert communication. A model activating more experts per token can generate more traffic. Larger batches can combine messages more efficiently, while small decode batches can magnify fixed latency.
The phrase “aggregate rollout throughput” also needs context. AWS does not provide the absolute number of output tokens, trajectories, or completed requests per second in the public post. Readers cannot calculate the cluster’s total utilization or compare it directly with another provider.
The baseline matters just as much. “Without DeepEP” could mean a standard NCCL all-to-all implementation with particular tuning choices. Different message aggregation, expert placement, concurrency, or routing policies might narrow or widen the measured gap.
Amazon reports a controlled result from its own internal workload. The company does not claim the benchmark was independently audited, and the public material does not include repeated-trial variance. The correct language is therefore that AWS says throughput increased by 40%.
There is also a portability question. Earlier UCCL-EP research argued that expert-communication systems tied closely to GPU and network interfaces create substantial integration work. The paper specifically examined how differing ordering semantics complicate support for EFA and other non-InfiniBand networks.
That research predates Amazon’s newly described native EFA work. It remains relevant because it explains the technical barrier that AWS says it has now addressed through libfabric contributions. The two accounts describe different points in a rapidly changing implementation history.
UCCL-EP takes another route. It keeps routing decisions on GPUs but delegates network execution to multithreaded CPU proxies, using a control channel to bridge hardware differences. Its authors report gains on NVIDIA plus EFA systems, but those tests involve their own models, frameworks, and configurations.
Neither result invalidates the other. They show that transport design can change the outcome, and that “EFA support” does not identify one fixed execution path. Operators need to know whether a build uses GPU-initiated transfers, CPU proxies, message aggregation, or another compatibility layer.
DeepEP’s own published requirements and performance results have also evolved. Current project documentation reports strong bandwidth on supported RDMA configurations, yet it encourages users to benchmark larger expert-parallel deployments directly. That advice is especially important on cloud fabrics with different topology and congestion behavior.
Cluster scale introduces further uncertainty. The reported test used 48 P5en instances, while AWS discusses scaling the broader architecture toward roughly one thousand accelerators. A design that performs well across 48 nodes does not necessarily retain the same efficiency at every larger scale.
Network contention can emerge when multiple worker groups share infrastructure. Token routing can become more imbalanced as model or workload behavior changes. A single slow rank can also delay tightly synchronized training operations.
Reinforcement learning adds its own source of variation. Prompt lengths, response lengths, environment latency, sampling settings, and reward-model complexity all affect how much time rollout workers spend communicating. A 40% gain on a communication-heavy workload can shrink when generation or environment execution dominates.
The result says even less about online serving. Production inference often uses smaller batches and strict per-request latency targets. A high-throughput kernel tuned for rollout generation does not automatically reduce time to first token or time per output token for interactive users.
Cost remains unstated as an absolute measure. Higher throughput on the same cluster usually improves useful work per accelerator-hour, but the post provides no total training bill. It also does not compare the optimized configuration with alternative instance types or network libraries.
These omissions do not make the result unimportant. They define where it is useful. The benchmark is evidence that AWS’s DeepEP integration can remove a meaningful bottleneck in one large MoE reinforcement learning pipeline.
EKS and Spot Capacity Change the Rest of the RL System
The communication gain becomes operationally useful only when the scheduler, buffer, storage, and failure model keep the faster rollout fleet supplied.
Amazon EKS allows the architecture to assign different node types to distinct tasks. GPU node groups can scale around training and inference demand. CPU groups can expand for environment workers, while memory-oriented systems absorb short-lived experience data.
This heterogeneity is especially relevant to GRPO and RLHF. Rollout workers can generate large amounts of temporary data, but policy trainers consume it in synchronized batches. If production and consumption rates diverge, one side waits while the other accumulates a queue.
A shared in-memory experience buffer decouples those rates over short periods. Rollout workers publish completed samples, and trainers pull batches when ready. Checkpoint caches help distribute updated weights without forcing every transfer through durable object storage.
Amazon S3 serves a different role. It holds datasets, recoverable checkpoints, completed model artifacts, and final weights. Keeping that durable path outside the most frequent sample exchange prevents object-storage latency from controlling every training step.
This separation also clarifies the value of EKS. Kubernetes is not accelerating matrix multiplication or expert kernels. It is coordinating the collection of services required to keep the accelerators productive.
EKS manages placement, restarts, scaling policies, and node-group boundaries. It can schedule stable policy-training capacity separately from more elastic rollout workers. That boundary supports Amazon’s second optimization, using EC2 Spot Instances for some rollout generation.
Spot capacity can be interrupted when AWS needs the underlying instances back. That risk is difficult for tightly synchronized policy training because losing one worker can stall or restart a coordinated job. Rollout tasks are easier to partition and retry.
Amazon recommends giving rollout workers bounded units of work and publishing samples frequently. When an interruption notice arrives, a worker can drain active requests and return unfinished tasks to a queue. Other workers continue without restarting the complete policy-training group.
The strategy does not make interruption free. Lost partial generations waste some compute, and replacement nodes need containers, model weights, and communication libraries. Autoscaling decisions must also consider queue depth, model-loading time, and available Spot capacity.
Still, the topology isolates two failure domains. Policy trainers remain on stable capacity, while rollout generation uses a cheaper but less predictable pool. That design matches the different synchronization requirements of the two stages.
The 40% DeepEP throughput increase can alter this balance. Faster inference workers may deliver experience more quickly than trainers consume it. Operators then need to resize node groups, adjust batch scheduling, or reduce inference capacity to avoid paying for idle production.
The opposite can happen after a policy update. Weight distribution and worker restart time can temporarily starve the experience buffer. A useful production dashboard must therefore track end-to-end policy iteration, not just tokens generated per second.
Teams also need reproducible build information. DeepEP, NCCL, CUDA, libfabric, EFA drivers, framework versions, and GPU architecture all influence the data path. Changing one component can silently select a slower fallback.
This operational evidence should live with model and experiment records. Engineering teams can preserve configuration decisions, benchmark notes, and failure reports in a searchable technical knowledge base. That practice becomes valuable when a later image rebuild changes throughput without changing the model.
Three Signals Will Show Whether the Gain Generalizes
The next test is reproducibility across models, cluster sizes, and complete policy iterations.
The first signal is a public benchmark package with absolute throughput. Useful results would include tokens or trajectories per second, latency distributions, expert-load imbalance, network utilization, and repeated-run variance.
That package should specify the baseline collective, all relevant software versions, and the precise DeepEP transport path. It should also disclose model dimensions, active experts per token, batch sizes, prompt lengths, response lengths, and expert-parallel degree.
If independent teams reproduce a similar gain, the AWS claim becomes stronger. If results vary widely, the integration remains useful but workload-specific. Either outcome would help operators decide when its added complexity is justified.
The second signal is scaling efficiency beyond the published 48-instance configuration. Results at several cluster sizes would show whether throughput grows proportionally or loses ground to synchronization, congestion, and expert imbalance.
A meaningful scaling study should hold the workload definition constant while increasing resources. It should report both aggregate output and per-accelerator efficiency. Aggregate throughput alone can rise even while every added GPU contributes less useful work.
Strong efficiency toward roughly one thousand accelerators would support AWS’s broader architectural claim. A sharp decline would indicate that DeepEP removed one bottleneck while another emerged at larger scale.
The third signal is end-to-end policy iteration time under real failures. Rollout throughput matters because training workers need fresh experience, not because generating isolated tokens is the final objective.
Future measurements should include environment execution, reward evaluation, buffer delays, policy updates, checkpoint publication, and weight redistribution. They should also show how Spot interruptions affect completed samples and recovery time.
A shorter full iteration would confirm that the communication optimization improves reinforcement learning progress rather than shifting idle time elsewhere. If iteration time barely changes, teams should investigate training, storage, or synchronization before adding more rollout capacity.
Amazon EKS MoE reinforcement learning now has a credible route for combining Kubernetes orchestration, EFA networking, and specialized expert communication. The reported 40% increase makes that route worth testing, but it remains a starting measurement rather than a transferable constant. Infrastructure teams should reproduce the comparison with their own model, routing profile, and RL loop before standardizing the stack. The practical question is not whether DeepEP can produce a faster chart. It is whether the same cluster completes more validated policy updates, at acceptable reliability and cost, after every part of the system is counted.



