top of page

F5 AI Load Balancing Hits 3.24x in a Lab, but Only Under Extreme Pressure

7 days ago
12 min read

F5 AI load balancing reportedly delivered 3.24 times more completed work than an Envoy-based gateway during the most demanding test in an F5-sponsored lab visit. The result came from BIG-IP Next for Kubernetes running on NVIDIA BlueField-3 data processing units, or DPUs. Yet the lighter workload showed a much smaller difference. That contrast matters more than the headline number.

The test used Supermicro servers equipped with eight NVIDIA H100 GPUs each. Every GPU served the Qwen3-32B model using FP8 numerical precision. BIG-IP Next for Kubernetes managed traffic through the DPUs, while the comparison gateway ran on host processors.

This was not a broad verdict on every Kubernetes gateway or inference cluster. It was a specific comparison under controlled conditions, reported by ServeTheHome after a sponsored visit. Still, it exposes an increasingly important contest: traffic routing based on live GPU conditions versus routing that remains largely detached from accelerator state.

The central question is no longer whether a cluster owns enough GPUs. It is whether its software can keep those expensive accelerators productive when requests become long, concurrent, and difficult to place.

The F5 BIG-IP Next for Kubernetes Test Put Routing Under Stress

The reported advantage emerged when the cluster’s key-value cache became oversubscribed, not during the easier baseline.

The lab test compared two paths serving the same Qwen3-32B model. F5’s control and Endpoint Picker components ran on BlueField-3 DPUs. The alternative used Envoy AI Gateway, since renamed Agent Router, on the host.

The test tracked P90 latency across 60-minute runs. NVIDIA’s AI Perf Tool generated requests at different concurrency levels and prompt lengths. Four traffic patterns covered requests with no shared prefix, multi-turn conversations, mixed traffic, and heavy prefix reuse.

The baseline combined 150 concurrent requests with 10,000 input tokens per request. ServeTheHome reported that the two gateways performed relatively closely in that case. The cluster used 46 percent of its available key-value cache capacity.

A key-value cache, usually shortened to KV cache, stores attention data that a model can reuse while generating tokens. It improves inference speed but consumes substantial GPU memory. Long prompts and many simultaneous users can push that memory resource beyond comfortable limits.

The demanding test increased concurrency from 150 to 200, a 33 percent rise. It also doubled each request’s input length to 20,000 tokens. The resulting workload demanded 1.24 times the cluster’s available KV cache capacity.

That oversubscription changed the result. ServeTheHome reported a 3.24x gain for the F5 path because it distributed the difficult workload more effectively. Its published charts also examined completed requests, output tokens per second, and time to first token.

The 3.24x figure should therefore be read as a high-pressure result. It does not mean every cluster will process 3.24 times more traffic after installing F5 software. The same report calls that number an extreme within the tested setup.

That qualification does not make the test irrelevant. Production inference systems must survive peaks, long contexts, and uneven accelerator loads. A gateway that behaves similarly at low utilization can become far more valuable near saturation.

The important change is architectural. Load balancing has moved closer to the model-serving state, including queue depth, GPU utilization, and cache pressure. The gateway is no longer making decisions from network connections alone.

F5 calls BIG-IP Next for Kubernetes, or BNK, an AI service plane. It sits between clients and GPU infrastructure while combining traffic management, security, routing, and usage controls. The product can run on host processors or supported BlueField-3 DPUs.

Placing it on a DPU shifts networking and security tasks away from the server’s main processors. The DPU is a programmable infrastructure processor designed to handle networking, storage, and security work. That leaves host resources available for model serving and cluster operations.

The test therefore measured two connected ideas. One was routing quality under GPU memory pressure. The other was moving infrastructure work onto hardware designed to process it outside the host.

Why F5 AI Load Balancing Improves Under Heavy Demand

F5’s mechanism depends on seeing accelerator conditions that conventional network metrics cannot describe.

Traditional load balancers can distribute traffic through round-robin, connection counts, or fixed priorities. Those methods work well when backend servers have predictable capacity. Large language model inference violates that assumption.

One request might contain a brief question. Another might include 20,000 tokens of documents and conversation history. A third might reuse a prefix already stored in one GPU’s KV cache.

These requests can produce different processing times even when they reach identical GPUs. Queues also change rapidly as models batch requests, allocate memory, and stream generated tokens. A healthy network endpoint can still be a poor destination for the next prompt.

F5’s load-balancing documentation describes an Analyzer component that watches GPU and model-serving telemetry. It recommends new traffic weights for each backend. F5’s Traffic Management Microkernel then applies those weights in the data plane.

The documented inputs include inference latency, queue depth, GPU memory consumption, thermal state, and error rates. F5 also supports telemetry from NVIDIA Inference Microservices, NVIDIA Data Center GPU Manager, and vLLM.

That feedback loop explains why the gap can widen under pressure. A static policy lacks a direct view of which GPU is approaching a memory limit. A telemetry-aware controller can reduce traffic to a struggling endpoint before its queue becomes the cluster’s bottleneck.

F5 also describes the routing as prefix-aware and KV-cache-aware. Prefix awareness tries to send related prompts toward a backend that already holds reusable context. Avoiding unnecessary cache rebuilding can reduce compute work and memory churn.

Load awareness serves a different purpose. It spreads requests according to available capacity instead of assuming every endpoint is equally ready. The strongest result should appear when those assumptions diverge, which is what the lab’s oversubscribed workload created.

The software does not make the H100 GPUs intrinsically faster. It attempts to waste less of their available processing time. That distinction is essential when evaluating claims about GPU performance.

Improved scheduling can increase total cluster throughput without changing model weights or accelerator silicon. It can also reduce the number of requests trapped behind unusually expensive prompts. However, the benefit depends on workload diversity and telemetry quality.

A uniform batch with short prompts offers fewer routing opportunities. A highly variable stream creates more chances for intelligent placement to matter. The lab results followed that pattern, with smaller differences under lighter conditions.

F5’s public documentation cites throughput improvements of 30 to 40 percent against round-robin routing. Separately, F5 said testing validated by The Tolly Group produced up to 40 percent higher token throughput. The same announcement claimed 61 percent faster time to first token and 34 percent lower overall request latency.

Those figures are more restrained than 3.24x because they describe different testing. They also remain vendor-published performance claims, even when an outside testing organization conducted the measurements. Buyers should examine the underlying configurations before comparing percentages.

F5’s system can sit in front of external model routers such as LiteLLM, RouteLLM, and NVIDIA Router. It can send a request through a model-selection layer before directing it toward a virtual address for the chosen backend.

That means BNK is not necessarily replacing every routing component. It can become the traffic and policy layer surrounding them. This broader position lets F5 connect GPU placement with security, metering, and network enforcement.

The architecture matters because inference gateways are becoming control points for scarce resources. They can decide which model handles a request, which user receives capacity, and when traffic must be slowed. A poor decision wastes more than network bandwidth.

The Real Contest Is GPU-Aware Routing Versus Opaque Backends

The pressure falls on gateways that treat every available inference endpoint as an interchangeable server.

F5’s primary opponent is not one company. It is an older traffic-management model that sees connections but not the internal condition of each accelerator. The lab used Envoy AI Gateway as the representative comparison.

The Agent Router project associated with the evolving cloud-native ecosystem reflects a wider push toward specialized AI routing. The naming and project landscape continue to change, which complicates simple product comparisons.

Envoy itself remains a widely used proxy foundation. The F5 test does not establish that Envoy cannot support smarter inference routing. It compares particular implementations, locations, policies, and configurations.

F5’s differentiation combines several layers. Its Endpoint Picker uses live telemetry for backend selection. The DPU deployment places traffic processing outside the host. Its broader platform adds controls for security, tenant isolation, and token consumption.

Moving those functions onto BlueField-3 creates a second competitive axis. A host-based gateway consumes CPU cycles and memory bandwidth on the server. A DPU-based gateway uses a dedicated processor while remaining physically close to the workload.

NVIDIA’s AI factory guide lists F5’s integration as one option for offloading proxies, load balancing, encryption, firewalling, and API protection. The same guide identifies integrations from security vendors including Fortinet and Palo Alto Networks.

That context shows why the market will not reduce to F5 against Envoy. Infrastructure suppliers are competing to place security and traffic intelligence inside the DPU layer. Open-source projects are also adding model-aware routing features.

The practical decision concerns ownership. Some operators want a commercial service plane with integrated support and policies. Others prefer composable open-source components that their platform teams can inspect, modify, and operate.

Commercial integration can reduce the work required to connect telemetry, routing, networking, and security. It can also deepen dependence on a particular vendor’s control plane and supported hardware matrix. That tradeoff becomes significant across large fleets.

Open components can offer flexibility and portability. They also require engineering teams to assemble observability, policy enforcement, routing logic, and lifecycle management. The cost of that work rarely appears in a simple throughput chart.

F5’s position is strongest where GPU fleets serve many tenants with uneven workloads. Shared infrastructure increases the need for isolation, rate limits, usage accounting, and predictable service levels. It also makes inefficient request placement more expensive.

Its position is less obvious for small or lightly loaded clusters. If endpoints rarely approach their limits, static or simpler routing can remain adequate. The added infrastructure must justify its operational footprint.

F5 says no model changes are required for its routing and DPU offload. That lowers one adoption barrier because teams can retain existing model servers. Yet deployment still involves new infrastructure components, telemetry pipelines, policies, and failure modes.

The company’s documentation says AI load balancing is disabled by default. Operators must configure the feature and its data path. They also need Prometheus and compatible telemetry when using the built-in analyzer.

Only NVIDIA GPU metrics have built-in plugin support in the current documentation. Organizations using other accelerators may need custom logic. Even NVIDIA environments can vary across model servers, networking layouts, and orchestration practices.

The hardware requirements are also specific. F5’s DPU requirements identify supported BlueField-3 hardware, minimum memory, dual-network interfaces, and required software components.

The same requirements state that the DPU must be dedicated to BNK. They warn that other DPU software can create performance problems or Kubernetes instability. Only one DPU per chassis is supported for BNK in that documented configuration.

Those constraints turn the buying decision into more than a gateway benchmark. Teams must decide how to allocate DPUs, manage firmware, integrate networking, and recover failed components. They must compare that work against the host capacity saved.

What the 3.24x Performance Claim Does Not Establish

The lab result is a useful stress signal, but it is not independent proof of a universal production advantage.

ServeTheHome explicitly disclosed that F5 sponsored the California lab visit. That transparency helps readers interpret the report, but it does not remove the need for independent reproduction.

The hardware, model, precision, prompt sizes, and request patterns were tightly defined. Each of those variables can change routing behavior. A different model or serving engine might manage cache pressure differently.

The strongest result came from the 200-concurrency, 20,000-token condition. That workload demanded 1.24 times the available KV cache. It deliberately placed the cluster beyond a comfortable resource boundary.

Such overload is valuable for exposing scheduler behavior. It can also magnify a product’s best-case differentiation. Buyers need results across normal utilization, peak utilization, and sustained overload.

The comparison also combined routing location with routing intelligence. F5 ran on a DPU, while the alternative ran on the host. Therefore, the test does not isolate the performance contribution from each design choice.

A more revealing evaluation would compare several configurations. F5 could run on the host and DPU with the same policy. Competing gateways could run with both static and telemetry-aware routing. The cluster could then expose the contribution from offload and scheduling separately.

The public article provides many charts but not every raw log or configuration detail needed for reproduction. It mentions that artificial intelligence helped transform logs into visual displays. That presentation choice increases the importance of publishing machine-readable results.

F5’s March 2026 performance announcement offers another evidence point. It reports lower gains from separate testing and attributes validation to The Tolly Group.

Multiple tests showing the same direction strengthen the mechanism’s plausibility. They do not make the percentages interchangeable. Different baselines, workloads, and success metrics can generate very different headline improvements.

Completed requests, token throughput, and latency each answer a different question. A system can produce more aggregate tokens while giving some users slower initial responses. It can reduce average latency while leaving tail latency unstable.

The lab examined mean and P99 time to first token alongside throughput. Production buyers should also study failed requests, retry rates, response quality, and fairness across tenants. Those measures reveal whether higher throughput arrives through undesirable prioritization.

Model-serving optimizations can also influence output consistency. Routing a prompt toward a smaller model may reduce resource use but change quality. F5 describes policy-based routing between larger and smaller models, although that function was not the core of this comparison.

Security functions create another measurement problem. A gateway handling encryption, firewall rules, token controls, and inspection does more work than a minimal router. Fair comparisons must align enabled features or explain their operational value.

DPU offload can preserve host resources, but those DPUs are not free capacity. They consume power, require management, and occupy part of the server architecture. The relevant economic measure is total cluster output against total infrastructure cost.

Vendor claims about “freeing GPU cycles” also require careful wording. Network services often compete directly for host CPU resources rather than executing on the GPU itself. Better routing can raise GPU utilization, but the DPU does not create new accelerator cores.

The 3.24x result is most credible as evidence of a bottleneck-management advantage in one extreme scenario. It should not become a blanket multiplier in capacity plans. Even ServeTheHome described it as being near the upper end of observed benefits.

The report offered a more modest example: a 1.25x improvement resembles receiving the output of five GPUs from a four-GPU baseline. That analogy communicates the economic stakes, but production gains will depend on each cluster.

Teams should recreate their own prompt-length distribution, concurrency curve, cache reuse, model mix, and service objectives. They should then compare consistent configurations over long runs. Short demonstrations cannot capture every operational failure.

A credible pilot should include telemetry outages and stale metrics. If the routing controller loses its view of GPU conditions, operators need to know how quickly it detects the problem. They also need a predictable fallback policy.

The test should cover DPU failure, control-plane disruption, and network partitioning. It should show whether active requests survive and whether new traffic moves safely. Performance during perfect operation represents only part of production readiness.

Three Signals Will Show Whether the Lab Advantage Travels

The next test is whether F5 can turn a compelling overload result into repeatable gains across ordinary production workloads.

The first signal is independent workload reproduction. Buyers need tests that publish raw data and complete configurations across several models, serving frameworks, and prompt distributions. Results should separate DPU offload from telemetry-driven scheduling.

Consistent improvements under moderate load would strengthen F5’s case. Benefits that appear only during deliberate cache oversubscription would narrow the addressable use case. Both outcomes would still provide useful capacity-planning information.

The second signal is broader deployment evidence. F5 and NVIDIA describe enterprises and GPU service providers as target users, but named production examples would make the operating model clearer. Useful accounts should explain cluster size, traffic variation, and observed failure modes.

Production evidence should also show whether teams keep the promised capacity gains after enabling full security controls. Token governance, encryption, tenant isolation, and auditing all add work. Their combined impact matters more than a stripped-down benchmark.

The third signal is the response from open gateway and inference-routing projects. If those projects add comparable GPU telemetry, prefix awareness, and cache-aware placement, F5’s routing advantage may become a standard capability.

That outcome would shift competition toward operational integration, DPU support, security policy, and vendor service. It would also benefit users by making AI-aware traffic management available through more deployment models.

F5 retains a meaningful position because it already combines those layers. Its platform overview frames BNK as unified Kubernetes traffic management across application delivery, security, and policy. The DPU option extends that model into AI infrastructure.

Still, platform breadth does not remove the burden of proof. Cluster operators should demand workload-specific measurements before redesigning their ingress and service planes. They should measure cost per completed request, not just peak tokens per second.

For developers, the development is a reminder that model code no longer determines inference performance alone. Request placement, cache locality, queue management, and infrastructure isolation can materially change how much work identical GPUs finish.

For enterprise buyers, the story is about utilization before expansion. A smarter control layer can be more practical than acquiring additional accelerators when power, rack capacity, or delivery schedules constrain growth.

The F5 AI load balancing result makes its strongest argument at the cluster’s worst moment. That is valuable because peak pressure often determines capacity purchases and user experience. It is also exactly where careful validation matters most.

Before adopting the 3.24x figure, reproduce the conditions that created it. Compare ordinary traffic, sustained peaks, and failure recovery using your models and policies. Then ask the decisive question: does smarter routing delay the next hardware purchase without adding more operational risk than it removes?

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page