top of page

AMD Nvidia Rivalry Sharpens as EPYC 9996 Challenges Vera and Xeon

AMD has detailed a 256-core EPYC 9996 processor that challenges Nvidia Vera and Intel Xeon with unusually aggressive performance claims. The amd nvidia contest now extends beyond accelerators into the CPUs that organize AI workloads, feed GPUs, and run agent tools.

The Zen 6 processor, code-named Venice, combines 256 Zen 6c cores, 512 threads, and 1,024 MB of L3 cache. AMD says it delivers up to 3.4 times the performance of Intel's 128-core Xeon 6980P in a selected AI benchmark.

AMD also claims 20 percent higher per-core performance than Nvidia Vera in a preliminary comparison. That result matters because Nvidia presents Vera as an AI-first host CPU for its Rubin systems, not merely another general-purpose server chip.

These figures have not received broad independent validation. They come from AMD-controlled tests, selected workloads, and comparisons that sometimes use different Venice configurations. The specifications are concrete, but the competitive verdict remains open.

EPYC 9996 Turns Venice Into a Full Server Portfolio

AMD is not launching one showcase processor. It is dividing Venice into distinct platforms for dense cloud computing, enterprise servers, and AI host systems.

The leading EPYC 9996 uses 256 Zen 6c cores and supports 512 simultaneous threads. Zen 6c is AMD's density-focused core design, built to place more cores within practical socket and power limits.

The processor runs at a 2.55 GHz base frequency and boosts to 4.1 GHz, according to the disclosed specifications. Its thermal design power reaches 600 watts, placing it firmly in the high-density data center category.

That power envelope supports 1,024 MB of native L3 cache. L3 is the processor's shared last-level cache, which reduces trips to slower system memory for frequently accessed data.

The cache is not AMD's vertically stacked 3D V-Cache. AMD instead gives each 32-core Zen 6c chiplet access to 128 MB of L3, equivalent to 4 MB per core.

AMD has confirmed that Zen 6 uses TSMC's 2-nanometer-class N2 manufacturing process. Its EPYC 9006 specifications describe up to 16 memory channels and MRDIMM speeds reaching 12,800 MT/s.

MRDIMMs, or multiplexer combined-rank memory modules, increase effective data rates by combining transfers from multiple memory ranks. AMD claims this configuration supplies 1.6 TB/s of memory bandwidth per socket.

That is a major architectural change from EPYC Turin. Turin supports 12 memory channels, DDR5 speeds up to 6,400 MT/s, and roughly 576 GB/s of bandwidth in the cited comparison.

Venice also introduces 128 PCIe 6.0 lanes in a single-socket SP7 server. Dual-socket systems can expose 160 lanes, giving system builders more capacity for accelerators, networking, and high-speed storage.

Core density is only one part of the lineup. Standard Zen 6 models reach 128 cores, while high-frequency models stop at 96 cores and boost to 5 GHz.

The 96-core EPYC 9686F and 64-core EPYC 9586F both reach that 5 GHz peak. Those processors target workloads where latency and per-core speed matter more than maximum thread density.

AMD plans a smaller SP8 platform with configurations ranging from eight to 128 cores. SP8 reduces memory support to eight channels but retains 128 PCIe 6.0 lanes.

A separate processor called Verano targets AI host-node duties. AMD says it will combine up to 72 cores, 5 GHz peak clocks, and 24-channel LPDDR5X memory using SOCAMM2 modules.

Venice-X is expected later, with 3D V-Cache for cache-sensitive technical workloads. Reported plans point to 96 cores, 1,152 MB of total L3 cache, and a 5.15 GHz boost frequency.

This segmentation reveals AMD's larger strategy. It wants one architecture to cover cloud density, conventional enterprise applications, high-performance computing, and CPU-heavy portions of AI systems.

That breadth creates the central tension. Nvidia is optimizing Vera around a tightly integrated AI platform, while AMD is betting that a flexible x86 portfolio remains more attractive.

The AMD Nvidia Fight Is Moving to the Host CPU

The important contest is no longer limited to whose GPU performs more matrix operations. The host CPU increasingly determines whether expensive accelerators stay productive.

AI agents perform more than model inference. They run Python code, query databases, call external tools, search vector indexes, compile software, and manage many concurrent execution environments.

Those steps often happen on CPUs. A slow host can leave GPUs waiting for data, orchestration decisions, or results from tools outside the model.

Nvidia designed Vera around that bottleneck. The processor uses 88 custom Olympus cores, supports 176 threads, and supplies up to 1.2 TB/s of LPDDR5X memory bandwidth.

Its memory sits in SOCAMM modules, compact memory packages designed for high bandwidth and lower power. Nvidia also gives Vera a direct connection to Rubin GPUs through NVLink-C2C.

That connection provides up to 1.8 TB/s of coherent bandwidth between CPU and GPU. Coherent memory allows processors to share data while maintaining a consistent view of its state.

Nvidia says Vera completes selected agentic workloads 1.8 times faster than comparison x86 processors. Its Vera CPU launch also claims twice the performance per watt for certain rack-scale configurations.

AMD is answering with different strengths. EPYC 9996 has almost three times Vera's core count and 1.6 TB/s of claimed socket memory bandwidth.

However, AMD's SP7 platform connects accelerators through PCIe rather than Nvidia's proprietary CPU-to-GPU fabric. PCIe 6 offers broad compatibility, but it cannot match NVLink-C2C's advertised coherent bandwidth.

This creates a clear strategic split.

Memory architecture

  • AMD: Up to 16 DDR5 channels, 12,800 MT/s MRDIMMs, and 1.6 TB/s claimed socket bandwidth.

  • Nvidia: LPDDR5X SOCAMM memory, 1.2 TB/s bandwidth, and a focus on efficiency per core.

CPU design

  • AMD: Up to 256 x86 Zen 6c cores, plus lower-core models tuned for frequency.

  • Nvidia: 88 custom Arm-based Olympus cores designed around agentic processing and AI infrastructure.

Accelerator connection

  • AMD: PCIe 6 connectivity supports accelerators from multiple vendors.

  • Nvidia: NVLink-C2C tightly couples Vera with Rubin GPUs using a coherent interface.

Deployment model

  • AMD: A broad server portfolio spanning conventional databases, cloud services, HPC, and AI.

  • Nvidia: An integrated AI factory spanning CPU, GPU, networking, storage, and software.

Nvidia's approach can reduce integration friction for customers standardizing on its complete platform. The same vendor controls the processors, interconnects, network adapters, system designs, and much of the software.

AMD offers another kind of flexibility. Customers can deploy EPYC with AMD Instinct accelerators, Nvidia GPUs, or other PCIe devices without adopting one vertically integrated architecture.

That choice matters for cloud providers and enterprises that operate mixed fleets. Existing x86 applications can move to Venice without changing instruction sets or rebuilding entire software environments for Arm.

Vera still has substantial support. Nvidia says major cloud providers and AI companies plan to deploy it, alongside systems from Dell, HPE, Lenovo, and Supermicro.

The competitive question therefore extends beyond benchmark speed. Buyers must decide whether an open server platform or an integrated AI stack produces better results across the complete workload.

Zen 6 Makes Core Density a Bandwidth Problem

EPYC 9996 works because AMD increased memory throughput, cache capacity, and input-output bandwidth alongside the core count. Cores alone would create a crowded and inefficient processor.

A server CPU with 256 cores must keep those cores supplied with instructions and data. Otherwise, many cores spend time waiting for memory rather than completing useful work.

AMD's first answer is the larger memory subsystem. Sixteen channels provide more parallel paths to memory than Turin's 12-channel design or Intel Xeon 6980P's 12 channels.

The second answer is faster MRDIMM support. A theoretical 12,800 MT/s data rate gives each channel more transfer capacity, though real applications rarely sustain the headline maximum.

The third answer is cache. EPYC 9996's 1,024 MB of L3 provides a much larger shared working area than Intel's Xeon 6980P, which includes 504 MB.

Intel's Xeon 6980P specifications list 128 cores, 256 threads, 12 memory channels, 96 PCIe 5.0 lanes, and a 500-watt TDP.

Venice doubles that Intel processor's core count while raising socket power by 20 percent. That comparison favors AMD on theoretical core density, but it does not establish application-level efficiency.

Software must expose enough parallel work to use 256 cores. Databases, web services, virtual machines, compilation farms, and cloud containers can often do this effectively.

Other applications depend on a limited number of fast threads. Those workloads can benefit more from AMD's 5 GHz models or Intel's higher-frequency offerings than from the largest EPYC configuration.

The chiplet design adds another consideration. AMD distributes cores across multiple compute dies and connects them to two input-output dies.

This arrangement improves manufacturing yield and lets AMD build several products from reusable components. It can also create nonuniform latency when a core accesses memory or cache resources located elsewhere.

AMD's benchmark choices reflect the workloads best suited to the architecture. NGINX web serving, Redis, vector search, databases, and agent sandboxes all support substantial concurrency.

In NGINX testing with the WRK load generator, AMD reported 28,789,170 requests per second for EPYC 9996. The company measured 24,320,476 for EPYC 9965 and 10,162,179 for Intel Xeon 6980P.

AMD therefore claims a 1.2-times generation-over-generation increase and a 2.8-times advantage over Intel. These are vendor results, not neutral lab findings.

The FAISS vector-search test presents a similar pattern. FAISS is an open-source library that searches large collections of numerical embeddings for similar items.

AMD reported 751,453 queries per second for EPYC 9996. Its figures showed 472,079 for EPYC 9965 and 316,069 for Xeon 6980P.

Vector retrieval is relevant to AI systems using retrieval-augmented generation. It helps an application find documents or records before asking a language model to produce an answer.

The processor also targets code compilation and media processing. Both workloads can spread independent tasks across many cores, making dense processors attractive for shared infrastructure.

Cloud operators can consolidate more services onto each socket when utilization remains high. They can also reduce the number of licensed servers, network connections, and motherboard components required for a fixed workload.

However, consolidation increases the impact of one server failure. It also raises cooling density and places more demand on memory capacity, networking, and storage.

EPYC 9996 is therefore not simply a faster CPU. It is a bet that data centers can keep a very dense socket fed, cooled, and busy enough to justify its design.

AMD's 3.4x Claim Needs More Context

AMD's benchmark results establish a credible performance case, but they do not yet establish a universal 3.4-times advantage over Intel or Nvidia.

The largest Intel comparison comes from TPCx-AI, a suite covering data science and machine-learning workflows. Its principal result is expressed as AI use cases processed per minute.

AMD reported 5,982.91 AIUCpm for EPYC 9996. The same presentation listed 3,458.79 for EPYC 9965, 2,704.19 for EPYC 9755, and 1,750.36 for Xeon 6980P.

Those results produce the cited 3.4-times Intel advantage and a 1.7-times generation-over-generation improvement. The TPCx-AI benchmark provides a recognized framework, but implementation details still affect every result.

Compiler settings, memory configuration, software versions, storage, tuning, and server firmware can influence performance. Vendor-run comparisons also select products and workloads that support the vendor's argument.

AMD provided actual figures for several tests, which makes its claims more useful than unexplained normalized bars. Yet some combined scores remain harder to interpret.

Its enterprise-tools comparison combines TPC-H, TPC-C-style, and Redis workloads into a geometric mean. AMD claims 2.6 times Intel's throughput and 1.6 times the performance of EPYC 9965.

Another test replayed five agent personas and produced a combined throughput score. AMD reported 4.451 for EPYC 9996, compared with 2.97 for EPYC 9965 and 1.779 for Xeon 6980P.

The public results provide limited detail about what one unit of that score represents. Without a reproducible test harness, buyers cannot easily translate the figure into response time or operating capacity.

The Nvidia comparison needs even more care. AMD used Nvidia's published SPEC CPU 2026 results as its comparison baseline, then ran Venice with the same GNU 15.2 compiler.

AMD says its 256-core Zen 6c model achieved 2.2 times Vera's integer throughput. That is unsurprising in one respect because the EPYC configuration has 256 cores, while Vera has 88.

The more important claim compares per-core performance. AMD says a 96-core high-frequency Venice processor produced 1.2 times Vera's per-core result.

That figure supports AMD's argument that its lead does not come entirely from core count. Still, it compares preliminary vendor submissions rather than independently tested production systems.

AMD used a 600-watt setting for both Venice configurations. Nvidia lists configurable Vera CPU power between 250 and 450 watts in its current platform information.

Power differences matter because data centers operate within rack-level electrical and cooling limits. A faster socket does not automatically produce a faster rack when each socket consumes more power.

Nvidia emphasizes this point in its own Vera performance analysis. The company focuses on sustained per-core speed, memory bandwidth per core, latency, and performance under full load.

AMD emphasizes total throughput, x86 compatibility, and workload consolidation. Each framing is reasonable, but each highlights the measurements most favorable to its architecture.

The comparisons also omit several ownership factors. Memory capacity, software licensing, server density, cooling design, accelerator utilization, and application migration can outweigh a benchmark lead.

EPYC 9996's 600-watt TDP is not a measurement of average application power. TDP guides cooling design, while actual consumption changes with configuration and workload.

The same caution applies to Nvidia's efficiency claims. Vera's final results will depend on shipping servers, firmware, compilers, memory populations, and software maturity.

Intel's position can also change. Its newer high-density Xeon products increase core counts and memory capabilities, while future platforms will answer Venice more directly.

AMD has supplied enough detail to make EPYC 9996 a serious contender. It has not supplied enough independent evidence to declare every Intel or Nvidia alternative outclassed.

Intel Faces the Immediate Density Pressure

Nvidia defines the strategic fight, but Intel faces the clearest near-term pressure because AMD targets conventional x86 workloads with twice the cores per socket.

Xeon 6980P remains a large server processor with 128 cores, 12 memory channels, and support for MRDIMMs up to 8,800 MT/s. It also carries 504 MB of cache and a 500-watt TDP.

EPYC 9996 doubles the core count, roughly doubles the cache, expands memory to 16 channels, and moves input-output support from PCIe 5.0 to PCIe 6.0.

That combination can reduce the number of servers needed for highly parallel workloads. It also gives AMD an appealing story for organizations refreshing aging x86 infrastructure.

The comparison is less decisive for lightly threaded software. Intel can remain competitive where application tuning, accelerator support, database certification, or existing fleet management carries greater weight.

Intel also offers platform features that do not reduce to core counts. Its acceleration engines target cryptography, compression, data movement, analytics, and networking workloads.

Customers rarely replace server platforms based on one chart. Procurement teams evaluate availability, support contracts, validated configurations, software licensing, security updates, and long-term supply.

AMD's challenge is converting architectural advantages into broad shipments. A processor that leads in tests still needs motherboards, firmware, memory qualifications, cloud instances, and reliable volume.

Nvidia approaches the market differently. Vera is not primarily trying to displace every x86 server. It is trying to control the CPU role inside AI factories built around Nvidia accelerators.

That creates pressure on both incumbent x86 vendors. If Nvidia succeeds, some CPU purchasing decisions move from general server teams into a complete Nvidia platform order.

AMD's answer is to make EPYC relevant to agentic AI before that transition becomes established. Its July 2026 messaging repeatedly connects Venice to web serving, vector search, databases, code execution, and agent orchestration.

This positioning is deliberate. These workloads sit around the model and can consume substantial CPU resources even when GPUs perform the main inference calculations.

AWS adds another competitive route through Graviton. AMD's presentation included Graviton5 results across web serving, vector search, data-heavy AI, and enterprise workloads.

Graviton can deliver economic and efficiency benefits inside AWS, but customers cannot deploy it across their own data centers. EPYC offers broader portability across cloud and on-premises environments.

Arm server CPUs still weaken x86's historical assumption of software compatibility as an unbeatable advantage. Many Linux cloud applications now run comfortably on either instruction set.

Vera builds on that progress while adding Nvidia's software and accelerator relationships. Its adoption by major system manufacturers gives buyers more options than Nvidia's earlier tightly packaged systems offered.

AMD must therefore fight on two fronts. It needs to preserve x86's role in general servers while preventing Nvidia from defining AI host CPUs as a separate category.

EPYC 9996 gives AMD a credible weapon because it combines extreme density with familiar software. The unresolved question is whether its broad design beats processors optimized around narrower AI environments.

What Buyers Should Watch After the AMD Nvidia Claims

The next verdict will come from independent systems, rack-level measurements, and real software deployments, not another round of vendor-selected slides.

The first signal is reproducible benchmark data from shipping EPYC 9996 servers. Independent reviewers and cloud providers need to test identical software across Venice, Vera, and current Xeon platforms.

Those tests should disclose compiler settings, memory configurations, power limits, firmware, and cooling. They should also report latency distributions, not only total throughput.

A strong result across databases, compilation, vector search, and web services would support AMD's broad-platform argument. Large variation between workloads would weaken any universal performance claim.

The second signal is rack-level efficiency. Buyers should compare completed work within fixed power, cooling, and floor-space limits.

A 600-watt processor can still improve rack efficiency if it replaces multiple lower-density sockets. It can lose that advantage if memory, networking, or cooling prevents full utilization.

Nvidia's vertically integrated systems deserve the same scrutiny. NVLink-C2C can improve data sharing, but customers need evidence that those gains raise complete application throughput.

The third signal is actual deployment. Cloud instance availability, OEM shipment volume, customer references, and production software support will show whether the hardware moves beyond controlled demonstrations.

Vera's second-half availability gives Nvidia an opportunity to establish AI-focused systems before every Venice variant reaches the market. AMD's existing x86 ecosystem gives it a faster path for conventional applications.

Buyers should also examine workload placement carefully. A CPU-heavy agent system can benefit from more cores, but a GPU-bound service might gain little from the largest EPYC model.

The ideal choice can change even within one AI application. Retrieval, orchestration, code execution, inference, and storage may each favor different hardware characteristics.

Organizations evaluating the amd nvidia competition should start with traces from their own production workloads. Request rates, memory stalls, GPU idle time, and tail latency reveal where infrastructure actually waits.

Synthetic benchmarks remain useful because they isolate architectural differences. They become misleading when a normalized score replaces the operational metric that customers care about.

EPYC 9996 changes the server discussion by making 256 cores, 1 TB of cache, and 16 memory channels part of one socket. That is a substantial engineering statement regardless of the final competitive ranking.

AMD's boldest claims remain claims. Nvidia's efficiency narrative also awaits wider independent testing, while Intel retains established deployments and another product response.

The useful question is not whether AMD, Nvidia, or Intel wins every benchmark. It is which platform removes the most expensive bottleneck in a specific data center.

Before selecting a system, ask vendors for full configurations and repeatable test methods. Then measure complete workload throughput, rack power, latency, and accelerator utilization under realistic load.

That evidence will determine whether EPYC 9996 reshapes AI infrastructure or simply wins a carefully chosen set of charts.

Get started for free

A local first AI Assistant w/ Personal Knowledge Management

For better AI experience,

remio only supports Windows 10+ (x64) and M-Chip Macs currently.

​Add Search Bar in Your Brain

Just Ask remio

Remember Everything

Organize Nothing

bottom of page