AMD EPYC Venice Benchmarks Challenge Nvidia, but the Footnotes Matter
AMD has published detailed AMD EPYC Venice benchmarks claiming its 256-core server CPU delivers more than twice Nvidia Vera's integer throughput.
The company also says a 96-core, high-frequency Venice processor offers about 20 percent greater per-core performance than Nvidia's 88-core CPU. Those claims expand the competitive case AMD introduced when it unveiled its sixth-generation EPYC family in July 2026.
The headline numbers give AMD a strong answer to Nvidia's attempt to redefine the data center CPU around agentic AI. However, AMD labels the relevant SPEC CPU 2026 figures as internal estimates, not independently published results.
That distinction shapes the real story. AMD is challenging Nvidia across total throughput, per-core speed, memory bandwidth, and AI infrastructure design. Nvidia, meanwhile, argues that completed AI sessions matter more than core counts or aggregate benchmark scores.
What AMD's New Benchmark Numbers Actually Claim
AMD's new disclosure turns a broad performance promise into several specific, testable claims.
The central comparison uses SPECrate 2026 Integer Base. This benchmark measures throughput by running multiple copies of integer workloads across a system.
AMD estimates a score of 2,070 for a dual-socket EPYC 9996 system. The processor has 256 Zen 6c cores per socket, giving the tested server 512 physical cores.
For a dual-socket Nvidia Vera system, AMD lists an estimated score of 925. Each Vera processor contains 88 custom Arm cores, creating a 176-core server configuration.
Those figures give EPYC 9996 about 2.24 times Vera's estimated integer throughput. AMD describes the result as roughly 2.2 times greater performance.
The comparison favors AMD's highest-density design, but it is not a core-for-core contest. A dual-socket EPYC 9996 machine contains almost three times as many physical cores as the Vera system.
AMD therefore published another comparison involving its 96-core high-frequency Venice processor. Its estimate gives that dual-socket system a score of 1,210, against 925 for Vera.
The resulting per-core figures are 6.3 for Venice and 5.3 for Vera. AMD presents this as approximately 1.2 times Vera's per-core performance.
Both processors used GCC 15.2 for this particular comparison, according to AMD's disclosure. That detail matters because compiler versions and optimization settings can materially affect CPU benchmark results.
The latest figures also sharpen AMD's generational story. The 256-core EPYC 9996 reportedly scores about 78 percent higher than the 192-core EPYC 9965 in the same throughput comparison.
That gain does not come from architecture alone. Venice increases core density, memory capability, and platform power alongside the transition from Zen 5 to Zen 6.
AMD's server CPU overview identifies every SPEC CPU 2026 number as a preliminary engineering projection. The company says those estimates remain subject to change.
The distinction between an estimate and a submitted result is critical. SPEC permits estimated performance metrics, but it requires vendors to label them clearly as estimates.
No independently reviewed SPEC submission currently settles the AMD versus Nvidia comparison. The new figures are more detailed than AMD's earlier claims, but they remain vendor projections.
That makes these AMD EPYC Venice benchmarks valuable as a statement of competitive intent. It does not make them a final verdict on shipping systems.
Why AMD EPYC Venice Benchmarks Put Nvidia on Defense
AMD is attacking the idea that Nvidia should control both the accelerator and host-processor layers of an AI server.
Nvidia built its data center position around GPUs, networking, and an increasingly integrated software stack. Vera extends that strategy into the general-purpose CPU market.
The 88-core Arm processor is designed to accompany Nvidia's Rubin generation of accelerators. It also gives Nvidia more control over memory behavior, data movement, and system-level optimization.
AMD has a different commercial interest. It wants cloud providers and enterprises to continue treating the host CPU as an independent purchasing decision.
A credible Venice advantage would help preserve that choice. It would also let AMD sell EPYC processors into servers built around accelerators from several vendors.
This is why AMD emphasizes workloads beyond model training and inference. AI systems also route requests, retrieve documents, query databases, execute code, call APIs, and manage storage.
Those stages often run on CPUs. They can determine total response time even when a GPU performs the model's main reasoning step.
AMD describes agentic AI as a systems workload with rapidly changing demands. Its testing covers web serving, retrieval, databases, tool execution, and response processing.
The company's 256-core strategy targets parallel work across those services. More cores can support additional requests, isolated runtime environments, database transactions, or simultaneous tool calls.
The high-frequency 96-core design addresses a different concern. It attempts to preserve per-core speed for tasks that cannot spread efficiently across hundreds of cores.
AMD's portfolio therefore challenges Nvidia on two fronts. The EPYC 9996 emphasizes throughput, while the high-frequency model addresses latency-sensitive execution.
The platform specifications strengthen that pitch. EPYC 9006 supports up to 256 cores and 512 threads per socket, according to AMD's Zen 6 platform details.
It also supports 16 DDR5 memory channels, MRDIMMs rated for 12,800 megatransfers per second, and PCIe 6 connectivity. AMD claims up to 1.6 terabytes per second of socket memory bandwidth.
Vera uses eight channels of LPDDR5X memory through SOCAMM2 modules. Nvidia lists approximately 1.2 terabytes per second of bandwidth.
Memory capacity, latency, power, and bandwidth do not reduce to one specification. Still, AMD's bandwidth claim weakens the argument that a conventional server CPU must concede this category to Nvidia.
AMD also has a deployment advantage rooted in familiarity. EPYC uses the x86 software environment already common across enterprise and cloud infrastructure.
Vera introduces an Arm platform into environments where software validation, monitoring, virtualization, and procurement processes may center on x86. Arm adoption is substantial, but migration costs vary by workload.
Nvidia can offset that friction through tightly integrated rack systems. Customers buying a complete Nvidia AI platform may value unified hardware and software more than CPU interchangeability.
AMD is betting that many buyers will resist that consolidation. The EPYC Venice performance story offers them a technical reason to keep the CPU layer open.
Core Density Meets Nvidia's Critical-Path Argument
The primary disagreement is not simply AMD versus Nvidia; it concerns which CPU behavior determines useful AI infrastructure performance.
AMD presents concurrency as a defining requirement for agentic systems. One user request can generate retrieval operations, database lookups, policy checks, code execution, and independent sub-agents.
Those operations create bursts of parallel work. A high-core-count CPU can absorb more of that work without forcing every service into a separate machine.
Nvidia focuses on the sequential dependencies between those bursts. An agent cannot take its next step until required tool calls or sub-tasks return.
That waiting chain is the critical path, meaning the sequence that controls total completion time. Faster individual threads can shorten it even when the system has spare parallel capacity.
Nvidia's own Vera performance argument claims the processor delivers 1.5 times Venice's per-core performance across four selected agent-oriented tests.
Those tests include compiler, static-analysis, Python, and memory-sensitive workloads. Nvidia compares Vera with AMD's dense 256-core EPYC 9996, rather than the 96-core high-frequency Venice model.
AMD selects a different comparison. Its strongest per-core claim pits the high-frequency 96-core Venice configuration against Vera in the complete SPECrate Integer suite.
Both vendors are choosing a workload and competing processor that support their architectural message. Nvidia targets AMD's dense flagship, while AMD answers with a frequency-optimized model.
This does not automatically invalidate either analysis. Server processors are designed for different operating points, and customers frequently compare multiple models within one product family.
However, the selections prevent a simple declaration that one CPU has universally better per-core performance. The answer changes with the Venice model and workload set.
Total throughput presents a similar complication. EPYC 9996 wins AMD's estimate partly because it places 256 cores in each socket.
That density has real value when software can keep the cores busy. It matters less when an application depends on a small number of fast threads.
A database fleet, web-serving cluster, or sandbox service can often exploit broad parallelism. A serial orchestration stage may respond more strongly to per-thread latency.
Many production systems contain both patterns. Requests arrive concurrently, but each request also contains steps that must run in order.
This mixed behavior makes fleet economics more important than one peak score. Buyers must consider completed jobs, tail latency, memory use, software licensing, rack power, and utilization.
AMD's EPYC Venice performance case is strongest when a workload scales across many cores. Nvidia's case strengthens when sequential latency dominates user experience.
Neither vendor has yet provided a neutral, end-to-end test covering representative agent sessions across comparable shipping systems. Until that arrives, architecture-specific benchmarks show direction rather than universal superiority.
The contest also extends beyond individual nodes. AMD estimates that Venice can deliver greater rack-level performance under a fixed power budget.
Its earlier methodology modeled a 100-kilowatt rack and reduced the number of Venice nodes to account for higher node power. The company still calculated higher aggregate throughput.
That analysis depends on performance projections and modeled system power. Real racks include cooling, networking, storage, accelerators, power conversion, and utilization effects.
The most useful question is therefore not which vendor wins a normalized bar chart. It is which system completes the required work within a buyer's latency, power, and software constraints.
What the Benchmark Footnotes Change
The AMD versus Nvidia Vera comparison remains provisional because several results combine estimates, vendor testing, and nonidentical configurations.
SPEC CPU offers a controlled workload suite, but valid interpretation still requires configuration details. Processor count, compiler, memory, threading, firmware, and optimization flags can all affect the score.
The SPEC reporting rules allow estimated performance results. They require each estimate to carry a clear label rather than hiding that status in a general disclaimer.
AMD follows that principle in its detailed footnotes. Its 2,070, 1,210, and 925 figures are all estimates.
The absence of a submitted result means readers cannot yet inspect a standard SPEC result page for every setting. That limits independent reproduction.
AMD's white paper also includes comparisons that use different compiler releases. Some Venice systems use AMD's AOCC 5.1 compiler, while certain competing systems use GCC 13 or Intel OneAPI.
Compiler choice is part of real platform performance. Vendors often optimize software for their own processors, and customers may deploy those optimized toolchains.
However, mixed compilers make architectural conclusions less direct. A score can reflect the processor, compiler maturity, flags, libraries, or interaction among all four.
Memory configurations differ as well. AMD's newer systems can use 16 channels of fast MRDIMMs, while the compared Vera system uses eight SOCAMM2 channels.
Those are platform choices, not accidental deviations. Yet they mean the benchmark evaluates complete system configurations rather than isolated CPU cores.
The EPYC 9996 also operates with a 600-watt default CPU power specification in AMD's disclosed Redis testing. That exceeds the power levels associated with many earlier server processors.
Higher power does not make a performance result irrelevant. Data center operators routinely accept higher socket power when consolidation improves throughput per rack.
It does require a broader calculation. Cooling density, electrical provisioning, idle behavior, and accelerator power can determine whether consolidation benefits survive deployment.
AMD's agentic workload tests introduce another limitation. The company uses recognizable components such as NGINX, FAISS, database workloads, and a replayed multi-persona agent.
Some tests derive workloads from TPC-H or TPC-C. AMD notes that these derived results are not comparable with officially published scores from those benchmark families.
That methodology can still reveal useful behavior across AMD's own controlled systems. It should not be presented as a standardized industry ranking.
The comparison with AWS Graviton5 adds cloud variables. AMD tested Graviton5 through a publicly available bare-metal instance rather than an equivalent local reference system.
Cloud firmware, storage, networking, and platform services can influence results. AMD discloses that those factors may affect the comparison.
Independent reporting has highlighted these caveats. The original benchmark analysis notes that AMD mixes data sources and sometimes uses different GCC generations.
The safest interpretation is narrow. AMD estimates that specific Venice configurations lead the compared systems under its stated workloads and assumptions.
That is more meaningful than an unsupported marketing slogan. It is less conclusive than independently reproduced tests on generally available hardware.
Why AI Data Centers Care Beyond the CPU Charts
The strategic value of Venice depends on whether its density improves real infrastructure economics, not whether it wins every isolated benchmark.
AI clusters dedicate most attention and capital to accelerators. CPUs still coordinate the surrounding work that keeps those accelerators productive.
A retrieval-augmented application must find documents, filter permissions, transform data, and assemble context before inference begins. Those steps place pressure on memory, storage, networking, and general-purpose compute.
An agent may also create temporary code environments or browser sessions. Each isolated task consumes CPU time and memory while the model waits for a result.
High core density can consolidate these supporting services. Fewer servers may provide the same aggregate capacity when workloads scale efficiently.
Consolidation can reduce rack space, networking ports, operating-system instances, and management overhead. It can also concentrate failures and increase cooling requirements.
The software layer decides how much hardware density becomes useful capacity. Schedulers must place work effectively, while applications must avoid locks, memory bottlenecks, and overloaded shared services.
Memory bandwidth becomes important as more cores request data concurrently. AMD's 16-channel design attempts to feed a much larger core population without proportionally increasing stalls.
Cache capacity also supports the density argument. EPYC 9996 includes one gigabyte of L3 cache, distributed across its chiplet architecture.
Large aggregate cache can reduce memory traffic for suitable workloads. It does not guarantee low access latency across every core and data location.
Nvidia's monolithic design emphasizes consistent access and strong loaded per-core performance. The company argues that this balance reduces variability across agent workflows.
These competing philosophies create practical purchasing choices. A cloud provider can prioritize maximum tenant density, while an enterprise may value predictable application latency.
Software licensing can alter the calculation again. Products licensed per core may penalize a 256-core deployment, even when its total throughput is higher.
Other applications license by server, user, or consumption. Those models can reward consolidation.
Compatibility also matters. Venice extends the x86 environment used by previous EPYC systems, but its new socket, memory design, and power profile still require qualified platforms.
Nvidia can simplify deployment for customers buying complete Rubin systems. Its CPU, accelerators, networking, and software arrive as parts of one integrated roadmap.
That advantage carries concentration risk. Buyers become more dependent on one vendor's pricing, availability, release schedule, and software direction.
AMD offers an alternative centered on CPU choice and standards-based infrastructure. Its manufacturing progress supports that strategy.
The company said Venice entered production ramp on TSMC's 2-nanometer process in May 2026. AMD called it the first high-performance computing product to reach that stage on the process.
Production ramp is not the same as broad availability. OEM qualification, firmware maturity, memory validation, and volume shipments determine when enterprises can deploy systems.
AMD says major OEM platforms are on track to launch, with leading cloud providers beginning deployments later in 2026. Those deployments will provide more meaningful evidence than engineering projections.
Three Signals Will Decide the AMD Versus Nvidia Vera Race
Independent results, production availability, and real agent throughput will determine whether AMD's benchmark lead survives outside its white paper.
The first signal is a complete set of published SPEC CPU 2026 results. Comparable submissions should disclose compilers, flags, memory, firmware, power settings, and system topology.
A verified 2.2-times throughput advantage would reinforce AMD's density argument. A narrower result would show that preliminary estimates overstated Venice's lead.
Per-core submissions matter just as much. The most informative comparison would include EPYC 9996, the 96-core high-frequency model, and Vera under aligned software conditions.
The second signal is system availability at meaningful scale. AMD has begun production, but customers need qualified servers, stable firmware, supported memory, and consistent supply.
Cloud instances will make Venice easier to test without a major hardware purchase. Their launch timing and available configurations will reveal which models reach customers first.
Nvidia faces the same test with Vera. Product claims carry limited operational weight until customers can obtain systems and measure sustained behavior.
Broad availability would strengthen the vendor that converts specifications into dependable capacity first. Delays would give rivals more time to improve hardware, compilers, and pricing.
The third signal is completed agent work under fixed constraints. Tests should measure successful user sessions, tail latency, throughput, power, and accelerator utilization together.
Those results must include both narrow critical paths and wide bursts of parallel tools. Otherwise, the test simply repeats one vendor's preferred architectural story.
Watch how systems behave when retrieval, databases, APIs, code sandboxes, and model inference operate at once. That mixture better represents production agent services than a single CPU suite.
For infrastructure teams, the right next step is not to accept either company's normalized chart. Define the workload, latency target, software stack, rack limit, and licensing model first.
Then compare shipping systems with identical data and service-level objectives. The AMD EPYC Venice benchmarks give buyers a serious reason to include AMD in that evaluation.
They also give Nvidia a clear target. If Vera can outperform Venice on completed AI sessions despite having fewer cores, Nvidia's critical-path argument gains credibility.
If Venice converts its projected throughput into lower infrastructure cost and strong latency, AMD will have challenged more than one CPU. It will have challenged Nvidia's plan to own the entire AI server platform.



