AMD Nvidia Benchmark Fight Narrows as Official EPYC Results Challenge Vera's Lead
NVIDIA presented Vera as much as 50 percent faster in selected comparisons, but the AMD Nvidia benchmark fight changes when official EPYC Turin results enter the calculation.
NVIDIA’s July 2026 Vera CPU material compares its new 88-core processor with AMD’s 128-core EPYC 9755. However, NVIDIA used internal measurements for Vera and an estimated result for the AMD system. ServeTheHome replaced that AMD estimate with published SPEC CPU 2026 data and found a much narrower contest.
That normalization does not turn Vera into a weak processor. It changes the question buyers should ask. Vera still appears to lead some performance-per-core comparisons, while AMD retains stronger options for total socket throughput. The real conflict is between vendor-selected framing and workload-matched purchasing decisions.
NVIDIA’s Vera Numbers Look Different With Official AMD Results
The most important change is not a new processor result, but a better baseline for interpreting NVIDIA’s existing claim.
NVIDIA released new architectural details and internal SPEC CPU 2026 measurements for Vera on July 21. SPEC CPU 2026 is a standardized collection of processor workloads designed to compare compute performance across systems. The company positioned the results around agentic AI, where CPUs handle code execution, retrieval, orchestration, databases, and other tasks surrounding GPU inference.
Vera uses 88 custom Olympus cores and supports 176 hardware threads through NVIDIA Spatial Multithreading. NVIDIA reported a two-socket SPECrate 2026 integer base score of 925. SPECrate measures throughput by running multiple workload copies, making it sensitive to core count, memory behavior, compiler choices, and system configuration.
The comparison in NVIDIA’s whitepaper used AMD’s EPYC 9755, a 128-core Turin processor launched in 2024. NVIDIA estimated the competing AMD system’s performance rather than using the strongest published, reviewed SPEC submission available for that processor.
That choice matters. A published two-socket EPYC 9755 system recorded a SPECrate 2026 integer base score of 1,070 in the official SPEC result. NVIDIA’s estimate was about 16 percent below that reviewed result, according to the subsequent benchmark normalization.
The official score also places the comparison in an unexpected position. Vera’s reported score of 925 trails the 1,070 score from the dual-socket EPYC 9755 system in total integer throughput. NVIDIA still has fewer cores, so the result supports its per-core argument. It does not support a simple conclusion that Vera is the faster processor in every meaningful sense.
NVIDIA’s charts emphasized ratios against the estimated AMD baseline. Once the official result replaces that estimate, the apparent distance between the processors contracts. ServeTheHome calculates that Vera reaches roughly 5.3 points per core in this comparison. AMD’s performance-oriented EPYC 9575F lands around 4.8 to 5.0 points per core.
That leaves Vera with an estimated per-core advantage of roughly 6 to 10 percent against AMD’s high-frequency Turin options. This is materially different from comparisons suggesting advantages of 50 percent or more.
The distinction is central to the AMD Nvidia debate. A processor can lead per core while losing per socket. It can also lead in selected latency-sensitive work while offering less throughput for parallel services. Each statement can be accurate, but none describes the entire purchasing decision.
NVIDIA has not submitted the Vera figure as an independently reviewed SPEC result. Its July measurements remain internal company data. That does not make the number invalid, but it places it in a different evidentiary category from a published SPEC submission.
The normalized result therefore creates the article’s central tension. NVIDIA appears to have built a competitive Arm server CPU with strong individual cores. Its marketing comparison makes that advantage look broader than the available reviewed data supports.
The AMD Nvidia Comparison Depends on Which Unit Buyers Value
Vera looks strongest when the unit of comparison is an active core, while EPYC looks stronger when the unit is a complete throughput-oriented socket.
NVIDIA designed Vera around a specific theory of agentic computing. An AI agent does more than generate tokens. It may retrieve context, execute Python code, query a database, invoke an external service, inspect the result, and repeat the loop.
Many of those steps are sequential. A slow tool call can hold up the entire response, even when hundreds of other tasks run elsewhere in the system. NVIDIA argues that strong loaded single-thread performance, meaning the speed of each thread while the socket is busy, therefore deserves more weight.
The company’s Vera architecture details describe a wide, out-of-order Olympus core. Out-of-order execution allows a processor to rearrange independent operations instead of waiting for every instruction to finish sequentially.
Olympus includes a 10-wide decode engine, advanced branch prediction, and a neural predictor for difficult branch patterns. These features target interpreters, compilers, agent runtimes, graph software, and other code with irregular control flow.
That is a coherent design goal. It also explains why NVIDIA prefers a per-core or per-thread lens. If an agent’s progress depends on one branch-heavy execution path, adding more slower cores does not necessarily reduce the latency of that path.
AMD’s Turin portfolio was designed for a wider set of server roles. It includes models optimized for high frequency, balanced performance, and maximum core density. Comparing Vera with only one EPYC model obscures those differences inside AMD’s own lineup.
The EPYC 9575F is a particularly important counterpoint. It has 64 Zen 5 cores, a maximum boost frequency of 5.0 GHz, and a 400-watt default thermal design power. AMD positions it for workloads that value high per-core speed, including host processing beside accelerators.
The EPYC 9755 takes another route. Its 128 cores increase total parallel throughput, but each core receives a smaller share of the socket’s power and memory resources. That makes it an imperfect reference for claims centered on single-thread performance.
The EPYC 9965 moves even farther toward density. AMD’s EPYC specifications list 192 cores, 384 threads, 384 MB of L3 cache, and a 500-watt default thermal design power. A published two-socket system using that processor has reached a SPECrate 2026 integer base score of 1,230.
That system-level result exceeds both Vera’s reported 925 and the EPYC 9755 result. It does not mean each 9965 core is faster. It shows why a single ratio cannot settle a portfolio comparison.
A cloud provider running thousands of small web services may value the number of requests completed per rack. A platform serving interactive agents may value consistent latency for each sandbox. A database buyer may care about software licensing per core. An AI system builder may care about keeping expensive GPUs supplied with work.
Each buyer is measuring a different economic unit. The processor, socket, node, rack, licensed core, and completed agent task are not interchangeable.
This creates pressure on both companies. NVIDIA must show that its per-core gains improve complete agent workloads, rather than only selected CPU benchmarks. AMD must show that its greater core density does not introduce latency or memory bottlenecks under the irregular workloads NVIDIA emphasizes.
The result is more useful than declaring one vendor the winner. It identifies the point where architecture becomes operational economics. Buyers need to select the unit that matches their service before they select the benchmark.
Memory Bandwidth Makes the Vera Lead Real but Conditional
Vera’s memory system is a substantial architectural advantage, but part of that advantage reflects a newer memory generation and a lower core count.
Vera combines 88 Olympus cores with up to 1.2 TB per second of LPDDR5X memory bandwidth. NVIDIA uses SOCAMM2 modules running at 9,600 MT/s, according to its architecture material. That bandwidth helps serve the many memory accesses generated by interpreters, databases, retrieval systems, and concurrent agent sandboxes.
AMD EPYC Turin supports 12 channels of DDR5 memory. The EPYC 9965 product page lists memory speeds up to 6,400 MT/s and 614 GB per second of socket bandwidth. The EPYC 9755 uses the same generation of memory interface.
At the socket level, Vera therefore offers almost twice Turin’s peak memory bandwidth. At the core level, the difference becomes larger because Vera divides that bandwidth among 88 cores. A 128-core or 192-core EPYC processor divides a smaller total among more cores.
NVIDIA says Vera provides more than three times the per-core memory bandwidth of traditional x86 server processors in its chosen comparisons. That framing is mathematically plausible when Vera is measured against a high-core-count Turin model. However, the result combines architecture, memory speed, channel design, and core count.
ServeTheHome notes that a large part of the socket-bandwidth comparison reflects memory vintage. Turin reached the market in October 2024 with DDR5-6400 support. Vera systems are scheduled for the second half of 2026 and use faster LPDDR5X memory.
This does not erase Vera’s advantage. NVIDIA made a deliberate architectural choice to attach more memory bandwidth to fewer, stronger cores. It does mean the comparison should not be interpreted as a pure measurement of core design.
Memory capacity adds another tradeoff. NVIDIA’s early Vera test system included eight 96 GB modules, providing 768 GB. EPYC platforms can support much larger memory footprints, which can matter for databases, analytics, virtualization, and CPU-hosted data surrounding large GPU clusters.
Latency also depends on more than raw bandwidth. Vera uses a monolithic compute die, a unified cache, and NVIDIA’s Scalable Coherency Fabric. NVIDIA says the fabric supplies 3.4 TB per second of on-die bandwidth and supports a single NUMA domain across a two-socket system.
NUMA, or non-uniform memory access, describes systems where memory access time depends on which processor or chiplet owns the data. Software that places memory poorly across NUMA domains can suffer unpredictable delays.
AMD’s high-core-count EPYC processors use chiplets. Chiplets allow AMD to assemble many cores economically, but communication between compute dies travels through an I/O die and fabric. That design can impose latency penalties for certain cross-die access patterns.
NVIDIA is using this contrast to argue that Vera delivers predictable latency under concurrency. For agent workloads, predictability can matter as much as average speed. One delayed tool execution can extend the response time experienced by a user.
Still, the relevant test is not peak bandwidth alone. A workload must generate enough useful memory traffic to benefit from it. Software also needs sufficient parallelism and a working set that stresses the memory subsystem.
Some compilation, cryptography, and branch-heavy workloads remain limited by core execution. Other database and analytics workloads can become sensitive to capacity, cache behavior, or storage access. Peak bandwidth cannot predict every outcome.
The compiler can also shift the result. NVIDIA’s internal SPEC measurements used LLVM and Clang for several comparisons, while published AMD results often use GCC. Compiler optimization affects scheduling, vectorization, code layout, and library selection.
A valid purchasing comparison should therefore hold the compiler and software stack constant when possible. If a buyer plans to deploy GCC-built software, a Clang advantage may never appear. If software can move to the faster compiler without compatibility problems, excluding that option would also distort the result.
Vera’s memory system strengthens NVIDIA’s case for agentic workloads. It does not automatically translate into a fixed advantage across every service. The gain must survive application code, compiler choices, data placement, and complete system testing.
What the NVIDIA Vera Benchmarks Still Do Not Prove
The available evidence shows that Vera is competitive, but it does not yet establish a universal lead over AMD EPYC Turin.
The first limitation is validation. NVIDIA measured the July SPEC CPU 2026 result internally. The company disclosed configurations in its whitepaper, but the result has not received the same public review as the submitted AMD scores in SPEC’s database.
That distinction becomes especially important because NVIDIA selected both the comparison processor and the normalization method. A vendor can follow benchmark rules while still choosing a framing that favors its product.
The second limitation is workload selection. SPEC CPU offers a controlled way to compare processor behavior, but it is not an agent platform. It does not measure an entire sequence involving retrieval, code execution, networking, databases, model inference, and GPU scheduling.
NVIDIA’s claim is ultimately about AI factories. The company argues that faster CPU-side work keeps GPUs busy and improves end-to-end agent responsiveness. A CPU throughput score is supporting evidence for that claim, not a direct measurement of it.
The third limitation is platform access. Phoronix tested Vera at NVIDIA’s Santa Clara facility before broad system availability. Its independent Linux tests included Vera, Grace, several AMD EPYC configurations, and Intel Xeon systems.
Those results were encouraging for NVIDIA. Vera produced strong performance in selected compilation, code, and server workloads. Across the chosen tests, it competed closely with or exceeded several x86 processors.
However, NVIDIA controlled the allowed benchmark set and provided only limited testing time. Phoronix explicitly described the suite as the workloads permitted for the session. Independent testers could not yet run every application, alter firmware extensively, or reproduce the platform in their own labs.
The fourth limitation is power measurement. Vera’s tested configuration had a peak 450-watt processor power target. AMD’s compared processors ranged from 300 watts to 500 watts. Processor thermal design power does not equal total node power, and different memory technologies change the platform balance.
A fair efficiency comparison needs wall power for complete systems under the same workload. It also needs throughput, latency, and energy measured simultaneously. Dividing a headline score by a listed processor wattage leaves too many platform variables unresolved.
The fifth limitation is software compatibility. Vera uses the Arm instruction set, while AMD EPYC uses x86-64. Linux and many cloud applications support both, but enterprises still run native extensions, proprietary agents, security software, and operational tooling compiled for x86.
Porting friction will vary. Containerized open-source services may move easily. Older commercial software may require vendor certification or code changes. A processor can win a benchmark and still lose a deployment if software support delays production.
The sixth limitation is comparison timing. Vera arrives roughly two years after Turin. NVIDIA’s whitepaper appeared as AMD prepared to discuss its next-generation EPYC Venice family, making Turin the outgoing comparison point rather than the concurrent architecture.
Comparing against a shipping processor is reasonable. Buyers can purchase Turin systems now, while Vera availability is planned for the second half of 2026. Yet long-term architectural claims should also be tested against hardware that will compete during Vera’s main deployment window.
AMD has already said Venice will expand to 16 memory channels and increase socket bandwidth. Early company projections also emphasize greater rack density and throughput. Those claims require the same caution as NVIDIA’s internal Vera numbers.
AMD’s published rack methodology estimates performance under a 100 kW rack limit. It projects an advantage for EPYC 9965 over Vera across six workloads and a larger lead for Venice.
However, AMD estimates Vera’s performance by scaling results from Grace, and it estimates portions of the future Venice comparison. That method answers NVIDIA’s favorable framing with AMD’s favorable framing. Neither vendor model replaces matched testing on shipping systems.
The correct skeptical position is symmetrical. NVIDIA’s charts should not be treated as proof of broad superiority. AMD’s rack projections should not be treated as proof that Vera loses across deployed AI infrastructure.
Both companies are selecting the level where their designs look strongest. NVIDIA emphasizes per-core progress, memory bandwidth, and latency. AMD emphasizes total throughput, density, and the number of cores available under a rack power limit.
Buyers should treat those narratives as testable hypotheses. The decisive data must come from reproducible systems running software that resembles the intended service.
The Real Contest Is Agent Latency Versus Rack Throughput
The AMD Nvidia rivalry is becoming a contest over which performance unit defines an AI data center’s economics.
NVIDIA’s strategy starts with the GPU. Vera is part of the Vera Rubin platform, where the CPU, GPU, interconnect, networking, and software stack are designed together. The CPU’s role is to remove delays around expensive accelerators.
That makes per-thread progress strategically valuable. If a faster CPU step reduces GPU idle time, Vera can create value beyond its own benchmark score. NVIDIA can also use NVLink-C2C coherency to connect the CPU tightly with the rest of its platform.
AMD approaches the market from a broader server base. EPYC serves cloud instances, databases, analytics, virtualization, storage, high-performance computing, and GPU hosts. Its high core counts let operators consolidate more workloads onto each socket or rack.
For many services, aggregate throughput still dominates. Web servers, key-value stores, batch analytics, virtual machines, and container platforms can distribute work across hundreds of cores. More completed jobs per rack can outweigh lower performance from each individual core.
The two strategies overlap in agentic AI because agent platforms need both qualities. They need quick sequential steps for user-facing responsiveness. They also need massive concurrency when many users, tools, and environments operate at once.
A small interactive deployment might favor Vera’s design. Consider an agent that edits code, launches a sandbox, compiles a project, reads an error, and tries again. Faster branch-heavy processing and compilation can shorten every loop.
A large cloud platform might reach another conclusion. It may run thousands of agents alongside databases, retrieval services, security filters, and unrelated customer workloads. Higher socket throughput and memory capacity can improve fleet utilization.
Licensing can reverse the economics again. Some enterprise software is licensed per processor core. Stronger performance from fewer cores may reduce the required license count. Open-source services without per-core fees may reward maximum core density instead.
Architecture is only one part of the decision. Procurement relationships, system availability, firmware maturity, virtualization support, and operational familiarity also affect deployment. AMD has an established x86 server base, while Vera enters as a newer Arm platform.
NVIDIA has another advantage through platform integration. Organizations already buying Rubin systems may prefer the CPU designed for that rack, especially if management tools and performance tuning arrive as one package. That decision can happen even when an isolated CPU benchmark remains close.
AMD can counter through flexibility. EPYC systems support accelerators from multiple vendors and fit established server configurations. Buyers concerned about platform concentration may value that option independently of a benchmark lead.
This is why the normalized results matter. A supposed 50 percent processor advantage can end a discussion before it starts. A 6 to 10 percent per-core lead creates a more realistic evaluation.
At that smaller margin, memory capacity, total throughput, software compatibility, rack power, and platform integration can determine the outcome. The comparison becomes a procurement problem rather than a victory chart.
Vera still deserves attention. NVIDIA has designed its first custom Arm core around an emerging workload and appears to have reached competitive server performance. That is significant in a market where mature x86 architectures have benefited from decades of optimization.
AMD also has a credible response. Turin’s reviewed results show that its current processors remain competitive, while the portfolio offers different core-count and frequency options. Buyers are not limited to the one EPYC SKU chosen in NVIDIA’s whitepaper.
The market pressure extends beyond these two companies. Intel must defend Xeon against AMD’s density and NVIDIA’s vertically integrated AI platform. Arm server vendors must now compete with an NVIDIA-designed core backed by the dominant accelerator supplier.
Still, the primary opponent remains NVIDIA’s per-core agent strategy against AMD’s socket and rack throughput strategy. Other vendors provide context, but they do not define the central decision created by these results.
Three Signals Will Decide Whether Vera’s Lead Holds
The next phase needs reproducible platform evidence, matched application tests, and direct competition with AMD’s current generation.
The first signal is a reviewed Vera submission to the SPEC CPU 2026 database. NVIDIA’s reported score of 925 provides a useful starting point, but independent review would confirm the benchmark configuration and create a consistent record beside AMD’s published results.
A reviewed result near 925 would strengthen the normalized conclusion. It would confirm that Vera delivers excellent performance per core while trailing higher-core-count Turin systems in total throughput. A materially different score would require the comparison to be recalculated.
The second signal is unrestricted testing on generally available Vera systems. Reviewers need enough access to change compilers, firmware settings, memory placement, thread counts, and software versions. They also need to measure power at the wall.
The most useful tests will follow complete agent loops. These should include retrieval, sandbox creation, code execution, compilation, database access, network calls, and GPU coordination. Results should report median latency, tail latency, throughput, energy, and accelerator utilization.
If Vera keeps its advantage across those workflows, NVIDIA’s architecture argument will become much stronger. If the lead disappears outside selected benchmarks, the whitepaper will look more like narrow product positioning.
The third signal is a matched Vera comparison against AMD EPYC Venice. Turin is available today, but Venice will compete during much of Vera’s commercial life. The comparison should use identical compilers, operating systems, workload versions, power limits, and node configurations.
That test must separate per-core, per-socket, and per-rack results. Combining them into one number would repeat the framing problem that triggered this debate.
A Venice win in per-core agent workloads would weaken NVIDIA’s central case. A Vera win in end-to-end latency, especially at lower node power, would validate NVIDIA’s decision to prioritize fewer strong cores and higher bandwidth.
For infrastructure buyers, the practical action is straightforward. Keep NVIDIA’s Vera claims and AMD’s responses in separate columns until both run the same workload. Record whether each number is measured, estimated, submitted, or independently reviewed.
Then choose the unit that maps to the service. An interactive agent platform should measure completed agent steps and tail latency. A cloud fleet should measure useful throughput per rack. A licensed database should measure performance per paid core.
The AMD Nvidia benchmark fight is already producing better questions than the original charts. Will Vera’s per-core lead survive independent review? Will that lead reduce full agent response time? Will AMD’s greater density deliver more useful work within the same power envelope?
Those three answers, not a vendor-selected ratio, should decide the next server purchase.



