OpenAI Jalapeño ASIC Deployment Chooses AMD Turin, Not Nvidia Vera
OpenAI’s Jalapeño ASIC deployment pairs its new inference chips with AMD EPYC Turin hosts, despite the company’s deep infrastructure relationship with Nvidia. Each host carries two Turin-class processors and 1.5TB of DRAM. Nvidia’s new Vera CPU did not make the first production design.
That choice is more than a component swap. OpenAI designed Jalapeño as an application-specific integrated circuit, or ASIC, optimized for language-model inference. Yet it built the surrounding host layer with a mature x86 server platform instead of Nvidia’s purpose-built Arm CPU.
Richard Ho, OpenAI’s vice president and head of hardware, described the Turin decision as “pragmatic.” He told Tom’s Hardware that standalone Vera remained “a little bit behind” at the required maturity level. The comment places deployment certainty ahead of tighter ownership of the computing stack.
OpenAI still depends heavily on Nvidia accelerators elsewhere. Jalapeño also remains an internal platform whose early performance claims require broader validation. Even so, the host decision shows how hyperscalers can challenge Nvidia selectively without abandoning its hardware entirely.
OpenAI Jalapeño ASIC Deployment Starts With a Two-Rack System
OpenAI has turned Jalapeño from a chip announcement into a rack design built around separate AMD host and custom accelerator layers.
The architecture uses one CPU host rack beside one Jalapeño ASIC rack. SemiAnalysis describes 16 “Katsu” CPU trays in the host rack, matched with 16 “Vindaloo” accelerator trays in the neighboring rack.
Each Katsu tray contains two AMD EPYC Turin-class CPUs and 1.5TB of DRAM. It also includes local storage and 400-gigabit frontend networking. Eight external PCIe cables connect each CPU tray to its corresponding accelerator tray.
The neighboring rack holds 128 Jalapeño chips across its 16 accelerator trays. Eight “Chana” switch trays connect those accelerators inside the rack and across larger installations.
The published rack architecture can extend its scale-up network across 16 racks. That configuration connects as many as 2,048 Jalapeño accelerators through copper and optical links.
The host processors do not replace the inference ASICs. They handle the supporting CPU work required to feed, coordinate, schedule, and manage accelerator workloads. The custom silicon remains responsible for the language-model calculations targeted by Jalapeño.
This division matters because an accelerator rarely operates as an isolated device. Production inference needs tokenization, request handling, storage access, networking, model orchestration, safety services, and other CPU-bound operations.
Agentic applications increase those demands. An agent can alternate between model inference, Python execution, retrieval, database access, and external tools. Poor host performance can leave expensive accelerators waiting while those stages complete.
OpenAI therefore needed more than a fast inference chip. It needed a host platform with sufficient memory capacity, network support, software compatibility, and operational history. Turin offered those qualities without adding another immature component to the program.
Power also illustrates the system-level challenge. SemiAnalysis estimates the host rack uses about 31 kilowatts in production, while the accelerator rack draws roughly 130 kilowatts. Together, the paired system consumes about 160 kilowatts.
Those figures show why Jalapeño cannot be assessed through chip specifications alone. Rack networking, host utilization, cooling, software, and workload placement all influence the useful work delivered by the installation.
The same principle applies to OpenAI’s benchmark claims. A favorable accelerator result matters only if the complete system can repeat it under production traffic. Choosing an established host platform reduces one source of uncertainty during that transition.
OpenAI says Jalapeño will begin its initial deployment by the end of 2026. The disclosed rack configuration provides the clearest view yet of how that deployment will work.
It also creates the article’s central tension. OpenAI built custom inference silicon to gain more control, but it avoided extending that experiment into the CPU layer.
Turin Reduces Risk in a Nine-Month Chip Program
AMD EPYC Turin hosts gave OpenAI a known platform while the company attempted an unusually compressed accelerator development schedule.
OpenAI and Broadcom say Jalapeño moved from initial design to manufacturing tape-out in nine months. Tape-out is the point when a completed chip design is sent for fabrication.
That schedule is a company claim, not an independently established industry record. However, it still describes a program with little room for avoidable integration problems.
OpenAI designed the accelerator architecture, while Broadcom contributed silicon implementation, networking, and connectivity expertise. Celestica worked on the board, rack, and complete system design.
The official Jalapeño announcement frames it as the first generation of a longer compute roadmap. Initial deployment is planned for late 2026, followed by expansion across later generations.
Ho said the team wanted aggressive performance and cost goals without accepting unnecessary risks. Turin met the program’s requirements, and OpenAI’s partners already had relevant experience with the platform.
That experience can be valuable during bring-up. Engineers must validate firmware, memory behavior, PCIe connectivity, operating systems, drivers, telemetry, failure handling, and workload scheduling before a rack enters regular service.
A new CPU architecture would expand that validation surface. Differences in instruction sets, compiler behavior, management tooling, and application compatibility can create delays even when the underlying processor performs well.
Turin belongs to AMD’s fifth-generation EPYC server family. It uses the established x86 instruction set and supports twelve memory channels per socket, giving system designers substantial memory bandwidth and capacity.
AMD’s own Turin architecture documentation describes production configurations using DDR5 memory. OpenAI’s rack design places 1.5TB of DRAM beside each pair of processors.
That memory pool serves a different role from the high-bandwidth memory attached directly to Jalapeño. Host DRAM can retain application state, prepare requests, manage data, and support CPU-side services surrounding inference.
OpenAI also had to prepare software for an accelerator without an existing developer ecosystem. Nvidia benefits from years of CUDA adoption, optimized libraries, deployment tools, and operator familiarity.
Jalapeño begins without that installed base. OpenAI can control its internal software environment, but its engineers must still build compilers, kernels, monitoring systems, and scheduling logic for the new platform.
Using familiar host hardware keeps that work concentrated on the custom accelerator. It lets the team separate Jalapeño-specific defects from problems caused by an additional CPU transition.
The decision therefore reflects schedule discipline rather than a sweeping judgment about x86 and Arm. OpenAI selected the component that reduced integration risk for this generation.
This distinction is important. Ho did not argue that Turin will remain the best host for every future OpenAI system. He said it satisfied the immediate program’s needs and maturity requirements.
The first Jalapeño generation is consequently a hybrid strategy. OpenAI takes architectural risk where specialization promises meaningful inference gains, while retaining commodity server technology where maturity provides greater value.
AMD EPYC Turin Hosts Put Pressure on Nvidia’s Full-Stack Pitch
The immediate pressure falls on Nvidia’s attempt to make Vera the default CPU around next-generation AI infrastructure.
Nvidia presents Vera as a processor designed for the CPU work surrounding agents. Its target workloads include Python runtimes, sandboxed code, orchestration, analytics, and other tasks that occur between accelerator calls.
The processor uses 88 custom Olympus cores and an LPDDR5X memory subsystem. Nvidia says that subsystem provides as much as 1.2TB per second of bandwidth.
Vera also connects to Rubin GPUs through NVLink-C2C. Nvidia states that this link supplies up to 1.8TB per second of coherent bandwidth between the CPU and GPU.
That tight coupling supports Nvidia’s larger sales argument. Customers can buy CPUs, GPUs, networking, interconnects, libraries, and rack systems as one coordinated platform.
Nvidia says Vera systems will become available through system builders and cloud partners during fall 2026. Its listed supporters include major server manufacturers and cloud infrastructure providers.
OpenAI’s choice exposes a timing problem inside that strategy. Vera may offer attractive specifications, but Jalapeño needed a host platform that partners could integrate during an accelerated development cycle.
A processor can be technically complete before its wider operating environment becomes mature. Server boards, firmware, management software, validation procedures, deployment experience, and supply readiness develop on separate schedules.
Ho’s criticism focuses on that difference. He did not say Vera lacks the performance required for AI hosting. He questioned the maturity of Vera as a standalone CPU for the Jalapeño program.
That qualification matters because Vera also serves as the host processor inside Nvidia’s integrated Vera Rubin systems. The standalone role creates additional requirements beyond a tightly controlled CPU-GPU configuration.
OpenAI uses its own accelerator and scale-up network, rather than Rubin GPUs and Nvidia’s native interconnect arrangement. A standalone Vera host would need to fit into that outside architecture without relying on the complete Nvidia platform.
Turin offers a more conventional relationship. OpenAI can connect an established server CPU to its custom accelerator through PCIe and preserve control over the rest of the rack.
This outcome weakens Nvidia’s full-stack pitch at one strategic customer, but it does not establish a broad market defeat. OpenAI continues to use Nvidia hardware, and Vera has commitments from numerous infrastructure providers.
OpenAI itself has emphasized that Nvidia remains an important partner. Its custom ASIC program appears designed to supplement a large and diverse computing fleet, rather than replace every Nvidia deployment.
AMD’s gain is also narrower than a direct accelerator victory. Turin supplies the host layer, while OpenAI’s own chips perform the specialized inference work.
Still, host CPUs occupy a valuable position. They control data preparation and orchestration around accelerators, and their memory systems influence how efficiently the full installation operates.
The OpenAI design gives AMD a place inside a high-profile custom silicon platform. It also demonstrates that an x86 host can support a large inference rack without sharing the accelerator vendor’s architecture.
For cloud providers and AI laboratories, this creates a credible modular alternative. They can develop or purchase specialized accelerators while retaining familiar CPUs, operating systems, and server management practices.
Nvidia wants Vera to make the opposite route more attractive. Its proposition is that a coordinated CPU, GPU, networking, and software stack can deliver better total system performance.
Jalapeño therefore creates a practical contest between modularity and integration. The winner will depend on deployment results, software quality, and total operating cost, not processor benchmarks alone.
The Real Reversal Is Selective Control, Not a Complete Nvidia Exit
OpenAI is breaking apart Nvidia’s platform where custom design offers leverage, while preserving mature components wherever replacement adds risk.
The strongest interpretation of Jalapeño is not that OpenAI has abandoned Nvidia. The company is instead separating the AI rack into layers and deciding which layers justify custom control.
Inference is the obvious starting point. OpenAI operates ChatGPT, Codex, its API, and other products that generate enormous volumes of model-serving traffic.
A custom inference accelerator can target those recurring workloads more narrowly than a general-purpose GPU. OpenAI can tune the chip, memory hierarchy, network, kernels, and scheduling system around its own models.
Jalapeño’s architecture addresses both major phases of language-model inference. Prefill processes the user’s prompt and is relatively compute intensive. Decode generates tokens and depends more heavily on memory bandwidth.
Moving model state between specialized resources can add communication delays. OpenAI says Jalapeño keeps important state, including the KV cache used during generation, close to the active compute resources.
The company reports that Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput than its comparison systems. It also claims 1.7 to 3.6 times lower end-to-end latency.
Those first benchmark results covered GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. OpenAI used SemiAnalysis’ public InferenceX benchmark across several operating points.
OpenAI rates each Jalapeño chip at 700 watts. It says measured sustained consumption stayed at or below 550 watts during the tested workloads.
These are notable figures, but they remain early results presented by the chip’s designer. OpenAI selected the configurations, workloads, software, and comparison methodology described in its publication.
SemiAnalysis says its team observed testing on real silicon. That adds more evidence than a simulation or projected specification, but it does not replace independent benchmarking across varied production conditions.
The comparison also centered on commercially available Nvidia systems rather than Vera Rubin. That limits what the results establish about Nvidia’s next architecture.
Jalapeño’s advantage may be strongest on workloads resembling OpenAI’s internal serving patterns. That specialization is its purpose, but it also reduces the reach of broad claims about overall accelerator leadership.
A general-purpose GPU must support many models, frameworks, numerical formats, and research workloads. An internal ASIC can sacrifice some flexibility to improve efficiency on a narrower operational target.
OpenAI can tolerate that trade because it controls models, serving software, and demand. An enterprise buying infrastructure for unknown future workloads faces a different calculation.
This is where the Turin decision becomes revealing. OpenAI pursued specialization in the accelerator because inference efficiency can directly affect product latency and computing demand.
It did not specialize every surrounding layer. The host CPU remained a place where compatibility, availability, and partner experience outweighed architectural novelty.
The strategy resembles a controlled decomposition of the AI platform. OpenAI keeps ownership of the parts most connected to its model roadmap and buys proven components for the rest.
Broadcom and Celestica remain essential within that model. Custom silicon does not mean self-sufficiency, because implementation, networking, manufacturing, boards, and rack integration require experienced suppliers.
AMD benefits from the same selective approach. Its processor becomes part of OpenAI’s system because it works as a mature building block, not because OpenAI adopted AMD’s complete accelerator platform.
Nvidia faces pressure because its business increasingly depends on selling an integrated AI factory. Customers that unbundle those layers can shift value toward their own chips and alternative suppliers.
However, integration retains important advantages. A single vendor can optimize coherent memory, interconnects, software, diagnostics, and support across the entire machine.
OpenAI must recreate or coordinate many of those capabilities around Jalapeño. Its early hardware efficiency claims will matter only if the operating stack remains dependable as deployments grow.
The reversal is therefore limited but significant. OpenAI is no longer accepting the accelerator vendor’s full architecture as an indivisible package.
Early Benchmarks Do Not Settle the Production Question
Jalapeño still has to prove reliability, utilization, and economic value across sustained internal workloads.
Benchmark leadership can disappear when a system encounters real traffic. Production demand varies by model, prompt length, output length, batch size, latency target, and geographic region.
Interactive services also experience sharp demand changes. A rack must sustain useful utilization without making users wait, even when request patterns differ from the benchmark configuration.
OpenAI’s tests used public models and a defined inference suite. Those results support the claim that the silicon works and can run large language models efficiently.
They do not establish long-term reliability across thousands of accelerators. Hardware failures, network congestion, thermal limits, software defects, and maintenance requirements become clearer only during extended deployment.
The two-rack structure adds physical complexity. Each accelerator tray relies on a corresponding host tray, multiple PCIe cables, and a separate switching layer.
That modularity can simplify replacement and preserve component choice. It can also create more connections that engineers must monitor, service, and validate.
Power figures need similar caution. Chip ratings help normalize benchmark results, but data centers pay for complete systems, cooling, networking, storage, and unused capacity.
A 160-kilowatt paired rack must deliver enough sustained throughput to justify its infrastructure. Peak benchmark performance does not guarantee that operating result.
Software is another open question. Nvidia’s advantage extends beyond silicon into mature libraries, profilers, compilers, orchestration tools, and a large pool of experienced developers.
OpenAI can build software around a controlled set of internal models. Yet each new model architecture or numerical format may require additional optimization before it uses Jalapeño efficiently.
The company says AI assisted parts of Jalapeño’s design and programming process. That may shorten optimization cycles, but it remains a company-reported benefit without a public productivity comparison.
Model evolution creates a longer-term risk. A fixed-function design started years before deployment must remain useful as inference techniques change.
Jalapeño is not fully fixed in the narrowest sense, and it supports multiple large models. However, its economic advantage still depends on workloads matching the assumptions behind its architecture.
Nvidia’s general-purpose approach offers more insurance against unexpected changes. Customers can repurpose GPUs across training, inference, simulation, and other accelerated workloads.
OpenAI can offset that disadvantage through scale. If ChatGPT and API demand remain large, even a narrower accelerator can stay highly utilized on predictable serving work.
The Turin hosts add another uncertainty. They were selected partly for maturity, but OpenAI has not disclosed every workload performed on the CPU layer.
Without that breakdown, readers cannot determine whether 1.5TB of host memory is necessary for model serving, operational flexibility, or future software requirements.
The decision also does not prove that Vera is unsuitable. Nvidia’s CPU is entering the market through multiple system vendors and cloud partners, and broader deployment evidence is still emerging.
Ho’s assessment applies to one project’s schedule and risk limits. Vera’s maturity can improve after Jalapeño’s first system design has already been fixed.
Nvidia may also demonstrate advantages when Vera operates beside Rubin through its coherent NVLink connection. That integrated configuration differs from using Vera as a host for third-party silicon.
A fair comparison must therefore examine complete systems under equivalent workloads. It should measure latency, throughput, power, availability, software effort, and total deployment cost.
Until such evidence arrives, OpenAI Jalapeño ASIC deployment remains a promising internal platform rather than a settled replacement for mainstream GPU infrastructure.
Three Signals Will Show Whether OpenAI’s Bet Holds
Deployment scale, production behavior, and Vera’s independent results will determine whether Turin was merely safer or strategically better.
The first signal is OpenAI’s planned Jalapeño ramp through the end of 2026. The company has promised initial deployment, but it has not publicly quantified the share of inference traffic moving onto the platform.
A meaningful production ramp would strengthen the case that Jalapeño works beyond controlled tests. Evidence could include broader model coverage, stable rack availability, or visible product latency improvements.
A slow or limited rollout would not automatically indicate failure. Supply constraints, data center readiness, and software qualification can delay otherwise functional hardware.
Still, repeated schedule changes would weaken the nine-month development narrative. Fast tape-out matters less if system integration requires a long period before useful service begins.
The second signal is operational evidence from the two-rack design. OpenAI should eventually provide measurements based on sustained production use, rather than only accelerator power ratings.
Useful figures would include complete rack power, utilization, failure rates, service intervals, and performance under mixed request patterns. Those measurements would test whether the modular host arrangement preserves Jalapeño’s efficiency.
Watch how frequently OpenAI updates its kernels and model support. Rapid optimization across new architectures would show that its software stack can keep pace with model development.
Also watch whether later Jalapeño generations retain AMD hosts. Continued use would suggest that modular x86 infrastructure offers lasting value beyond the first program’s deadline.
A change to Vera, another Arm processor, or a custom OpenAI CPU would point toward a transitional role for Turin. OpenAI has described Jalapeño as the beginning of a multigenerational platform, leaving future host choices open.
The third signal is Nvidia Vera’s performance outside Nvidia-controlled accelerator systems. Standalone deployments will test the exact maturity concern raised by Ho.
Nvidia has announced support from cloud providers and major server manufacturers. Their production availability, software compatibility, and independent benchmarks will show how quickly Vera closes the perceived gap.
If standalone Vera systems deploy smoothly and outperform x86 hosts on agent workloads, OpenAI’s choice will look increasingly schedule-specific. Nvidia’s integrated platform argument would remain intact.
If those deployments encounter delays or limited adoption, Turin’s selection will look more strategic. It would show that established x86 infrastructure can retain its role even as AI accelerators become more specialized.
Developers and enterprise buyers should care because host architecture affects more than benchmark charts. It shapes software portability, infrastructure availability, operational complexity, and the range of vendors available for future systems.
AI product users may see the outcome indirectly. Successful custom inference could reduce waiting, support longer agent workflows, and make service capacity more predictable during demand spikes.
None of those benefits is guaranteed by a chip announcement. They depend on the complete system delivering stable, economical inference at scale.
The key question is now concrete: can OpenAI turn its custom accelerator and AMD host design into a repeatable production platform before Nvidia’s Vera ecosystem matures?
Track the deployment rather than the headline comparison. If Jalapeño expands across models and data centers while retaining its efficiency, OpenAI’s selective hardware strategy gains credibility. If Vera closes the maturity gap first, Nvidia’s tightly integrated route will remain difficult to displace.



