top of page

OpenAI Jalapeño ASIC Is Staying Internal, but Its Nvidia Challenge Points Further

1 hour ago
12 min read

OpenAI has confined its first custom inference chip to internal deployments, despite presenting the OpenAI Jalapeño ASIC as capable of running models beyond its own portfolio.

That distinction now defines the chip’s place in the market. OpenAI needs Jalapeño to reduce the power and capacity pressures created by ChatGPT, Codex, its API, and increasingly demanding agentic products. Yet its public benchmarks deliberately put the chip beside Nvidia’s GB200 and GB300 systems.

Richard Ho, OpenAI’s head of hardware, told Tom’s Hardware that supporting OpenAI will remain the priority for “a good long time.” He also said the processor could serve other organizations and was not hard-coded for OpenAI models.

That leaves OpenAI in an unusual position. It is not offering an Nvidia alternative to cloud buyers, developers, or competing AI laboratories. However, it is building and publicly measuring something that increasingly resembles one.

The real contest is therefore not OpenAI against Nvidia in the merchant chip market. It is OpenAI’s internal-first strategy against the wider potential of the programmable inference platform it has created.

The OpenAI Jalapeño ASIC Has a Clear First Customer

OpenAI built Jalapeño to serve its own inference demand, and that demand is large enough to absorb the initial supply.

Jalapeño is an application-specific integrated circuit, or ASIC, designed around large language model inference. Inference is the process that produces an answer after a trained model receives a request.

OpenAI unveiled the processor with Broadcom on June 24, 2026. The companies described it as OpenAI’s first “Intelligence Processor” and the beginning of a multi-generation compute platform.

The Jalapeño announcement tied the chip directly to OpenAI’s operating workloads. Those include ChatGPT, Codex, the OpenAI API, and future products built around AI agents.

OpenAI designed the architecture, while Broadcom contributed silicon implementation and networking technology. Celestica supports the boards, racks, and system integration needed to move from a chip into a deployable computing platform.

That division of labor matters. OpenAI is not simply purchasing a processor and adapting its software afterward. It is aligning chip architecture, memory movement, networking, kernels, scheduling, and product workloads around the same inference goals.

The immediate objective is efficiency. Ho told Tom’s Hardware that this was the central design priority, especially within power-constrained data centers.

Modern AI infrastructure cannot expand by adding processors alone. Operators must also secure electricity, cooling, networking, memory, packaging capacity, and physical data-center space.

A processor that completes more useful inference work within a fixed power envelope can create capacity without waiting for another facility. For OpenAI, that means serving more requests from the infrastructure it can actually deploy.

This explains why an external launch is not the current priority. OpenAI already has a captive customer with persistent demand, direct access to its software stack, and strong incentives to improve utilization.

Ho said the company will have its “hands full” meeting its own requirements. His comments establish an order of operations rather than a permanent restriction.

Jalapeño will first be installed where OpenAI controls the models, serving software, and infrastructure decisions. That environment also gives engineers the shortest feedback loop for identifying bottlenecks and improving later generations.

Initial deployment is scheduled for the end of 2026. The company has not announced a public cloud instance, accelerator card, licensing program, or sales channel for outside customers.

That makes Jalapeño an internal platform today. It does not make the architecture inherently internal.

OpenAI says engineering samples have already run machine-learning workloads at their target production frequency and power. The company also presented results from models developed outside OpenAI, which complicates any narrow interpretation of the processor.

The chip’s first customer is settled. The boundaries of its future market are not.

Efficiency Turns Internal Silicon Into a Strategic Asset

Jalapeño gives OpenAI another way to expand inference capacity without relying entirely on general-purpose accelerator roadmaps.

OpenAI and Broadcom first disclosed their broader infrastructure partnership in October 2025. Their agreement covers 10 gigawatts of OpenAI-designed accelerators and associated networking systems.

Under the 10-gigawatt collaboration, rack deployments were targeted to begin during the second half of 2026. Completion was scheduled by the end of 2029.

That scale makes Jalapeño more than an isolated chip experiment. It is intended to become part of a large infrastructure program spanning multiple processor generations and data-center partners.

OpenAI says the first processor moved from design to manufacturing tape-out in nine months. Tape-out is the point when a completed chip design is released for fabrication.

The schedule is an OpenAI and Broadcom claim, not an independently established industry record. Even so, the timeline shows why software and semiconductor roadmaps are starting to converge.

OpenAI knows which operations consume time and energy across its deployed products. Those observations can influence the hardware before manufacturing, instead of being handled entirely through later software optimization.

The architecture targets reduced data movement and closer alignment among compute, memory, and networking resources. Moving model data often consumes substantial energy, so reducing unnecessary movement can improve system-level efficiency.

That focus also reflects how inference differs from training. Training creates or updates model parameters, while inference repeatedly uses those parameters to answer live requests.

Jalapeño does not replace Nvidia hardware across both workloads. OpenAI has positioned it for inference, leaving training outside its stated role.

However, inference becomes more important as usage grows. Every ChatGPT response, API call, or multi-step Codex task adds recurring serving work after model training is complete.

Agentic products can increase that pressure. A single user instruction might trigger many model calls, tool actions, checks, and revisions before returning a completed result.

For developers and enterprise buyers, the effect matters more than the chip label. Better inference efficiency can support lower latency, steadier availability, or improved economics when demand spikes.

OpenAI has not committed to passing any savings directly to customers. It has also not published enough operational data to calculate a verified cost reduction per request.

The strategic benefit is still visible. Owning an inference architecture gives OpenAI another lever when capacity, electricity, or supplier schedules constrain product growth.

It also puts pressure on Nvidia without requiring OpenAI to become a direct chip vendor. Every workload shifted to Jalapeño is one workload that does not require an Nvidia inference accelerator.

That pressure has clear limits. Nvidia supplies a mature software environment, training systems, networking, and accelerators available across multiple cloud providers.

OpenAI cannot reproduce that ecosystem simply by producing an efficient ASIC. It does not need to do so for Jalapeño to become valuable internally.

An internal chip only has to outperform the available alternatives on the workloads OpenAI runs frequently enough to justify the engineering and deployment costs.

That is why efficiency is the pivotal metric. Jalapeño’s value depends less on winning every benchmark than on delivering useful tokens within OpenAI’s power, latency, and supply constraints.

OpenAI Jalapeño vs Nvidia Is Not a Clean Contest

OpenAI’s benchmarks establish Jalapeño as a credible inference design, but they do not establish a complete replacement for Nvidia’s platforms.

At Hot Chips 2026, OpenAI compared Jalapeño with Nvidia GB200 and GB300 rack systems using SemiAnalysis’s public InferenceX benchmark suite.

The tests covered OpenAI’s GPT-OSS 120B model, DeepSeek R1 670B, and Moonshot AI’s Kimi K2.5. Including outside models supported OpenAI’s claim that Jalapeño is programmable.

According to the published results, Jalapeño produced between 1.5 and 1.9 times more throughput per kilowatt. OpenAI also reported between 1.7 and 3.6 times lower end-to-end latency.

The published benchmarks compared a 700-watt Jalapeño package with Nvidia accelerators rated at 1,200 and 1,400 watts.

OpenAI said sustained Jalapeño power remained at or below 550 watts during testing. Its primary calculations nevertheless normalized results using each processor’s published thermal design power.

The chip combines one compute die with six HBM4 stacks. It provides 216 GiB of high-bandwidth memory and 15.4 terabytes per second of memory bandwidth.

High-bandwidth memory places memory close to the processor and supplies data much faster than ordinary server memory. That bandwidth is critical when large model parameters must remain available during inference.

OpenAI’s design reportedly concentrates on exposing more of the available bandwidth to running workloads. It does not merely add memory bandwidth and assume applications will use it efficiently.

Those specifications help explain Jalapeño’s efficiency claims. They do not remove several important qualifications.

First, OpenAI chose the models, configurations, latency points, and comparison framework. SemiAnalysis participated in running the tests, but the results still came from OpenAI’s laboratory and hardware.

Second, the main comparisons used package power ratings. A supplementary utility-power comparison produced narrower differences because it included more of each system’s actual energy consumption.

Third, Jalapeño and the Nvidia systems did not always use identical decoding techniques. Nvidia deployments can use multi-token prediction, which generates several candidate tokens during one large-model pass.

Jalapeño used single-token prediction in the highlighted comparisons. That makes the results informative, but it prevents readers from reducing the entire contest to one performance ratio.

Fourth, Nvidia’s newer Vera Rubin platform was not included. Benchmarking against GB200 and GB300 shows competitiveness against important Blackwell systems, not leadership over every available Nvidia generation.

Most importantly, Jalapeño does not train models. Nvidia’s systems support both training and inference, giving customers greater flexibility when workloads change.

OpenAI also continues to need Nvidia equipment. Jalapeño is a diversification strategy, not evidence that the relationship has ended.

The OpenAI Jalapeño vs Nvidia framing therefore requires precision. Jalapeño challenges Nvidia’s efficiency on selected inference workloads under OpenAI’s published conditions.

It does not match Nvidia’s total market role, software reach, deployment availability, or training capabilities.

The benchmarks still matter because custom ASIC comparisons rarely look this direct. Specialized chips often remain hidden behind internal cloud services, with limited public evidence about competitive performance.

OpenAI instead named current Nvidia systems, supplied power-normalized results, and ran three large models. That decision invited the market to evaluate Jalapeño as more than captive infrastructure.

The company says engineers ported DeepSeek and Kimi to the A0 silicon sample within roughly two months. A0 refers to the first manufactured revision returned for testing.

Ho said the demonstration was meant to reject the assumption that the chip only supports OpenAI models. His more consequential claim was that Jalapeño is programmable and general-purpose within its intended inference domain.

That distinction keeps the external-market question alive. A processor tied to one private model stack would have little relevance outside its owner.

A processor that can run major external models has a wider addressable use case, even if OpenAI has not chosen to sell access.

Programmability Leaves the Door Open, but Supply Keeps It Narrow

Jalapeño’s software flexibility supports a broader rollout, while manufacturing capacity and OpenAI’s own demand argue against one soon.

The strongest evidence for wider use is not Ho’s speculative wording. It is the combination of outside-model support, a multi-generation roadmap, and planned gigawatt-scale deployment.

OpenAI could eventually expose the hardware in several ways. It might offer Jalapeño-backed API capacity without letting customers manage the chip directly.

It might work with an infrastructure partner to provide dedicated instances. It could also make later generations available through selected data-center operators.

None of those routes has been announced. Treating them as plans would overstate what OpenAI has said.

The internal-use interview establishes a more restrained position. OpenAI believes other organizations could use the processor, but its own capacity needs come first.

Supply is the most immediate constraint. Advanced accelerators compete for fabrication capacity, high-bandwidth memory, packaging, networking components, electricity, and completed data-center space.

Jalapeño does not escape those bottlenecks because OpenAI designed it. It enters many of the same production queues as competing AI processors.

Its six HBM4 stacks are especially relevant. Scaling that configuration across a substantial deployment would turn OpenAI into a major consumer of scarce high-bandwidth memory.

The system also needs more than packaged processors. Broadcom provides networking technology, while Celestica contributes boards, racks, and system integration.

Every layer must arrive in the correct sequence. A shortage in one component can delay useful capacity even when the accelerator itself is ready.

OpenAI says it is in good shape regarding internal supply. That statement does not demonstrate that it has enough capacity to support outside customers with predictable service guarantees.

External sales would create additional obligations. Buyers would need documentation, software support, workload qualification, maintenance, scheduling, and stable access to future generations.

A successful merchant platform also needs trust that extends beyond benchmark performance. Customers must believe the vendor will provide predictable capacity and maintain compatibility across software updates.

OpenAI currently controls both sides of that relationship inside its own infrastructure. It can change models, kernels, and scheduling software alongside the chip.

An external customer would not necessarily share OpenAI’s workload mix. Different model architectures, quantization formats, context lengths, and latency targets might expose limitations absent from the published tests.

That is the central skeptical angle. Running three outside models shows flexibility, but it does not prove universal efficiency across production inference environments.

The Hot Chips presentation supplied meaningful architectural detail. An independent deployment would still offer stronger evidence than results produced in OpenAI’s laboratory.

A wider rollout could also distract from the internal objective. Supporting third parties would consume engineering time while OpenAI is trying to deploy its first generation and prepare successors.

The company does not need merchant revenue to justify the project. Lower internal inference costs, higher capacity, or better response latency may provide sufficient returns.

That makes a near-term commercial launch less likely, even if the chip performs well. OpenAI gains strategic flexibility by preserving the option without committing supply.

The tension may remain unresolved across several generations. OpenAI can keep improving Jalapeño internally while making its software surface more portable.

If outside access eventually appears, it will probably follow proven internal deployments rather than precede them.

The more immediate impact falls on infrastructure suppliers. Broadcom gains a central role in OpenAI’s compute stack, while Nvidia faces another customer developing specialized silicon.

Google’s TPU program provides the clearest historical pattern. Google first used its Tensor Processing Units internally, then exposed them through Google Cloud rather than selling conventional accelerator cards.

Amazon followed a related path with Inferentia and Trainium. Those processors strengthen AWS services and give customers access through Amazon’s cloud environment.

OpenAI lacks an equivalent general-purpose cloud business. That difference makes its route to external availability less obvious.

Microsoft and other data-center partners could provide a channel. Still, the parties have not described whether customers will be able to select Jalapeño-backed capacity.

Until that changes, the processor remains infrastructure for OpenAI products rather than a purchasable alternative to Nvidia.

Three Signals Will Show Whether Jalapeño Moves Beyond OpenAI

Deployment evidence, broader workload validation, and an external access model will determine whether Jalapeño becomes a platform or remains private infrastructure.

The first signal is a verified production deployment by the end of 2026. OpenAI has said the initial platform will begin deployment before the year closes.

Readers should watch for confirmation that Jalapeño is serving real ChatGPT, Codex, or API traffic. Production use would test reliability, utilization, latency, and integration beyond controlled benchmarks.

The most useful disclosure would separate installed hardware from active serving capacity. A processor sitting in a laboratory or partially completed rack does not provide the same evidence as customer traffic.

Production operation would strengthen OpenAI’s case that its full-stack approach works. A delay would weaken the aggressive schedule suggested by its nine-month development cycle.

The second signal is independent performance testing across more models and decoding configurations. OpenAI’s Hot Chips results created a credible starting point, not a final verdict.

A stronger test would compare total system power, equivalent precision, matching latency targets, and current software stacks. It would also include Nvidia Vera Rubin when comparable systems become available.

Reviewers should examine performance across long contexts, varying batch sizes, mixture-of-experts models, and agentic workloads. Average efficiency can hide weak results at operating points that matter to interactive products.

The Hot Chips analysis shows why system details matter. Jalapeño is presented as a complete inference platform, not merely a standalone processor.

Independent validation across that whole platform would strengthen claims about general use. Narrower results would support the view that Jalapeño remains optimized primarily for OpenAI’s workload profile.

The third signal is a concrete external access mechanism. That means more than a statement that other organizations could use the chip.

A real move would include a named cloud partner, instance type, capacity commitment, developer program, or supported software environment. It would also define who operates the hardware and who supports customers.

If OpenAI announces such a route, Jalapeño will begin to pressure Nvidia as an accessible alternative. Without one, its competitive impact will remain indirect.

That indirect impact can still be substantial. Internal custom silicon reduces the share of future inference demand automatically flowing to established suppliers.

It also gives OpenAI leverage in purchasing negotiations. A credible alternative changes discussions about delivery schedules, system design, and long-term capacity, even when Nvidia remains essential.

Developers should not redesign applications around Jalapeño today. There is no public instance, direct programming environment, or deployment commitment available to them.

They should instead watch for changes in OpenAI’s service behavior. Lower latency, increased rate limits, steadier availability, and new agentic features might reveal the operational effects first.

Enterprise buyers should ask a different question. They need to know whether OpenAI’s infrastructure diversification improves service reliability or introduces another immature dependency.

OpenAI’s control over more of the stack can accelerate optimization. It can also concentrate technical decisions within one provider and make performance harder to compare with standard cloud hardware.

Knowledge workers will experience the outcome through products rather than processor specifications. Faster responses matter, but dependable multi-step work may matter more as AI systems handle longer tasks.

OpenAI Jalapeño explained through that lens is less about entering the semiconductor business. It is about controlling the cost and capacity behind every model interaction.

For now, the company has drawn a firm operational boundary. Jalapeño will serve OpenAI first because OpenAI expects to consume everything it can deploy.

Its public positioning points beyond that boundary. The processor runs outside models, follows a multi-generation roadmap, and has been benchmarked against Nvidia’s flagship Blackwell systems.

Those choices preserve a larger ambition without requiring an immediate commercial launch.

The next three months should make the distinction clearer. Watch for production traffic, independent benchmark replication, and a named route for outside access.

If all three appear, Jalapeño will have moved from internal infrastructure toward a broader inference platform. If only deployment arrives, OpenAI will still have achieved something strategically important.

It will have created a specialized source of compute for its fastest-growing workloads while keeping Nvidia in the stack where Nvidia remains strongest.

That mixed strategy is the current reality. OpenAI is not replacing Nvidia, and it is not selling Jalapeño to the market.

It is building enough technical and operational credibility to keep both possibilities open.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page