OpenAI's Self-Developed Chip Jalapeño Passes First Test: Per-Watt Performance Up to 1.9x NVIDIA's Flagships, 700W vs 1400W

On August 26, 2026, OpenAI released the first performance data for its self-developed inference chip, Jalapeño, showing 1.5-1.9x higher throughput per kilowatt and 1.7-3.6x lower end-to-end latency than NVIDIA GB200 and GB300 rack systems on the InferenceX benchmark, marking the company's formal entry into hardware-level competition.

On August 26, 2026, OpenAI released the first performance data for its self-developed inference chip, Jalapeño: on InferenceX, a public benchmark published by industry analysis firm SemiAnalysis, Jalapeño delivered 1.5 to 1.9 times higher throughput per kilowatt and 1.7 to 3.6 times lower end-to-end latency than NVIDIA GB200 and GB300 rack systems across tests on three mainstream large models. This is OpenAI's first public quantitative comparison since entering the custom silicon space, marking the AI company's formal extension of its technological competition into the hardware layer.

700W vs 1400W: The Power Gap Is Key to Understanding the Performance Data

Jalapeño's core design logic is to perform equal or more inference work at lower power consumption. According to OpenAI's official data, the Jalapeño has a rated power of 700 watts, with actual power consumption remaining below 550 watts during testing. By comparison, the NVIDIA GB200 has a rated power of 1,200 watts, and the GB300 has 1,400 watts. Under the same power budget, Jalapeño can run at a much higher instance density than NVIDIA's solutions, which is the fundamental source of its lead in "throughput per watt."

In terms of memory and bandwidth specifications, each Jalapeño package integrates six HBM4 memory stacks, with a total capacity of 216 GiB and bandwidth of 15.4 TB/s. This bandwidth figure is critical for large model inference—large models are typically memory-bandwidth-bound tasks during the inference phase, rather than compute-bound. High-bandwidth memory directly determines token generation speed in each inference pass and is the hardware foundation that enables Jalapeño to achieve both low latency and high throughput. The chip was co-designed with Broadcom as the ASIC partner, manufactured on TSMC's 3nm process, with Samsung reportedly supplying the HBM4.

Specific Numbers for the Three Models: The Gap Varies by Scenario

The InferenceX benchmark covers three open-source models ranging from mid-size to ultra-large scale, and the results show clear and systematic differences.

In the GPT-OSS 120B test, Jalapeño achieved approximately 1.9x peak per-kilowatt efficiency over the GB200—specifically 85,448 vs. 44,960 mixed TPS/kW, with end-to-end latency dropping from 1.8 seconds to 1.03 seconds. This is the scenario with the largest efficiency advantage across the three test groups, and also the most common production form of mid-size inference workloads.

In the DeepSeek R1 670B test, Jalapeño achieved a 1.7x efficiency advantage over the GB300 (19,641 vs. 11,781 mixed TPS/kW), with latency dropping from 5.99 seconds to 1.65 seconds, a 3.6x reduction—the most significant latency improvement across all the data. DeepSeek R1 is currently a heavyweight inference model widely deployed in the industry, and a 3.6x latency reduction carries substantive significance for the user experience of conversational products.

The third test group used the Kimi K2.5 trillion-parameter model, which is also one of the largest-scale inference scenarios in current public testing. OpenAI's published summary data shows that Jalapeño maintains its advantage, but specific TPS/kW figures for this model have not yet appeared in public sources.

Additionally, OpenAI disclosed a subset of data for "interactive workloads": for the most latency-sensitive task types, Jalapeño's performance metrics reached 2.1 to 4.1 times those of the GB200/GB300. This figure is higher than the overall benchmark, reflecting the chip's design orientation toward low-latency scenarios.

Impact on the Competitive Landscape: NVIDIA Is Safe in the Short Term, but Long-Term Pressure Is Beginning to Be Priced In

Looking at the test numbers alone, Jalapeño represents a technically meaningful challenge to NVIDIA. But understanding the real impact of this data on the industry landscape requires distinguishing between two time horizons.

In the short term, NVIDIA's position is barely affected. In its announcement, OpenAI explicitly stated that actual deployment volumes before the end of 2026 will be extremely limited, and the company will continue to purchase GPUs in large quantities from NVIDIA and other partners. Large-scale training tasks still rely entirely on NVIDIA hardware; Jalapeño is positioned as a pure inference chip and does not involve training scenarios. This means the practical impact on NVIDIA's revenue is nearly zero in the short term.

The medium-to-long term is another story. OpenAI has also disclosed that the second-generation Jalapeño chip may be only months away from tape-out, and the third-generation design has already been initiated. This cadence indicates that Jalapeño is not a one-off engineering demonstration but has entered a product roadmap of continuous iteration. If OpenAI's inference scale continues to expand, the proportion of internal chips replacing externally purchased GPUs will gradually rise, eroding precisely the data center business where NVIDIA's profit margins are highest.

For NVIDIA, the more direct signaling pressure comes from the market psychology level. Previously, the emergence of Google TPU and Amazon Trainium/Inferentia had already made the market aware that hyperscale AI companies are capable of developing their own chips, but cases of directly competing with and beating NVIDIA's flagship on public benchmarks are uncommon. By choosing to publish its numbers on InferenceX, a third-party framework, OpenAI's data is at least traceable, rather than a closed internal whitepaper.

What It Means for Developers and Enterprise Users

Jalapeño is not currently for sale, nor has any public cloud service integration plan been announced. For the vast majority of developers and enterprise users, the only change they may perceive in the short term is this: if OpenAI deploys Jalapeño into its API service infrastructure, inference latency is expected to drop, especially for real-time conversational scenarios that are sensitive to response speed.

More direct is the cost transmission path. Improvements in per-watt efficiency of inference chips directly correspond to lower electricity costs per unit of compute in data centers. If Jalapeño reaches large-scale production, OpenAI's marginal inference cost will decline, theoretically providing more room for API pricing compression. However, mass production timelines, actual deployment scale, and cost data have not been disclosed, so this path remains at the hypothesis stage.

For enterprises evaluating inference infrastructure, the most immediate reference value of Jalapeño is that it redefines the "industry feasible baseline": the inference efficiency achievable on a 3nm process with 700 watts of power, 216 GiB of HBM4, and 15.4 TB/s of bandwidth will become a new frame of reference for enterprises when evaluating NVIDIA procurement options and assessing cloud vendors' self-developed chip roadmaps.

The Boundaries of Self-Reported Data

There is a fundamental limitation to Jalapeño's first batch of data: all tests were conducted by OpenAI itself, using the InferenceX framework designed by SemiAnalysis, but test execution, hardware configuration, and tuning parameters were all controlled by OpenAI. This is similar in nature to NVIDIA's official GPU whitepapers—traceable methodology, but lacking independent reproduction.

Key information not yet disclosed includes: Jalapeño's results when compared against NVIDIA hardware running highly optimized inference engines (such as TensorRT-LLM), system-level TCO (total cost of ownership) comparisons, and actual performance on proprietary models (such as the GPT series) rather than open-source models. The benchmark's selection of DeepSeek R1 and Kimi K2.5—two open-source models from Chinese companies—is itself a signal: these models have public weights, facilitating third-party validation, while also covering the most closely watched inference workload types in the current industry.

Strategic Assessment: Which Signal to Track Next

The most likely strategic path for Jalapeño is to first replace part of NVIDIA's inference capacity within OpenAI's own API infrastructure, to control marginal costs, while not interrupting the training side's dependence on NVIDIA. This "internal absorption" model closely aligns with the early deployment path of Google's TPU—Google also first absorbed self-developed chip capacity in search and internal products, then gradually opened it up externally through Cloud TPU.

The tape-out milestone of the second-generation chip is the key signal most worth tracking. If the second generation can be completed and enter mass production readiness before 2027, it would mean that OpenAI's silicon engineering capability has formed a stable iteration cadence, at which point the impact on NVIDIA procurement volumes would move from theoretical to actual financial figures.

Another dimension to watch is pricing signals: if OpenAI lowers API inference prices after large-scale Jalapeño deployment, that would be direct evidence of efficiency dividends being passed through to users. Conversely, if API prices remain unchanged, it would indicate that efficiency gains are being used to expand profit margins rather than benefit the market. These two paths have completely different implications for enterprise technology selection strategies.