OpenAI Jalapeño Chip Benchmark Debut: 700W Showcases Efficiency Advantage Over GB300

OpenAI and Broadcom's Jalapeño inference ASIC posted its first benchmark results at Hot Chips on August 25, 2026, delivering 1.9x higher per-watt inference throughput than NVIDIA's GB300 at just 700W versus 1400W.

The Jalapeño inference-specialized ASIC, developed by OpenAI in collaboration with Broadcom, released its first batch of benchmark data at the Hot Chips conference on August 25, 2026. In the InferenceX tests spanning GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, it led NVIDIA's GB300 by 1.9x in inference throughput per watt, achieved 1.7-3.6x lower end-to-end latency, with the gap widening to 4.1x under high-interaction workloads—all while consuming just 700W of power compared to the competitor's 1400W.

Facts

According to TechCrunch, Jalapeño was developed jointly by OpenAI and Broadcom, utilizing HBM4 memory supplied by Samsung. Benchmark results show the chip reduces data movement during the prefill and communication phases, allowing model states and KV cache to be placed locally with corresponding compute, memory, and network resources activated on demand. OpenAI's head of hardware, Richard Ho, noted in a press call that the results represent a significant performance leap over existing state-of-the-art inference processors.

The deployment timeline is clear: small-scale integration into ChatGPT infrastructure by end of 2026, with expansion in 2027. The platform is positioned as a multi-generational product with coordinated development across models, chips, and memory.

Mechanism Breakdown

Jalapeño is designed around common bottlenecks in the inference pipeline. By minimizing data movement and communication latency, the system can dynamically combine resources across different stages rather than relying on a fixed architecture. This contrasts with traditional GPUs, which employ unified scheduling throughout the entire workflow—explaining why the chip can simultaneously improve per-watt throughput and reduce latency on identical models.

The 700W versus 1400W power comparison directly reflects the difference in workload-carrying capacity under power constraints. The integration of HBM4 memory further supports larger local retention of model states, reducing cross-chip communication overhead.

Industry Impact

This result will affect compute-economics assessments that rely on public benchmark rankings. Power consumption accounts for a significant share of inference cost structures, and the efficiency advantage demonstrated by Jalapeño could alter the marginal cost curve of services such as ChatGPT, with ripple effects on cloud service pricing and the competitive landscape currently dependent on NVIDIA's GB300 series.

Samsung's role as the HBM4 supplier reportedly underscores a broader trend toward supply chain diversification. Other large-model vendors may reassess their self-developed ASIC paths, particularly as inference workloads continue to grow as a share of overall compute demand.

Strategic Assessment

[Analysis] If Jalapeño ramps as planned in 2027, its multi-generational platform strategy could form a closed-loop advantage. However, NVIDIA's subsequent product iterations remain uncertain, and actual market share shifts will depend on deployment scale and real-world workload performance. OpenAI's full-stack co-design model offers a reference for the industry, but the difficulty of replication lies in the ability to synchronously iterate models and hardware.

[Analysis] For the broader AI infrastructure landscape, this development signals that inference-specialized chips are moving from the periphery to the core. Cost structure reshaping will transmit through to model training and service pricing, requiring stakeholders to make longer-term portfolio decisions in hardware selection.