Product signal

OpenAI’s Jalapeño Debuts Against Blackwell

OpenAI released initial Jalapeño inference-chip results, claiming higher work per watt and lower end-to-end latency than NVIDIA Blackwell across three public models.

OpenAI has published its first public performance results for the Jalapeño AI server chip and compared it directly with NVIDIA Blackwell. The company says Jalapeño delivered 1.5x to 1.9x more peak work per watt and 1.7x to 3.6x lower end-to-end latency across GPT-OSS, DeepSeek, and Kimi. That moves the program beyond the engineering-sample stage described in June.

From sample to measurable platform

When OpenAI and Broadcom introduced Jalapeño in June, they described an engineering sample running machine-learning workloads at target frequency and power, with final performance still being measured and initial deployment planned for late 2026. The new disclosure adds a 700W rated power figure, sustained test consumption of no more than 550W, and throughput and latency comparisons for three public models. Manufacturing deployment has not yet changed, but the performance claim is now specific enough for external scrutiny.

Inference efficiency is the entry point

Jalapeño targets inference rather than training. For services processing a sustained volume of model requests, work per watt and end-to-end latency shape server counts, data-center power demand, and user experience. By testing GPT-OSS, DeepSeek, and Kimi, OpenAI is attempting to show that the design is not limited to an internal model. If the results hold under operational loads, some inference demand could shift from general-purpose accelerator purchases toward infrastructure designed around OpenAI’s own serving patterns.

Production validation remains incomplete

SemiAnalysis said it witnessed InferenceX running at OpenAI but did not run the complete test suite or observe AgentX results. It also argued that HBM4 makes NVIDIA Rubin a more relevant comparison than Blackwell, while an 8k-input, 1k-output single-turn test cannot represent long-context and multi-turn workloads that stress cache, routing, and offload systems. The supported conclusion is that Jalapeño has made a competitive efficiency claim, not that it has completed a broad replacement test.

What to watch next

Observable next evidence includes whether OpenAI begins initial deployment by late 2026, releases long-context and multi-turn agent results at different concurrency levels, and reports production power, availability, and cost per token. Continued support across those measures would strengthen the case that Jalapeño can alter OpenAI’s sourcing mix. Delays or weaker operating economics would make it primarily a technical validation.

Sources