AdvertisementAdvertisementAdvertisement
Technology

OpenAI Publishes First Benchmark Results for Its Own Chip

8/25/2026, 05:55 PM • Evgenia Sliv

(edited: 08/25/2026)

OpenAI Publishes First Benchmark Results for Its Own Chip

OpenAI has published the first performance benchmark results for the Jalapeño processor — the company's first proprietary semiconductor chip, developed specifically for AI model inference tasks. According to the data presented, the new architecture demonstrates significant gains in both data processing speed and energy efficiency, enabling it to overcome the traditional trade-off between system throughput and response latency compared to leading commercial solutions in the AI hardware market.

Performance evaluation was conducted using the InferenceX framework — an open benchmark developed by analytics firm SemiAnalysis. The testing protocol covered three different large language model architectures: GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T.

In benchmark results, the Jalapeño chip delivered 1.5 to 1.9 times more AI compute operations per watt of power consumption at peak throughput. Additionally, end-to-end latency was reduced by 1.7 to 3.6 times compared to competing hardware systems across all three test models. For highly interactive workloads, performance was 2.1 to 4.1 times higher.

The most notable results came from testing on the Kimi K2.5 1T model, the largest of the publicly available architectures evaluated. In this scenario, Jalapeño demonstrated approximately 1.5 times higher peak performance per watt and 3.4 times lower end-to-end latency relative to competing systems. The chip's rated thermal design power is stated at 700 watts; however, during testing, actual sustained power consumption under load did not exceed 550 watts, confirming the high efficiency of its power management.

OpenAI representatives noted that the chip's architectural advantage becomes even more pronounced when running the company's advanced in-house models. This indicates that the microarchitecture scales and improves in efficiency as the complexity and volume of computational workloads increase.

The development process for Jalapeño, from initial design to tape-out, took nine months. A distinguishing feature of this cycle was the extensive use of proprietary AI tools for hardware design and optimization. Earlier generations of OpenAI models were applied during the initial design phase, while newer algorithm versions accelerated optimization and low-level programming processes.

In particular, using the Codex and GPT-Astra tools, the engineering team was able to bring three open-weight models to high performance levels in just two months. For individual compute units — such as attention mechanisms and the mixture-of-experts architecture in the GPT-OSS implementation — AI-generated code ran 1.5 to 1.8 times faster than existing implementations written manually by engineers.

OpenAI plans to begin integrating and deploying Jalapeño chips across its global computing infrastructure before the end of this year. In parallel, the company confirmed that development of a second generation of the processor is already underway, with a third generation currently in the conceptual stage. OpenAI also stated its intention to maintain strategic partnerships with existing hardware suppliers: Nvidia accelerators and those of other technology partners will continue to be used for model training and inference alongside the proprietary Jalapeño solutions, forming a heterogeneous computing environment.

Popular news