OpenAI has just published the first measured results of Jalapeño, its first chip designed specifically to run artificial intelligence models. The company claims significant gains in speed and energy efficiency. This strategic advancement could improve the responsiveness of AI agents and reduce ChatGPT's operating costs, but it does not signal the end of its dependence on Nvidia.
Key points to remember
- Jalapeño is the first AI chip designed by OpenAI for inference, that is, the execution of already trained models.
- OpenAI claims 1.5 to 1.9 times more AI work per watt and end-to-end latency up to 3.6 times lower in published tests.
- Nvidia remains indispensable : Jalapeño diversifies OpenAI's infrastructure but does not replace the accelerators used for training.
What is Jalapeño, OpenAI's AI chip?
The battle for artificial intelligence is no longer solely about the quality of the models. It is shifting towards the infrastructures capable of running them on a large scale: data centers, networks, memory, energy and, now, custom-designed chips.
It is with this in mind that OpenAI is developing Jalapeño with Broadcom. The chip was presented in June 2026. On August 25, the company published its first detailed results and claims to now have an operational processor, intended for gradual deployment in its own infrastructure.
Jalapeño is not a chip intended to train future OpenAI models. It focuses on the inference : the work done when a pre-trained model analyzes a request and generates a response. Every conversation with ChatGPT, every request sent to an API, and every step executed by an AI agent thus consumes inference resources.
More AI work for every watt consumed
According to measurements published by OpenAI, Jalapeño reportedly delivered between 1.5 and 1.9 times more AI work per watt at its maximum throughput level than the systems used for comparison. The company also announces end-to-end latency between 1.7 and 3.6 times weaker. For highly interactive uses, it claims a performance gain ranging from 2.1 to 4.1 times.
The tests were conducted on three publicly available models of different sizes and origins: GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. OpenAI states that it used InferenceX, a publicly available benchmark developed by the specialized analytics firm SemiAnalysis. SemiAnalysis explains that it tested the chip in OpenAI's labs with the company's engineers and considers Jalapeño to be particularly competitive among the systems evaluated.
These figures should nevertheless be presented with their scope in mind. They describe tests carried out on certain models, in a controlled environment, and before a large-scale deployment. They do not prove that Jalapeño will dominate all accelerators for all uses.
Why this chip is strategic for ChatGPT and AI agents
For a typical chatbot, a fraction of a second can already alter the perceived fluidity. For an agent that has to perform dozens or hundreds of actions, every delay adds up. A chip capable of producing more tokens while consuming less electricity could therefore make agents faster, more available, and less expensive to operate.
OpenAI is also seeking to better control the economics of its services. The financing of its ambitions and the burden of its infrastructure costs For several years, this has been a strategic priority. By designing models, software, memory, network, and processors together, the company can optimize the entire chain rather than adapting its products to general-purpose hardware. It also claims to have used its own AI models to accelerate certain design and optimization stages of Jalapeño.
This does not guarantee an immediate price reduction for ChatGPT or the API. Any potential savings will depend on the chip's manufacturing cost, its actual throughput at scale, data center consumption, and how OpenAI passes these savings on to its customers.
Nvidia remains indispensable
Jalapeño does not, at this stage, represent a complete break with Nvidia. The chip is designed for inference and not for training the most advanced models, an activity for which OpenAI continues to require considerable computing power.
In an interview with Axios, Richard Ho, OpenAI's vice president of hardware, indicated that the company will continue to combine multiple vendors and technologies, including Nvidia, AMD, and other specialized infrastructure. OpenAI plans to have a limited number of Jalapeño systems in operation in 2026, before increasing capacity the following year. A second generation is already under development.
The strategy therefore looks more like diversification than replacement: having an in-house solution for the largest inference volumes, while retaining partner chips for the workloads for which they are best suited.
The AI war is also becoming an energy war
Google has its TPUs, Amazon its Trainium and Inferentia chips, while Microsoft is also developing its own accelerators. With Jalapeño, OpenAI joins the ranks of players who want to control an increasing share of their infrastructure.
The stakes go beyond simply competing with Nvidia. Data centers' power needs are increasing rapidly, and access to energy is becoming a major obstacle to AI development. Generating more responses for every watt consumed can therefore become as important as gaining a few points on a reasoning benchmark.
Jalapeño does not signal the end of Nvidia. But its initial results show that OpenAI now wants to control the entire artificial intelligence value chain, from the model to the silicon that generates each response.


