Industry

OpenAI and Broadcom unveil Jalapeño inference chip designed for lower-energy AI serving

OpenAI has introduced Jalapeño, its first custom inference chip developed with Broadcom and tailored to the company’s serving workloads.

OpenAI and Broadcom unveil Jalapeño inference chip designed for lower-energy AI serving

OpenAI has revealed its first custom inference chip, developed and manufactured in collaboration with Broadcom. The processor, named Jalapeño, was specifically optimized to meet the particular requirements of OpenAI’s inference workloads.

According to the company, OpenAI’s own AI models were used in the chip’s development. The processor is currently in a testing phase, but OpenAI says early results indicate that it significantly outperforms current leading alternatives in energy efficiency — measured as performance per watt.

Reducing dependency on NVIDIA

The OpenAI–Broadcom partnership was announced officially last October, though industry speculation about the company’s chip ambitions had circulated earlier. The stated objective is to reduce OpenAI’s reliance on NVIDIA processors, which today form the dominant hardware foundation of the AI industry.

Google and Amazon Web Services followed a similar strategy in the past, developing their own AI accelerators — specialized chips designed to speed up machine learning tasks.

Optimized for OpenAI’s workloads

Greg Brockman, President of OpenAI, has said the company leverages deep knowledge of its own workloads. According to Brockman, the team looked for particular tasks that lack adequate hardware support today and aimed to build a solution that pushes the boundaries of what’s possible.

Jalapeño is purpose-built for inference, the process by which a pre-trained AI model executes computations to respond to user instructions. OpenAI emphasized that the chip can run models optimized for programming tasks in real time with very low operating costs.

Limits and economic impact

It is likely that more computationally intensive operations, such as pretraining and training new large models, will continue to run on NVIDIA hardware. Nevertheless, even modest reductions in inference costs can materially improve the company’s profitability, because inference is a major driver of operating expenses for AI services.

Industry trends suggest that optimizing inference systems will be a key factor in the economics of artificial intelligence going forward, and that this optimization will occur across all layers of technology infrastructure.

Pursuing control across the full technology stack

OpenAI states its strategy extends well beyond model development: the company intends to design the infrastructure beneath the models as well. Its statement lists chip architecture, low-level runtimes (kernels), memory systems, network infrastructure, scheduling mechanisms, deployment systems, and user experience among the components it designs.

OpenAI argues that this end-to-end visibility and control of the full technology stack is a competitive advantage: when every layer can be optimized for the same objectives, models can become faster, more reliable, and more affordable for users.

Current status

Jalapeño remains in testing, and OpenAI’s initial reported results point to a significant improvement in energy efficiency relative to existing top-tier alternatives. The company has not provided a detailed timeline for broader deployment.