From Software to Silicon

On June 24, 2026, OpenAI unveiled its first-ever custom-built processor — a decisive move from software leader to full-stack infrastructure player. Developed in collaboration with Broadcom over an accelerated nine-month cycle, the chip — codenamed Jalapeño — is a reticle-sized ASIC purpose-built for large language model inference.

Unlike Nvidia's general-purpose GPUs, Jalapeño is optimized exclusively for running pre-trained models in response to user commands. Early testing shows significantly better performance-per-watt than current state-of-the-art alternatives, OpenAI said.

Why Inference Matters

Inference — the process of running AI models to answer queries — is where the bulk of AI computing costs now lie. As OpenAI scales products like Codex, ChatGPT, and agentic AI tools, shaving even small percentages off inference costs has massive bottom-line impact. The company noted that Jalapeño's architecture is especially efficient for real-time coding model workloads.

"We have a deep understanding of the workload," OpenAI President Greg Brockman said. "We've really been looking for specific workloads that are underserved — how can we build something that will accelerate what's possible?"

Full-Stack Ambition

Jalapeño represents more than a chip: it's the physical manifestation of OpenAI's vertical integration strategy. The company now designs chip architecture, kernels, memory systems, networking, scheduling, and deployment — all optimized around a single goal: faster, more reliable, more affordable AI.

Training of frontier models will still rely on Nvidia hardware for now. But even modest inference savings compound enormously at OpenAI's scale. With a heavily anticipated public offering in 2026, owning its silicon sends a powerful signal to investors about long-term margin control.

Industry Context

OpenAI joins Google (TPU), Amazon (Trainium/Inferentia), and Microsoft (Maia) in building custom AI accelerators. What sets Jalapeño apart is its pure inference focus and its sub-year development timeline, accelerated by OpenAI's own models aiding the chip design process.

The chip is expected to begin deployment in late 2026, ramping up over subsequent years.