At the recent Hot Chips 2026, OpenAI presented the full architecture and the first published benchmarks for Jalapeño, its inaugural custom AI accelerator. OpenAI announced the chip’s name and confirmed Broadcom as its silicon partner earlier this year, following a strategic collaboration the two companies announced in 2025 to deploy up to 10 gigawatts of OpenAI-designed accelerators.
The Hot Chips presentation filled in details withheld at the original reveal, including per-chip compute and memory specifications, rack- and pod-level system topology, the manufacturing process node, and head-to-head performance comparisons against Nvidia’s GB200 and GB300 platforms.
Jalapeño targets inference workloads specifically, while OpenAI continues to rely on NVIDIA and other merchant GPU suppliers for model training.
Its architecture reflects a bet that a single “balanced” chip design can support prompt processing, token generation, and low-latency draft models used in speculative decoding, without a heterogeneous fleet of specialized parts.
This matters for two reasons. It is OpenAI’s clearest public step toward supplementing merchant GPU supply with silicon it controls end-to-end, and it confirms Broadcom’s position as the leading merchant partner for custom silicon for frontier AI labs and hyperscalers, a role it already holds with Google and Meta.
Limited production deployment is planned for late 2026, with broader rollout in 2027, and OpenAI says a second-generation Jalapeño is already deep into development, with a third generation underway.
Technical Details
OpenAI and Broadcom engineers jointly presented Jalapeño at Hot Chips, covering the accelerator die, the local and global interconnect fabric that ties individual chips into racks and multi-rack pods, and the AI-assisted design methodology OpenAI used to compress its development timeline
Compute and memory

Each Jalapeño accelerator delivers 13.4 petaFLOPS of MXFP4 matrix compute, paired with six HBM4 stacks totaling 216 GiB of capacity and 15.4 TB/s of memory bandwidth, with a rated thermal design power of 700W. OpenAI said that measured sustained power during testing remained at or below 550W.
Manufacturing
Broadcom implemented the design on a TSMC 3nm-class process and completed the RTL-to-tapeout cycle in nine months.
Rack-scale local domain
A base unit of 128 accelerators forms a local compute domain, delivering 1.7 exaFLOPS of 4-bit compute and 27.5 TB of aggregate HBM4 memory over a 600 GB/s local collective network. The KV cache is kept resident in this domain to minimize data movement between inference phases.
Pod-scale global domain
A full production pod scales to 2,048 chips across 16 racks, delivering 27 exaFLOPS of aggregate compute and 432 TiB of aggregate memory. Chips connect across racks via Broadcom Tomahawk6 Ethernet switches arranged in a half-flattened two-level Clos topology running at 200 GB/s.
Inference-phase balancing
Jalapeño activates or gates specific compute, memory, and networking blocks depending on whether it is running compute-bound prompt processing, memory-bandwidth-bound token generation, or low-batch draft models used in speculative decoding. OpenAI describes this as reducing idle power draw across mixed inference workloads on a single-chip design.
AI-assisted design

In one of its most interesting disclosures, OpenAI said it used its own models in a “continuous convergence” workflow built on the XLS hardware description language, and that this approach delivered a 56 percent performance improvement in BF16 multiply operations compared with a human-engineered baseline.
Performance
OpenAI reported 1.5x to 1.9x higher throughput per kilowatt than Nvidia’s GB200 and GB300 systems at matched operating points across a set of models that included GPT-OSS-120B, with gains rising to as much as 8.6x to 104.3x at peak-efficiency configurations.
OpenAI also reported 1.7x to 3.6x lower end-to-end latency and 2.1x to 4.1x faster performance on ultra-low-latency benchmarks. Measured at the full-system level, including utility power beyond chip TDP, OpenAI reported Jalapeño at 1.18kW, compared with 2.55kW for a comparable GB300 configuration.
Analysis
Jalapeño extends OpenAI’s broader full-stack infrastructure strategy, which already includes long-term compute commitments to Nvidia, AMD, and Cerebras, as well as its own data center buildout with partners such as Microsoft, Oracle, and SoftBank through the Stargate program.
Owning inference silicon gives OpenAI leverage against merchant GPU pricing and allocation constraints and lets it tune hardware to the specific prefill and decode characteristics of its own model family, a level of specialization that a general-purpose GPU cannot match.
OpenAI President Greg Brockman has framed the effort as part of a long-term strategy to make compute more abundant. OpenAI’s confirmation that a second-generation Jalapeño is already deep into development, with a third generation underway, points to a sustained multi-year roadmap.
The risk is one of scale. OpenAI hardware lead Richard Ho has said the company still lacks sufficient internal compute to walk away from competing hardware purchases. The 10-gigawatt Broadcom target, sizable in absolute terms, remains a fraction of the capacity OpenAI has committed to secure from other suppliers through the end of the decade, including a $30 billion Nvidia investment announced in early 2026, tied to 10 gigawatts of Nvidia-based capacity.
Final Thoughts
Jalapeño is a laudable technical achievement and a significant step in Broadcom’s expansion beyond networking into the center of the custom AI silicon market. The specifications OpenAI disclosed at Hot Chips, including a 700W-rated part delivering 216 GiB of HBM4 at 15.4 TB/s and scaling to 27 exaFLOPS across a 2,048-chip pod, are competitive on paper with the highest-end merchant GPU platforms. The nine-month design cycle is a notable engineering achievement, separate from the benchmark claims layered on top.
The open questions concern scale, independence, and durability. Jalapeño’s near-term production volume, even against a 10-gigawatt target spread through 2029, remains small compared with Nvidia, AMD, and Cerebras capacity that OpenAI has separately committed to secure, a gap that OpenAI’s own hardware leadership has publicly acknowledged.
Whether Jalapeño becomes a durable share of OpenAI’s inference fleet or stays a strategic hedge depends on how the second and third generations perform, and on whether Broadcom’s Tomahawk6-based interconnect scales cleanly beyond the 2,048-chip pod size shown at Hot Chips.
For Broadcom, Jalapeño confirms that its custom silicon franchise, already anchored by Google and Meta, can bring a frontier AI lab from a standing start to a working accelerator in under a year.
That capability, more than any single benchmark chart from Hot Chips, is what makes Jalapeño consequential for the rest of the AI infrastructure market, where every major model developer is now weighing custom silicon against continued reliance on a single merchant GPU supplier.



