[Explanation] OpenAI’s first AI chip, Jalapeño,

IT

OpenAI, in collaboration with major U.S. semiconductor company Broadcom, has announced its first independently designed AI inference chip, Jalapeño. This move is driven by a strategic goal to address severe GPU shortages and reduce costs, reduce dependence on the market-dominant NVIDIA, and build infrastructure optimized for its own AI models.

Unprecedented Rapid Development and Partnerships

OpenAI’s announced “Jalapeño” is a purpose-built integrated circuit (ASIC) designed from scratch to streamline inference processing for large language models (LLMs). This project was carried out in close collaboration with Broadcom, a major U.S. semiconductor company, and Celestica, a manufacturing support provider. Notably, it took an extraordinary speed—just about nine months—from initial design to manufacturing tape-out. While it is said that designing an ASIC from scratch typically takes 1.5 to 2 years, by optimizing parts of the development process using OpenAI’s own AI models, we achieved this ultra-short development cycle. Richard Ho, OpenAI’s Head of Hardware, emphasizes that deep collaboration with the company’s research team enabled designs that deeply reflect the fundamental principles of LLMs, such as model kernels, data movement, and networks.

A unique architecture specialized for inference processing

Jalapeño is not a general-purpose learning accelerator, but is built specifically for “inference” workloads like ChatGPT’s response generation. Inference, unlike model training, is a crucial process that influences the enormous computational costs of daily service operations. Jalapeño’s architecture is designed to minimize data movement and optimize the balance of computation, memory, and network resources, achieving real-world utilization rates very close to theoretical peak performance. Specifically, a design called the “systrick array,” suitable for repeated matrix operations, is adopted, efficiently executing mathematical processing unique to inference. This enables both high throughput and low latency, making it well-suited for the operation of “agent-type AI” that will perform complex inference in the future.

Physical structure and cutting-edge manufacturing processes

The manufacturing of Jalapeño is expected to use TSMC’s cutting-edge 3nm technology. According to analysis of released wafer images and other materials, this chip is structured around a single massive compute chiplet, surrounded by six high-bandwidth memory (HBM) modules. Please refer to the diagram below.

Figure 1

The size of the compute chip is estimated to be about 840 square millimeters, which is close to the reticle limit of semiconductor lithography equipment and is considered extremely large for inference chips. While many inference accelerators use inexpensive DRAM, Jalapeño chose expensive and high-speed memory like HBM3/4 to maximize data transfer speeds, which are bottlenecks during inference. Currently, engineering samples are already running in the lab, successfully processing the latest machine learning workloads, including GPT-5.3-Codex-Spark, at frequencies and power levels designed for mass production.

Market Impact and Strategic Significance

Outstanding Cost Efficiency and Performance Claims

The greatest benefit of implementing Jalapeño is the dramatic reduction in operating costs. In an interview, Broadcom CEO Hoc Tan stated that initial tests show that it is expected to achieve about 50% cost reduction compared to conventional general-purpose GPUs. It is also claimed to significantly surpass current cutting-edge hardware in terms of performance per watt, that is, power efficiency. With strong competitors like Nvidia’s Blackwell and AMD’s Instinct MI350, owning proprietary chips optimized for specific models offers OpenAI a significant economic advantage. Even a slight reduction in inference costs can be an inflection point at OpenAI’s scale, which processes billions of queries daily, significantly impacting the final profit margin.

Breaking Free from Dependence on NVIDIA and “Project Nexus”

The development of Jalapeño is positioned as part of a vast AI infrastructure initiative called “Project Nexus.” OpenAI has long been one of Nvidia’s largest GPU customers, but the shift to in-house designed chips serves as a strong check on Nvidia’s pricing power. In October 2025, OpenAI signed a multi-year agreement with Broadcom to deploy AI accelerators and rack systems with a maximum capacity of 10 gigawatts (GW), making Jalapeño the first-generation product aimed at in-house this strategic hardware stack. However, OpenAI still relies on NVIDIA for the “training” of its cutting-edge models, and rather than replacing all chips with its own chips, it is believed to aim to diversify suppliers in certain high-cost areas such as inference and enhance its bargaining power.

Implementation Challenges and Future Outlook

Deployment Schedule and Supply and Funding Barriers

Jalapeño’s full-scale rollout is scheduled to begin between late 2026 and 2027. In collaboration with Microsoft and other partners, deployment will be advanced to gigawatt-scale data centers. However, there are also several challenges ahead of this grand ambition. Even the first phase of the project alone requires a massive funding of approximately $18 billion, supported by financial frameworks such as Broadcom, Apollo Global Management, and Blackstone. Additionally, TSMC’s 3nm process production slots are fiercely competitive with other companies like Apple, making the focus on securing supply as scheduled. In fact, some slides have been reported to roll out the initial batch from the initial late 2026 to 2027.

Roadmap to Next-Generation AI Infrastructure

Jalapeño is not a one-off project, but just the beginning of a long-term hardware roadmap. Broadcom CEO Hoc Tan revealed that next-generation chip development is already planned for 2028, with annual updates planned thereafter. OpenAI is transforming into a “full-stack” company that vertically integrates everything from model development to chip design and product delivery. This move aligns with the strategies of hyperscalers like Google and Meta, which already operate proprietary AI chips (TPUs and MTIA). Jalapeño’s success will serve as a touchstone to see if we can achieve “speed,” “high reliability,” and “lower cost” of AI services through proprietary hardware, making intelligence more affordable and accessible to society.

[#OpenAI #Jalapeño #AIチップ #半導体 #ブロードコム #人工知能 #大規模言語モデル #エヌビディア]

コメント

Copied title and URL