News report
OpenAI Pairs Jalapeño ASIC Racks With AMD EPYC Turin Hosts
OpenAI's Jalapeño deployment uses dual-EPYC Turin host trays with 1.5TB of DRAM, prioritizing platform maturity while Vera remains under evaluation.
On this page
Jalapeño's production system uses AMD EPYC Turin host trays
OpenAI's first-generation Jalapeño inference ASIC is being deployed in a rack-scale system whose host side uses AMD EPYC Turin-class processors. SemiAnalysis documents a paired layout: a host rack contains 16 Katsu CPU trays, each corresponding to one of 16 Vindaloo accelerator trays in the adjacent ASIC rack.
Each Katsu host tray carries two Turin-class AMD EPYC CPUs and 1.5TB of DRAM, plus two E1.S and two M.2 SSDs. The host and accelerator trays connect through external PCIe links. The important point is architectural rather than a consumer-CPU win: Jalapeño is an accelerator, while the EPYC systems provide the mature host-compute layer around it.
| Layer | Reported configuration | Role |
|---|---|---|
| Host tray | 2× AMD EPYC Turin-class CPUs, 1.5TB DRAM | CPU host for a corresponding accelerator tray |
| Host rack | 16 Katsu trays | One host tray per Vindaloo accelerator tray |
| Accelerator tray | 8 Jalapeño ASICs | LLM inference compute |
| Accelerator rack | 16 Vindaloo trays | 128 Jalapeño ASICs per rack |
| Host-to-accelerator link | External PCIe connections | System I/O between paired trays |
OpenAI says the Turin choice was about reducing deployment risk
OpenAI hardware VP Richard Ho told Tom's Hardware that selecting Turin was a pragmatic decision intended to de-risk the Jalapeño system and move quickly. He said the Turin platform did what OpenAI needed and that its partners already had experience with it.
Ho contrasted that maturity with NVIDIA's standalone Vera CPU, describing Vera as being somewhat behind on maturity for this particular design decision. That should not be read as a general performance verdict on Vera versus EPYC: OpenAI is also among organizations exploring Vera, and the comments concern the host choice for the current Jalapeño deployment.
The host choice fills in a missing part of OpenAI's custom-silicon story
OpenAI and Broadcom unveiled Jalapeño in June as OpenAI's first Intelligence Processor, built specifically for LLM inference and intended for multi-generation, gigawatt-scale deployment. OpenAI's announcement established the accelerator program but did not make the CPU host the headline.
The newly detailed rack architecture shows that custom AI silicon does not eliminate the general-purpose host layer. OpenAI is combining its purpose-built inference accelerator with an established x86 server platform while the custom accelerator, networking and software stack mature together.
Jalapeño remains an inference accelerator, not a replacement for every GPU workload
OpenAI positions Jalapeño around LLM inference. SemiAnalysis has published detailed architecture and performance analysis after observing the system and benchmarking in OpenAI's lab, but it also notes important evidence boundaries: some benchmark inputs and numbers came from OpenAI, and the full preferred AgentX suite had not been run.
That distinction matters when interpreting the rack design. The EPYC host decision is concrete deployment information; it does not establish that x86 is universally superior for AI hosts, that Vera will not be used in later OpenAI systems, or that Jalapeño replaces NVIDIA and AMD accelerators across training and other workloads.
Sources
Primary and technical sources
These sources support the reporting and analysis above. Current stories are updated when later evidence materially changes the facts.
Related
Continue from here
Useful next steps selected from the same technical reference and publication system.
Technical guide
16 GB vs 32 GB vs 64 GB RAM for Gaming PCs
Choose 16 GB, 32 GB, or 64 GB of system RAM for a gaming PC by measuring the games and simultaneous workloads you actually run instead of relying on a universal capacity rule.
Tool
DDR Memory Latency Calculator
Convert DDR data rate and CAS latency cycles into CAS timing in nanoseconds.
Technical guide
Windows Page File Explained: Virtual Memory and Commit Limit
Understand what the Windows page file does, how it extends the system commit limit, how paging differs from RAM use, and why crash dumps can depend on it.
Tool
DDR Memory Bandwidth Calculator
Calculate theoretical peak DDR memory bandwidth from transfer rate, bus width per channel, and active channel count.