OpenAI publishes the first Jalapeño chip benchmarks
OpenAI reported that its custom inference chip could answer faster while doing more work for the same amount of power. A spicy name for a familiar engineering problem: speed without a bigger electricity bill.
What happened
- OpenAI tested Jalapeño on the public InferenceX benchmark with GPT-OSS 120B, DeepSeek R1 and Kimi K2.5.
- The company reported 1.5–1.9 times more work per watt at peak throughput and 1.7–3.6 times lower end-to-end latency than the comparison systems.
- In its August announcement, OpenAI said it planned to begin deploying the chip in its infrastructure by the end of 2026.
These are OpenAI’s benchmark results for the tested workloads. They are not a promise of the same gains for every model or customer.

Made with Supermeme
The surprised reaction captures the request every infrastructure team gets: make it faster and keep the bill down.

