Cerebras' wafer-scale chips now run OpenAI's most capable model at up to 750 tokens per second, 14 times faster than standard inference.
Cerebras' wafer-scale chips now run OpenAI's most capable model at up to 750 tokens per second, 14 times faster than standard inference.

Cerebras is powering a new Ultrafast tier in OpenAI's API that runs GPT-5.6 Sol at up to 750 output tokens per second, 14 times faster than Standard processing with no loss in intelligence, opening a new region of the speed-intelligence frontier for frontier-model workloads.
"GPT-5.6 Sol on Ultrafast is proof that speed and intelligence are no longer mutually exclusive," Andrew Feldman, CEO and co-founder of Cerebras, said.
Based on output speeds reported by Artificial Analysis, Ultrafast runs 5x faster than Claude Opus 4.8 in Fast mode and 11x faster than Claude Fable 5. On Humanity's Last Exam, a 2,500-question benchmark spanning graduate-level chemistry, economics and literature, GPT-5.6 Sol Ultrafast answered the full set in just over 11 hours, versus more than three days of continuous compute for Claude Fable 5, reaching comparable accuracy nearly 7x faster. On GDP-Val, a benchmark of economically valuable knowledge work such as legal briefs, financial models and engineering reports, Ultrafast delivered a 5.6x end-to-end speedup with no quality degradation.
The speed comes from Cerebras' Wafer-Scale Engine architecture, which keeps 44 GB of SRAM on each wafer-sized chip, holding model weights on-chip rather than shuttling them between on-chip memory and off-chip storage as GPU-based inference must. That eliminates the memory-bandwidth bottleneck that constrains frontier-model inference speed on conventional hardware.
Faster token generation translates directly into productivity for economically valuable work, from programming to drafting legal documents. Rohan Varma, product lead at OpenAI, said Ultrafast is a persistent edge for organizations using frontier AI to respond quickly to incoming information, from root-causing production outages to detecting cyberattacks. Jeffrey Wang, an OpenAI researcher, said Ultrafast delivers real-time insights so users don't have to context-switch across multiple parallel sessions.
Cerebras' contrarian approach packs 44 GB of SRAM on each wafer-sized chip, so weights stay on-chip and tokens flow uninterrupted through model layers pipelined across wafers. The architecture scales with model size, paving the way for a continued speed advantage on future frontier models. Sachin Katti, vice president of compute strategy and GPT-infrastructure at OpenAI, said the company is starting with a small group of customers to learn where the speed creates meaningful value before expanding the service.
Cerebras, which trades on Nasdaq under the ticker CBRS, counts OpenAI among a limited number of significant customers alongside Group 42 Holding, Mohamed bin Zayed University of Artificial Intelligence and AWS. The Ultrafast partnership deepens Cerebras' reliance on OpenAI while giving the chipmaker a marquee proof point for its wafer-scale inference technology against Nvidia's GPU-based data center business. OpenAI is beginning with a select group of customers, with access expanding as capacity grows.
The deal also shows how frontier-model providers are rethinking inference economics. OpenAI's decision to route its most capable model through Cerebras hardware rather than Nvidia GPUs suggests the memory-bandwidth bottleneck on conventional accelerators is becoming a competitive liability as models scale. For Cerebras, the OpenAI win provides a reference deployment that could pull in additional cloud and enterprise customers evaluating wafer-scale inference. The company did not disclose the commercial terms of the Ultrafast arrangement or the pricing OpenAI will charge for the tier.
This article is for informational purposes only and does not constitute investment advice.