OpenAI has rolled out a new mode called Ultrafast, designed to dramatically accelerate its most powerful model, GPT-5.6 Sol. The company says Ultrafast can work at 14 times the speed of standard processing, delivering up to 750 output tokens per second — tokens being the discrete pieces of text an LLM generates as it interacts with a human.

"Until now, getting real-time speed typically meant choosing a smaller or more specialized model," the company said in a blog post on Thursday. "Ultrafast points to progress in a new direction: more useful work per second."

OpenAI suggests the high-octane version of GPT-5.6 Sol can be deployed across a range of corporate workflows, notably incident response, customer service and support, financial market analysis, and e-commerce. The preview is being powered by OpenAI's partnership with chipmaker Cerebras, whose accelerated compute platform is widely used to serve large language models at very high throughput. Currently, the preview is only available to a small group of customers, although OpenAI says it will expand access as "capacity grows."

The launch continues a wider trend among frontier AI labs of offering speedier versions of their flagship models. OpenAI's competitor Anthropic, for instance, has a "fast mode" for its Claude models, though TechCrunch notes it does not deliver the same level of speed OpenAI is offering here.

For developers and enterprises, the appeal of Ultrafast is latency: near-instant responses that make AI assistants feel more natural in real-time applications like live customer support and financial trading workflows. The move also signals OpenAI's push to court enterprise users as competition over both price and performance intensifies — the same week Bloomberg reported that OpenAI's annualized revenue run rate has topped $40 billion, roughly doubling since the end of 2025.

Ultrafast launches first in the OpenAI API to a select group of customers, with expanded access to more businesses expected as infrastructure capacity ramps up.