OpenAI has launched Ultrafast, a new mode built for its GPT-5.6 Sol model. The mode aims to cut the time it takes the system to generate responses.
Ultrafast runs at 14 times the speed of standard processing. It can produce up to 750 output tokens per second. Tokens are the distinct pieces of text an LLM creates when it responds to a user.
The company put it this way in its announcement: “Until now, getting real-time speed typically meant choosing a smaller or more specialized model. Ultrafast points to progress in a new direction: more useful work per second.”
How Ultrafast Compares to Other Fast Modes
Other AI labs have released accelerated options. Anthropic’s Claude includes a fast mode. OpenAI says its version delivers higher speed than that alternative.
Corporate Uses for the Faster Model
OpenAI points to several business settings where the higher speed could fit. These include incident response, customer service and support, financial market analysis, and e-commerce.
Preview Access and the Cerebras Partnership
Ultrafast is out in preview right now. The mode draws on OpenAI’s work with chipmaker Cerebras. Access is limited to a small group of customers for the moment. The company plans to open it more widely as capacity grows.

