OpenAI began a limited preview of Ultrafast, a new speed setting in its API that runs the GPT-5.6 Sol model far faster than the standard service, using hardware from Cerebras. Early testers include firms working on coding, commerce, financial research and voice support. Access is restricted to a selected group of customers for now, and OpenAI says it will widen it as capacity allows.
What changed
Getting real-time response speed generally meant switching to a smaller or more specialized model.
What it unlocks
Running OpenAI's most capable model inside conversations and incident work where a reply has to arrive within a second or two.
- up to 14× faster than Standard processing
- up to 750 output tokens per second
What you need to act on it
- OpenAI API access
- selection into the limited preview, or joining the notification list
- openai.com2026-08-10