Release2026-08-10

OpenAI began a limited preview of Ultrafast, a new speed setting in its API that runs the GPT-5.6 Sol model far faster than the standard service, using hardware from Cerebras. Early testers include firms working on coding, commerce, financial research and voice support. Access is restricted to a selected group of customers for now, and OpenAI says it will widen it as capacity allows.

What changed

Getting real-time response speed generally meant switching to a smaller or more specialized model.

What it unlocks

Running OpenAI's most capable model inside conversations and incident work where a reply has to arrive within a second or two.

  • up to 14× faster than Standard processing
  • up to 750 output tokens per second

What you need to act on it

  • OpenAI API access
  • selection into the limited preview, or joining the notification list

Send this to someone who needs it

Shares the story and its sources. Nothing about you.

What does this mean for your job?

This is the story as everyone gets it. Once a week we send you the version written for your role — what changed, why it matters for the work you actually do, and one thing to try. Free while we tune it.