OpenAI previews Ultrafast GPT-5.6 Sol at up to 14x Standard speed
The API service tier is powered by Cerebras, with OpenAI reporting speeds of up to 750 output tokens per second and initial access limited to a select group of customers
OpenAI says GPT-5.6 Sol Ultrafast can run at up to 14 times the speed of Standard processing
OpenAI has previewed a new Ultrafast service tier for GPT-5.6 Sol that it says can run the model at up to 14 times the speed of Standard processing, bringing lower-latency access to its frontier model first through the OpenAI API.
Powered by Cerebras, Ultrafast can generate up to 750 output tokens per second, according to OpenAI. The company is initially testing the service with a limited group of customers across coding, commerce, financial research, customer support and other interactive applications.
Rather than using a smaller or more specialized model when response speed is critical, developers can use GPT-5.6 Sol at substantially higher inference speeds.
For research teams, OpenAI says that could turn work previously handled as an overnight batch process into something closer to an interactive session, allowing teams to run an experiment, examine its results and iterate again during the working day.
The company is also testing Ultrafast internally for incident response, where developers are using it to read logs, analyze traces, synthesize conversations and help prepare or validate fixes while a problem is still developing.
Speed becomes part of the model decision
AI model selection has typically involved trade-offs between capability, cost and latency. Ultrafast is OpenAI's attempt to change the speed part of that calculation for GPT-5.6 Sol.
The company says the new tier is intended for workloads where waiting for a response changes how useful the system can be in practice.
Its examples include analyzing market signals while conditions are moving, answering complex customer support questions during a live conversation, checking inventory and resolving checkout issues while a customer is still shopping, and assisting engineers during a system outage.
OpenAI is also using the preview to study whether the change in latency alters how people work with the model, rather than simply allowing the same tasks to finish faster.
One early customer is Jane Street. John Crepezzi, AI Assistants at Jane Street, says:
“The increase in speed brought by Cerebras is impressive. It enables different ways of using the models, and makes it practical for developers to work in a more focused and productive way alongside them.”
Podium, Basis and Rogo are also among the companies testing GPT-5.6 Sol in Ultrafast mode.
OpenAI tests faster research and incident-response workflows
Inside OpenAI, developers are using Ultrafast in workflows where information changes quickly and several rounds of analysis may be needed.
During incident response, teams use the model to process logs, traces and conversations, identify possible next checks and help prepare or validate a fix. OpenAI stresses that engineers remain responsible for judgment and deployment decisions.
Research is another test case. The company says its teams are using Ultrafast to search knowledge sources, query data and gather, organize and summarize information across connected tools.
A research workflow that previously involved launching experiments overnight and examining the results the following morning could instead support multiple rounds of experimentation during the same working day.
That is potentially the more consequential change behind the headline speed figure. If the model can return complex outputs quickly enough for repeated interaction, latency becomes less of an interruption to a developer or researcher's workflow.
Cerebras powers the new service tier
Ultrafast extends OpenAI's existing partnership with Cerebras, which is providing the infrastructure for the low-latency inference service.
OpenAI says Cerebras is supporting GPT-5.6 Sol at speeds of up to 750 output tokens per second.
The company describes the service as a new speed class rather than a separate model. Customers therefore continue using GPT-5.6 Sol, but through a processing tier designed for substantially faster output.
For now, access remains narrow. GPT-5.6 Sol Ultrafast is available in limited preview to a select group of customers, with OpenAI planning to expand availability as capacity grows. Businesses can register to receive updates when access expands.