OpenAI is previewing a new way to run its most capable GPT-5.6 model at dramatically higher speeds. The company says its new Ultrafast service tier can run GPT-5.6 Sol up to 14 times faster than standard processing.

Ultrafast mode launches first through the OpenAI API and is powered by Cerebras. It can generate up to 750 output tokens per second, potentially bringing frontier-level performance to workflows where latency matters as much as model intelligence.

OpenAI introduced GPT-5.6 Sol in June alongside the balanced Terra and speed-focused Luna models. The full family became broadly available in July, including through ChatGPT, Codex, and the API.

The company sees Ultrafast supporting live or near-production tasks including voice, customer support, commerce, developer agents, financial research, and security response.

OpenAI says its own developers have used it to analyze logs and traces during incidents, as well as compress research cycles that previously ran overnight into multiple iterations during the workday.

Access remains limited to a select group of customers while OpenAI evaluates how the added speed changes real-world products and expands capacity.

Businesses can join the Ultrafast waitlist by sharing their workload, latency requirements, expected usage, and other details.

Do more with your Apple products