GPT-5.6 Sol Ultrafast Hits 750 Tokens/s, Up to 14× Faster
Original: GPT-5.6 Sol Ultrafast Hits 750 Tokens/s, Up to 14× Faster View original →
Frontier intelligence moves toward real time
GPT-5.6 Sol now has an inference path that OpenAI says runs up to 14 times faster than Standard processing. The new Ultrafast tier generates as many as 750 output tokens per second on Cerebras infrastructure. It is aimed at latency-sensitive work where developers previously had to trade model capability for a smaller, faster system. Access starts with a limited set of OpenAI API customers and will expand as capacity grows.
“Previewing Ultrafast mode: GPT-5.6 Sol at up to 14x the speed. Launching first in the OpenAI API to a select group of customers with expanded access to more businesses as capacity grows.” — OpenAI
OpenAI’s account is the company’s primary feed for model, API, research, and deployment updates. Its accompanying product page confirms both headline figures: up to 14× Standard speed and up to 750 output tokens per second. The company describes Ultrafast as a service tier rather than a separate model, preserving GPT-5.6 Sol while changing the infrastructure and processing route behind it.
Where lower latency could matter
OpenAI lists voice and customer support, commerce, coding and design, financial research, security response, and live experimentation among the early targets. In an outage, a model could inspect changing logs, traces, code revisions, and engineer reports while the incident is still unfolding. In voice support, it could traverse several systems without forcing a caller to wait through a long pause. Those are cases where end-to-end response time affects whether the system is useful at all.
The preview leaves important operational details unanswered. OpenAI has not published pricing, a general-availability date, regional coverage, or a service-level commitment for the 750-token rate. Output throughput also measures only one part of latency; time to first token, network delay, tool calls, and long-context processing can still dominate a complete workflow.
Watch for capacity expansion, API pricing, and independent measurements across long prompts, tool use, and structured output. The source post is on X, while OpenAI’s Ultrafast product page provides the deployment context.
Related Articles
OpenAI tied GPT-5.6 Sol’s new “The Last Ones” cyber-range result to Codex Security, a plugin meant to find, validate, and fix vulnerabilities in real repositories. The important comparison is controlled benchmark success versus code review work that security teams can actually run.
OpenAI’s Astra is the company’s first model treated as potentially Critical for cybersecurity. GPT-5.6 Sol stayed at High, but Astra triggered stricter controls and a pause on internal work that does not meet them.
A cyber-specific model completed 95.0% of advanced security requests, up from 1.5% for GPT-5.6 Sol. Access is restricted to approved Daybreak Red researchers, and the model has already uncovered two previously unknown V8 flaws.