Skip to content

GPT-5.6 Sol Ultrafast Hits 750 Tokens/s, Up to 14× Faster

Original: GPT-5.6 Sol Ultrafast Hits 750 Tokens/s, Up to 14× Faster View original →

Read in other languages: 한국어日本語
LLM Aug 15, 2026 By Insights AI (Twitter) 2 min read 1 views Source

Frontier intelligence moves toward real time

GPT-5.6 Sol now has an inference path that OpenAI says runs up to 14 times faster than Standard processing. The new Ultrafast tier generates as many as 750 output tokens per second on Cerebras infrastructure. It is aimed at latency-sensitive work where developers previously had to trade model capability for a smaller, faster system. Access starts with a limited set of OpenAI API customers and will expand as capacity grows.

“Previewing Ultrafast mode: GPT-5.6 Sol at up to 14x the speed. Launching first in the OpenAI API to a select group of customers with expanded access to more businesses as capacity grows.” — OpenAI

OpenAI’s account is the company’s primary feed for model, API, research, and deployment updates. Its accompanying product page confirms both headline figures: up to 14× Standard speed and up to 750 output tokens per second. The company describes Ultrafast as a service tier rather than a separate model, preserving GPT-5.6 Sol while changing the infrastructure and processing route behind it.

Where lower latency could matter

OpenAI lists voice and customer support, commerce, coding and design, financial research, security response, and live experimentation among the early targets. In an outage, a model could inspect changing logs, traces, code revisions, and engineer reports while the incident is still unfolding. In voice support, it could traverse several systems without forcing a caller to wait through a long pause. Those are cases where end-to-end response time affects whether the system is useful at all.

The preview leaves important operational details unanswered. OpenAI has not published pricing, a general-availability date, regional coverage, or a service-level commitment for the 750-token rate. Output throughput also measures only one part of latency; time to first token, network delay, tool calls, and long-context processing can still dominate a complete workflow.

Watch for capacity expansion, API pricing, and independent measurements across long prompts, tool use, and structured output. The source post is on X, while OpenAI’s Ultrafast product page provides the deployment context.

Share: Long

Related Articles