GPT-5.6 Luna gets 80% cheaper as Sol adds 2.5x Fast mode
Original: Advancing the price-performance frontier with GPT-5.6 View original →
The GPT-5.6 story is now about how much useful work buyers can run per dollar, not only where the models land on benchmarks. In a July 30, 2026 update, OpenAI cut GPT-5.6 Luna pricing by 80%, reduced GPT-5.6 Terra pricing by 20%, and added a Fast mode for GPT-5.6 Sol in the API. For enterprises trying to scale agent workloads, that is a direct change to the unit economics of automation.
Luna is the fastest and lowest-cost member of the GPT-5.6 family. OpenAI positions it for high-volume work where tool use and multi-step workflows matter, but where the frontier model is not always necessary. Terra sits in the middle of the family, aimed at everyday work where quality, speed, and reliability need to be balanced. OpenAI says the lower prices also affect how usage is counted against paid subscriptions in Codex and ChatGPT Work.
Sol’s update targets latency rather than base price. Fast mode replaces Priority Processing and offers up to 2.5x faster speeds than Standard processing with no change in intelligence, at twice the Standard price. Existing API requests tagged as priority will automatically use Fast mode, which lowers migration friction for teams already paying for priority handling. The obvious target is work where latency itself determines user experience: coding agents, customer operations, incident response, and other interactive workflows.
The move follows OpenAI’s engineering post on GPT-5.6 efficiency. There, the company said GPT-5.6 Sol helped optimize production kernels, improve draft-model experiments, reduce end-to-end serving costs by 20%, and raise token-generation efficiency by more than 15%. This pricing update turns those internal gains into customer-facing levers. The next pressure point is whether rival labs answer with lower prices, more durable long-running agents, or both.
Related Articles
GPT-5.6 Sol moved from 13.3% to 38.3% on ARC-AGI-3 when OpenAI retained reasoning and used compaction in the harness. The result makes benchmark setup, not just model weights, part of the frontier-agent story.
OpenAI’s newest model family is shipping first to a small trusted group after US government review. The post matters because Sol, Terra, and Luna combine new pricing tiers with a policy-limited rollout, including Terra at 2x lower cost than GPT-5.5.
Starting July 2, organizations that have not run inference on a fine-tuned model in the past 60 days can no longer create new fine-tuning jobs. Active existing customers lose new job creation on January 6, 2027, while inference on existing fine-tuned models continues until the base model is deprecated.