Skip to content

GPT-5.6 Luna falls 80% as Terra gets a 20% API price cut

Original: GPT-5.6 Luna price drops 80% as OpenAI cuts Terra 20% View original →

Read in other languages: 한국어日本語
LLM Aug 1, 2026 By Insights AI (Twitter) 1 min read 1 views Source
GPT-5.6 Luna falls 80% as Terra gets a 20% API price cut

Lower token costs reach agent workflows

The frontier-model race is now as much about cost per useful call as raw capability. On July 30, 2026, OpenAI posted a concrete pricing change for its GPT-5.6 lineup. The substantive line in the tweet was: Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%, and offering a faster option for GPT-5.6 Sol in the API. The source tweet is available on X.

The largest number is the 80% reduction for Luna. In a related post, Sam Altman put the new Luna rate at $0.20 per million input tokens and $1.20 per million output tokens. Terra moves down by 20% to the $2 and $12 range, while Sol gets a Fast mode that OpenAI describes as up to 2.5x faster at 2x the price with the same intelligence. The company also said the lower Luna and Terra prices are reflected in how usage is counted inside Codex and ChatGPT Work.

OpenAI’s main account is the company’s canonical channel for product, API, ChatGPT, and Codex changes, so this tweet is more than a social post about a price sheet. Repeated model calls define the ceiling for code review agents, internal research assistants, customer-support copilots, and data-cleaning pipelines. A lower Luna tier could let teams keep more work inside the GPT-5.6 family instead of routing routine steps to smaller models or aggressive caching layers.

The next thing to watch is realized cost, not the headline discount alone. Context length, generated output, tool calls, retries, and review loops can all erase part of a token-price cut. Teams running Codex auto-review or ChatGPT Work should compare bills before and after the change, then test whether Sol Fast mode’s 2x price is justified by the promised speedup in their own latency-sensitive jobs.

Share: Long

Related Articles