OpenAI on July 30 announced major API price reductions for two models in its GPT-5.6 family, cutting the cost of GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%.
GPT-5.6 Luna, the company’s fastest and lowest-cost model, now costs $0.20 per million input tokens and $1.20 per million output tokens. Terra, positioned as the middle option for everyday workloads, drops to $2 per million input tokens and $12 per million output tokens.
OpenAI CEO Sam Altman announced the move on X, calling them “major price cuts today,” including the 80% reduction for Luna and the 20% reduction for Terra. The company said the reductions stem from improvements across its models, inference systems, and agent tools, enabling GPT-5.6 models to complete tasks more efficiently.
High-volume intelligence at low-cost margins
The pricing shift moves GPT-5.6 Luna directly into competition with lower-cost inference offerings from rival providers, including Google's Gemini 3.5 Flash-Lite and offerings from Chinese startups.
According to benchmarking cited by OpenAI, Luna matches the capabilities of frontier-class models from a year prior at roughly 6 cents on the dollar per task while executing nearly nine times faster. On the professional benchmark Agents' Last Exam, OpenAI claims Luna outperforms Anthropic's Claude Fable 5 “at an estimated cost per task nearly 99% lower.”
OpenAI attributed the price cuts to internal efficiency gains driven autonomously by GPT-5.6 Sol itself. Within a human-led engineering process, Sol modified production software kernels and ran token-generation experiments. These automated adjustments reduced end-to-end model serving costs by 20% and boosted token-generation efficiency by over 15%, according to the company.
The shift from model access to unit economics
As enterprise adoption shifts from experimental deployment to high-volume operational workflows, AI developers are facing heightened scrutiny regarding return on investment.
The steep drop in per-token costs reflects an evolving market strategy where access to baseline intelligence is rapidly commoditizing.
Rather than competing solely on raw performance metrics at the top end of the portfolio, AI vendors are increasingly competing on unit economics to retain enterprise clients running continuous, multi-step agentic systems.
For corporate technology buyers, these adjustments make complex agent workflows such as multi-stage code generation, document parsing, and automated routing economically viable at scale.
By aggressively lowering the floor on lightweight models like Luna while charging a premium for latency-focused processing on Sol, OpenAI is encouraging customers to orchestrate mixed-model architectures: assigning lower-cost nodes to handle routine steps while reserving high-tier reasoning for complex bottlenecks.
Check out our AI pricing guide comparing current costs for ChatGPT, Claude, Gemini, image generators, video tools, and voice AI.


