OpenAI cuts GPT-5.6 Luna API prices by 80% and Terra by 20%

AI

New API rates took effect on July 30, while ChatGPT Work and Codex subscribers receive lower credit usage without changes to subscription prices or quota budgets

Robotic hands hold an illuminated blue cube. The image represents OpenAI’s work to improve the efficiency of its GPT-5.6 models.

OpenAI has reduced GPT-5.6 Luna and Terra prices and introduced faster processing for GPT-5.6 Sol in the API

OpenAI has reduced API prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%, lowering the cost of its faster models for high-volume AI workloads.

The rates took effect on July 30, 2026. Terra now costs $2 per million input tokens and $12 per million output tokens, while Luna costs $0.20 per million input tokens and $1.20 per million output tokens.

The pricing change also applies to how Terra and Luna usage is counted against paid subscriptions in ChatGPT Work and Codex. Subscription prices and quota budgets remain unchanged, but use of the two models now consumes fewer credits.

Scott Rosecrans, Vice President of Strategic Pursuits at OpenAI, said on LinkedIn that the company was passing improvements in operating efficiency to customers. “As we get more efficient, we are passing the savings on to you!” he wrote.

Luna is positioned as OpenAI’s fastest and lowest-cost GPT-5.6 model. The company says it can use tools and complete multi-step workflows, making it suitable for high-volume work where speed and unit cost carry more weight than using the family’s most capable model.

OpenAI claims Luna delivers performance comparable to models that were considered frontier-class a year earlier at approximately six cents per dollar per task and nearly nine times the speed. It also says Luna outperforms Fable 5 on professional work measured by Agents’ Last Exam at an estimated cost per task almost 99% lower.

GPT-5.6 Sol gains faster API processing

OpenAI has also introduced Fast mode for GPT-5.6 Sol in the API, replacing its Priority Processing option.

The company says Fast mode can deliver speeds up to 2.5 times faster than Standard processing without changing the model’s intelligence. It costs twice the Standard processing rate, while Sol’s underlying API pricing remains unchanged.

Existing API requests tagged as priority will automatically use Fast mode, making the replacement backward compatible. The option also aligns with the /fast setting in Codex.

This gives customers a direct cost and speed choice. Luna and Terra become less expensive for routine or high-volume processing, while Sol users can pay a premium when response time is more important.

Sol contributed to OpenAI’s efficiency work

OpenAI attributes part of the pricing reduction to changes across its models, inference systems and the agentic software connecting models with tools and context.

Within a human-led process, GPT-5.6 Sol autonomously rewrote and optimized production kernels, designed and ran hundreds of experiments, and monitored training while intervening when problems arose, according to the company.

OpenAI says the kernel work helped cut the end-to-end cost of serving the model by 20%. Experiments involving its draft model increased token-generation efficiency by more than 15%.

Other changes covered routing, load balancing, prompt caching and context management. OpenAI says its agentic system limits unnecessary context growth, caps tool output at 10,000 tokens by default and preserves prompt prefixes so previous computation can be reused.

Terra and Luna remain available through ChatGPT Work, Codex and the OpenAI API. Free and Go users can access Terra in ChatGPT Work and Codex, while Plus, Pro, Business and Enterprise subscribers can select Terra or Luna. OpenAI said the pricing changes would also begin rolling out through AWS on July 30.

Previous
Previous

Google Cloud opens $75,000 Agentic Cinema hackathon for Gemini-powered media tools

Next
Next

Swiss AI Initiative adds image and audio capabilities to Apertus 1.5