Cheaper AI at Scale: OpenAI Slashes GPT-5.6 Luna and Terra Prices
OpenAI has sharply reduced GPT-5.6 Luna and Terra prices as businesses seek more affordable ways to scale AI agents and routine workflows.
OpenAI has announced major price reductions for two of its smaller GPT-5.6 models as businesses look for more affordable ways to adopt artificial intelligence at scale.
The company has reduced the price of GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%. Pricing for GPT-5.6 Sol, the flagship model in the family, remains unchanged.
The new prices took effect on July 30, 2026, and apply to the OpenAI API. OpenAI is also reflecting the reductions in how Luna and Terra usage is counted within paid Codex and ChatGPT Work subscriptions.
The move could make advanced AI agents, high-volume automation and routine business workflows significantly less expensive to operate.
What Has Changed?
The GPT-5.6 family includes three primary models:
- GPT-5.6 Sol: The flagship model for complex reasoning, coding and professional work.
- GPT-5.6 Terra: A balanced model for everyday business and development tasks.
- GPT-5.6 Luna: The fastest and most affordable option for high-volume workloads.
OpenAI initially launched the models with different capability and pricing levels. The latest update makes the two smaller models considerably more affordable while preserving Sol as the premium option.
GPT-5.6 API Pricing Comparison
The following table shows standard short-context API pricing per one million tokens:
| Model | Previous Input Price | New Input Price | Previous Output Price | New Output Price | Reduction |
|---|---|---|---|---|---|
| GPT-5.6 Luna | $1.00 | $0.20 | $6.00 | $1.20 | 80% |
| GPT-5.6 Terra | $2.50 | $2.00 | $15.00 | $12.00 | 20% |
| GPT-5.6 Sol | $5.00 | $5.00 | $30.00 | $30.00 | No change |
An input token represents information sent to a model, while an output token represents the response generated by it.
Actual costs can vary depending on context length, prompt caching, processing mode, tools and other API features.
GPT-5.6 Luna Receives the Largest Price Cut
GPT-5.6 Luna has received the most dramatic reduction.
Its standard API input price has fallen from $1 to $0.20 per million tokens, while its output price has dropped from $6 to $1.20 per million tokens.
OpenAI positions Luna as its fastest and least expensive GPT-5.6 model. It is intended for businesses that need to process large numbers of requests without using the most powerful model for every task.
Potential applications include:
- Classifying customer-support requests
- Extracting information from documents
- Generating structured data
- Creating product descriptions
- Processing background automations
- Writing routine tests and code changes
- Summarizing emails and reports
- Translating business content
- Routing tasks between AI agents
- Moderating or categorizing user content
These tasks may be relatively simple individually but expensive when repeated millions of times. An 80% reduction in token pricing could therefore have a meaningful effect on the operating costs of AI-powered products.
Luna Is More Than a Basic Text Model
Lower pricing does not mean Luna is limited to simple chatbot responses.
GPT-5.6 Luna can use tools and complete multi-step workflows. This allows developers to use it inside AI agents that search for information, call functions, process results and take additional actions.
For example, a customer-support agent could use Luna to:
- Read a customer message.
- Determine the request category.
- Retrieve the relevant account information.
- Search a company knowledge base.
- Draft a suggested response.
- Escalate the issue when human review is required.
When this process is repeated across thousands of customer conversations, model pricing becomes a major part of the product's operating cost.
Luna's new price may make these agentic workflows practical for more startups and small businesses.
GPT-5.6 Terra Becomes More Affordable
GPT-5.6 Terra has received a smaller but still meaningful 20% reduction.
Its standard input price has fallen from $2.50 to $2 per million tokens, while output pricing has declined from $15 to $12 per million tokens.
Terra is positioned between Luna and Sol. It is intended for workloads that require stronger reasoning and reliability than the lowest-cost model but do not need the full capability of the flagship.
Suitable Terra workloads may include:
- Workspace question answering
- Business-document analysis
- Everyday coding assistance
- Research and report preparation
- Sales and marketing automation
- Multi-step administrative tasks
- Internal knowledge assistants
- Financial document extraction
- Drafting professional communications
- Reviewing structured business data
Terra may be particularly useful when Luna is not sufficiently reliable but using Sol for every request would be unnecessarily expensive.
Why Businesses Are Focused on AI Costs
Artificial intelligence experiments can appear inexpensive at a small scale. Costs become more visible when a successful feature reaches thousands or millions of users.
An AI application's total cost can include:
- Input and output tokens
- Repeated prompts and instructions
- Long conversation histories
- Tool calls and web searches
- Failed attempts and retries
- Human review
- Data storage and retrieval
- Monitoring and evaluation
- Computing infrastructure
- Additional agent steps
An agent may make several model requests before completing one user task. If each request contains a long prompt and generates a detailed response, the cost can grow quickly.
Businesses are therefore moving beyond asking which model is the most capable. They are also asking which model can complete a specific task successfully at the lowest total cost.
Cost per Token Is Not the Only Measurement
A cheaper model does not automatically produce the lowest-cost outcome.
If a low-cost model requires several retries, produces more errors or needs extensive human correction, a more capable model may ultimately be more economical.
Businesses should evaluate the full cost of completing a successful task, including:
- Number of model requests
- Tokens used per request
- Response time
- Tool calls
- Failure rate
- Human-review time
- Customer satisfaction
- Business value created
A model costing twice as much per token may still be cheaper if it completes the task correctly in one attempt instead of four.
The best approach is to run evaluations using real tasks from the intended application rather than selecting a model only by its listed token price.
A Multi-Model Strategy Can Reduce Costs
The GPT-5.6 family allows businesses to assign different models to different parts of a workflow.
A company might use:
- Sol to plan complex work or resolve uncertainty.
- Terra to handle everyday reasoning and professional tasks.
- Luna to complete repetitive, well-defined actions at scale.
For example, a software-development agent could ask Sol to analyze a complicated feature request and create an implementation plan. Luna could then make routine changes, generate tests and perform repetitive checks. Terra could review the work and prepare the final summary.
This approach avoids paying flagship prices for every step while preserving stronger intelligence where it has the greatest effect.
Developers can also create routing systems that automatically select a model based on the complexity, risk or value of a request.
GPT-5.6 Sol Pricing Remains Unchanged
OpenAI has not reduced the standard price of GPT-5.6 Sol.
Sol continues to cost:
- $5 per million input tokens
- $30 per million output tokens
The model remains OpenAI's recommended GPT-5.6 option for complex reasoning, advanced coding and demanding professional work.
Instead of lowering Sol's standard price, OpenAI has introduced a new Fast mode for API customers.
Fast mode can provide up to 2.5 times faster performance than standard processing without changing the model's intelligence. It costs twice the standard API price and replaces OpenAI's Priority Processing option.
Existing requests marked for priority processing will automatically use Fast mode, preserving backward compatibility.
Fast mode may be useful for high-value situations where response time is more important than price, such as interactive coding agents, urgent analysis or customer-facing professional tools.
What Changes for Codex and ChatGPT Work Users?
Terra and Luna remain available across ChatGPT Work, Codex and the OpenAI API.
OpenAI has not reduced subscription prices or increased the listed quota budgets. Instead, Luna and Terra now consume fewer credits because of their lower underlying costs.
Free and Go users can access Terra in ChatGPT Work and Codex. Plus, Pro, Business and Enterprise customers can choose between Terra and Luna, along with other available models.
This means paid users may be able to complete more work within the same subscription allowance by selecting Luna or Terra for appropriate tasks.
Users should still check the model picker and current account limits because availability and usage policies may vary by product, plan and region.
How Startups Could Benefit
The Luna price cut may be especially important for startups building AI-powered SaaS products.
Early-stage companies frequently need to balance product quality with limited infrastructure budgets. A lower-cost model can make it easier to test features, support early customers and increase usage without creating unsustainable expenses.
Potential startup opportunities include:
- Affordable customer-support agents
- AI assistants for small businesses
- Document-processing platforms
- Coding and testing tools
- Research and summarization products
- Email and sales assistants
- Automated data-entry systems
- Industry-specific knowledge agents
- Content-management workflows
- Background monitoring tools
Lower API prices can also help companies offer more generous free plans or usage limits. However, pricing should still account for infrastructure, support, storage and other operational expenses—not only model tokens.
Practical Ways to Lower AI Spending
Changing models is only one way to reduce AI costs.
Businesses can also:
- Use shorter and clearer system prompts.
- Avoid repeatedly sending unnecessary conversation history.
- Use prompt caching for repeated instructions and context.
- Limit output length when detailed responses are unnecessary.
- Send simple tasks to Luna and reserve Sol for difficult work.
- Use Terra for tasks requiring a balance of quality and cost.
- Store completed results instead of generating them repeatedly.
- Monitor token usage by feature, customer and workflow.
- Set spending limits and alerts.
- Evaluate quality before changing production models.
- Reduce unnecessary agent loops and tool calls.
- Use batch or flexible processing for non-urgent work.
The objective should be to reduce the cost of a successful outcome rather than simply reduce the price of each API request.
Greater Competition Across the AI Industry
OpenAI's price reductions arrive as AI companies compete on both model intelligence and affordability.
Developers now have access to proprietary and open-weight models from several providers. This gives businesses more options but also increases the pressure on model companies to improve price, speed and reliability.
As model capabilities become more similar for everyday tasks, pricing may play a larger role in purchasing decisions.
OpenAI's decision to cut Luna's price by 80% signals that lower-cost models are becoming central to the next stage of enterprise AI adoption.
Final Thoughts
OpenAI's new GPT-5.6 pricing makes advanced AI automation more accessible, particularly for businesses operating high-volume workflows.
Luna's 80% reduction creates a much cheaper option for background agents, document processing, classification and routine development work. Terra's 20% reduction strengthens its position as the balanced choice for everyday professional tasks.
Sol remains the flagship option for the most complex work, with Fast mode available when customers need greater speed and are willing to pay a premium.
The most effective AI systems will likely use a combination of these models rather than relying on one model for every request. By matching model capability to task complexity, businesses can control spending while maintaining the quality their users expect.