Friday, July 31, 2026

News

OpenAI Cuts GPT-5.6 Luna Prices by 80 Percent Amid Pressure From China

MarketPatryk Raba
OpenAI Cuts GPT-5.6 Luna Prices by 80 Percent Amid Pressure From China
Fot. Steve Jennings / TechCrunch, Wikimedia Commons (CC BY 2.0)

OpenAI has cut prices for its cheapest model, GPT-5.6 Luna, by 80 percent and its mid-tier model, Terra, by 20 percent, responding to competition from Chinese AI labs and growing business complaints about AI bills.

Contents
  1. Three weeks after launch
  2. Pressure from China
  3. Bills outpacing budgets
  4. What this means for the market

OpenAI has announced a drastic price cut for access to two of its GPT-5.6 family language models. The decision came just three weeks after this generation of models launched and is part of an escalating price war in the AI market, driven by competition from China and growing dissatisfaction among businesses over the cost of AI deployments.

As of July 30, 2026, the cheapest and fastest model in the family, GPT-5.6 Luna, costs $0.20 per million input tokens and $1.20 per million output tokens. That's an 80 percent drop from the previous rate of $1 and $6. The mid-tier model, GPT-5.6 Terra, designed for everyday business tasks, got 20 percent cheaper, dropping from $2.50 and $15 to $2 and $12 per million tokens.

Three weeks after launch

GPT-5.6 became generally available on July 9, 2026, so the price correction came just three weeks later. That's an unusually short interval for a major model provider's pricing policy, and a sign of how quickly the competitive landscape is shifting in the industry. OpenAI didn't change the price of its most powerful model in the family, GPT-5.6 Sol, but did introduce a new Fast mode for it: processing up to 2.5 times faster than standard, though at double the cost, $10 per million input tokens and $60 per million output tokens. Fast mode replaced the previous Priority Processing option.

The company attributes the cuts to improvements in inference infrastructure and production software. Optimized GPU kernels cut the cost of handling queries by 20 percent, while improvements in token generation raised efficiency by more than 15 percent.

Pressure from China

Behind the decision is above all growing competition from Chinese labs, which offer models of comparable quality for a fraction of the price of their American rivals. Models like Moonshot AI's Kimi K3, with 2.8 trillion parameters, or Z.ai's GLM family are closing the quality gap while, according to some estimates, costing up to thirty times less than leading American models. Google, which has spent months promoting cheaper Gemini variants, including the latest Gemini 3.6 Flash, and Microsoft are also driving down prices.

Luna's new price moves the model from the mid-market segment into the budget tier, where smaller models from Google, Xiaomi, DeepSeek and MiniMax now compete.

Bills outpacing budgets

The second driver of the changes is growing dissatisfaction among business customers over the cost of AI deployments. Uber admitted that after rolling out Claude Code to five thousand engineers, it burned through its entire 2026 AI budget in just four months.

Earlier this year, this wasn't a topic that came up at all, people were completely happy with what they were spending. Suddenly it became a huge problem - Sam Altman, CEO of OpenAI

Sam Altman said as much outright at OpenAI's June Intelligence at Work event, describing the sharp shift in sentiment among enterprise customers over rising token bills.

What this means for the market

Cutting token prices can boost query volume, but it also squeezes model providers' margins at a moment when both OpenAI and Anthropic are preparing for stock market debuts. Both companies filed confidential IPO registration paperwork in June 2026. Analysts warn that a price war fought in full view of prospective investors could call into question the profitability of a business model built on selling access to tokens.

For Polish companies and developers using OpenAI's API, the cut means a tangible reduction in bills, especially where Luna is used for simpler tasks: text classification, summarization, or handling the first line of queries in AI agents. The higher cost of Fast mode for the Sol model, meanwhile, may prompt some teams to reconsider whether paying for extra speed is worth it, given that cheaper competing models are closing the quality gap.

The price war in the language model market is intensifying week by week, and the price per token is becoming as important a battleground for providers as answer quality itself. Further moves from OpenAI's competitors, including Google and Anthropic, likely won't take long to arrive.

Share: