Luna, a speed-focused model, now costs 20 cents per million input tokens and $1.20 per million output tokens, down from $1 and $6, the company said. OpenAI also cut its mid-tier GPT-5.6 Terra price by 20%, to $2 per million input tokens and $12 per million output tokens, according to the company. Pricing for the flagship GPT-5.6 Sol model remains unchanged.
The GPT-5.6 family consists of three tiers: Sol (highest capability), Terra (balanced), and Luna (fastest), according to the report. The reductions represent the clearest signal yet that OpenAI is being drawn into a price war driven by lower-cost competing models, the report said.
"Our strategy remains focused on advancing both capability and efficiency so each generation of intelligence can accomplish more work at a lower cost," OpenAI said in its announcement. The company said the lower prices follow efficiency improvements that reduced GPT-5.6 operating costs, with savings passed through to how usage is counted in Codex and ChatGPT Work subscriptions.
The new rates also apply on AWS, according to OpenAI. The company framed the cuts as the fruit of efficiency work rather than a margin sacrifice, the report said.
OpenAI CEO Sam Altman said at a recent event that costs had become "a huge issue," adding, "I think we'll have a lot of ways we can help people get more value for less spend," as quoted by the Wall Street Journal [1]. The company projected 2025 revenue of $12.7 billion – more than triple the estimated $3.7 billion in 2024 – though it anticipated it would not yet be profitable, according to a March 2025 report [2].
Companies that previously encouraged heavy AI usage are reviewing AI expenses that sometimes run into the billions, according to the report. Enterprises are seeking clearer returns before committing to the most expensive models and now have more alternatives than in the early ChatGPT era, the report stated.
Demand for reductions has come from senior executives. Palo Alto Networks CEO Nikesh Arora told CNBC that AI token prices need to fall as much as 90% before enterprise adoption can scale [3]. Palantir CEO Alex Karp said the business model of renting intelligence by the token is "effing insane" [4].
The Trends Journal reported that companies will build or buy smaller AI models customized to their own needs rather than purchasing ever larger systems [5]. JPMorgan's Mark Schilsky said most of his high-level investor discussions focus on "when will the party end?" [6]. A separate report said OpenAI missed revenue and user targets, and its chief financial officer expressed concern that $1.5 trillion in commitments cannot be paid [7].
Moonshot AI's Kimi K3, released in July, is a 2.8-trillion-parameter open-weight model that can run on company-owned infrastructure, according to the report. Blind developer testing placed K3 first in LMArena's front-end coding arena, ahead of Anthropic's Claude Fable 5, according to the report [8].
Artificial Analysis' cost-per-task index lists K3 at $0.94 per task and DeepSeek V4 Pro at $0.04, compared with $1.04 for OpenAI's Sol and $1.80 for Anthropic's Claude Opus 4.8, the report stated. Moonshot reported cache-hit rates above 90% in coding workloads, reducing K3's effective input cost to 30 cents per million, and has capped subscriptions and API access because of demand, according to a detailed analysis of Chinese models [9]. The company's daily revenue has grown roughly sixfold since launch as it seeks a $50 billion valuation ahead of a potential Hong Kong IPO, the report said.
Chinese companies such as DeepSeek and ByteDance have released AI models that rival U.S. giants like OpenAI and Google with significant cost and efficiency advantages, according to previous reports [10]. DeepSeek R1 was reported to rival OpenAI's models at a fraction of the cost while running on standard hardware [11].
Anthropic released Claude Opus 5, which it said approaches its top-tier Fable 5 model's performance at half the price, according to the report. Microsoft promoted MAI-Cyber-1-Flash, which AI chief Mustafa Suleyman said delivers "world-leading performance at 50% of the cost." Google launched three Gemini Flash models this month and said its top Flash model undercuts Kimi K3 on a per-task basis, the report stated.
The report said whether the price cuts slow the shift toward cheaper alternatives remains to be seen, but the era of unconstrained AI spending is giving way to a more pragmatic market. China's release of world-class AI models at low cost represents a significant challenge to the U.S. economy, according to an analysis by the Health Ranger Report [12]. The market for AI platforms, services, hardware and infrastructure is growing by 40% to 55% annually, with economic value projected to leap from $185 billion in 2023 to as much as $990 billion before 2028, according to the Trends Journal [13].