OpenAI cut prices on two GPT-5.6 models by as much as 80%, intensifying the artificial-intelligence industry's price war as enterprise customers demand better economics from increasingly expensive deployments. The move could accelerate model adoption and cloud usage, but it also raises questions about whether efficiency gains can offset lower revenue per token.

GPT-5.6 Luna's input price fell 80% to $0.20 per million tokens from $1, while output pricing dropped to $1.20 from $6. OpenAI reduced GPT-5.6 Terra pricing by 20%, with input tokens now costing $2 per million and output tokens priced at $12. Pricing for the flagship GPT-5.6 Sol remained unchanged at $5 for input and $30 for output.

The cuts arrived just three weeks after OpenAI launched the GPT-5.6 family on July 9. OpenAI attributed the reductions to model improvements and greater efficiency across its software and computing infrastructure, suggesting inference costs are falling faster than expected. ChatGPT and Codex subscription pricing was unchanged.

OpenAI also introduced a faster Sol API option for developers willing to pay twice the standard rate for responses delivered at up to 2.5 times the speed. The combination gives customers a broader choice between low-cost inference and premium performance.

Investor Takeaway

OpenAI remains privately held, leaving Microsoft NASDAQ:MSFT, Amazon NASDAQ:AMZN and Nvidia NASDAQ:NVDA as important public-market proxies. Microsoft retains a revenue-sharing relationship and primary cloud role, while OpenAI models are also available through Amazon Bedrock. Nvidia supplies critical computing systems for OpenAI's expanding infrastructure.

Investors should watch whether lower API prices produce enough additional usage to expand total revenue and cloud consumption. Rising token volumes would benefit infrastructure providers, but aggressive discounting could pressure AI software margins and force rivals to follow. The central test is whether OpenAI's efficiency improvements continue outpacing price reductions as enterprise workloads scale.