Three weeks is not a long time to live with a price tag. That's how long GPT-5.6 lasted before OpenAI reworked the numbers on Thursday, dropping the cost of its cheapest tier by 80% and trimming its mid-range option by 20%. Two of the three models in the family got cheaper. The flagship didn't move at all.

The pressure behind the decision isn't mysterious. Enterprise buyers have started reading their invoices closely, and cheaper Chinese rivals have made "good enough at a fraction of the cost" a real option rather than a talking point.

What Changed in GPT-5.6 API Pricing

Luna, the fastest model in the lineup, took the deepest cut. It now runs 20 cents per million input tokens and $1.20 per million output tokens, down from $1 and $6. Terra, the mid-tier option, fell to $2 and $12 from $2.50 and $15. Sol, the flagship, holds at $5 input and $30 output.

 

Model

 

 

Input (per 1M tokens)

 

 

Output (per 1M tokens)

 

 

Change

 

 

GPT-5.6 Luna

 

 

$0.20 (was $1)

 

 

$1.20 (was $6)

 

 

−80%

 

 

GPT-5.6 Terra

 

 

$2 (was $2.50)

 

 

$12 (was $15)

 

 

−20%

 

 

GPT-5.6 Sol

 

 

$5

 

 

$30

 

 

Unchanged

 

The split matters. Leaving Sol untouched keeps the premium tier positioned as a premium tier, while the cheaper models absorb the competitive fight. If you're running high-volume, latency-sensitive workloads, Luna just became roughly five times less expensive on input than it was in the morning.

Why OpenAI Says It Can Charge Less

The company pinned the reductions on infrastructure work rather than margin sacrifice, laying out the details in an engineering post published a day before the announcement.

The Efficiency Gains Behind the Discount

Four changes did the heavy lifting: speculative decoding, fine-grained load balancing, cache optimization, and custom GPU kernels. Together, OpenAI said, they trimmed end-to-end serving costs by 20% and pushed token-generation efficiency up by more than 15%.

In its release, the company framed this as a continuation of its existing approach — push capability and efficiency at the same time, so every new generation does more work for less money. It's a tidy story, and the engineering numbers support part of it. But a 20% reduction in serving cost doesn't mechanically produce an 80% price cut. The gap between those two figures is where competitive strategy lives.

The Competitive Pressure Behind the Repricing

The repricing lands in a month that has been unusually crowded.

Chinese startup Moonshot AI put out its open-weight Kimi K3 model earlier in July, and it beat some leading American models on industry benchmarks. Anthropic answered with Claude Opus 5. Google shipped three new models aimed squarely at undercutting rivals on cost. And on Wednesday's earnings call, Microsoft CEO Satya Nadella made a point of highlighting cost-effective models — not the most capable ones, the most economical ones.

That last detail is worth sitting with. When the CEO of the company with the deepest AI distribution footprint starts talking about price-performance on an earnings call, it tells you where the conversation with customers has moved.

Where the New Prices Leave Anthropic

At 20 cents per million input tokens, Luna undercuts Anthropic's cheapest published model by a factor of roughly five on input, according to Reuters. Terra now slots in beneath the $3 per million input tokens Anthropic charges for its mid-tier Claude Sonnet offering.

So OpenAI hasn't just gone cheaper — it has gone cheaper at both the budget and middle rungs of the ladder, the two places where volume actually accumulates.

Enterprises Are Finally Watching the Meter

For a while, the standard corporate posture was to hand employees AI access and encourage them to use it. That era is closing. Companies are now working to rein in bills that grew faster than anyone modeled.

Amazon's engineering organization capped AI spending after running past its cost projections. OpenAI itself shipped hard spend limits for API customers on July 22 — a feature that only exists because customers asked for a way to stop the meter.

A few things follow from that shift:

  • Per-token price became a purchasing criterion, not a footnote. Buyers are comparing tiers the way they compare cloud instances.
  • Cheaper models get a serious look for production work. Workloads that would have defaulted to a flagship a year ago are being tested on smaller tiers.
  • Guardrails are now table stakes. Spend caps and budget alerts have moved from nice-to-have to expected.

What the Speed of the Cut Signals

Price reductions in this industry usually arrive months after a model ships, once usage patterns settle and serving costs come down through ordinary iteration. Three weeks is fast enough that Axios flagged the timing itself as the notable part.

Analysts have raised the obvious tension. Lower prices should drive more usage — that's the whole bet. But they may also squeeze the finances of both OpenAI and Anthropic at a moment when each is heading toward an anticipated public listing. Cheap tokens are an excellent way to win developers and an awkward line item to explain to prospective investors.

For anyone building on these APIs, the practical read is straightforward. Rerun your cost math, because the assumptions you made three weeks ago are stale. And expect the numbers to move again — at this cadence, today's pricing page is a snapshot, not a contract.