OpenAI now reaches more than a billion active users and over two million businesses, a scale disclosure that landed exactly one day after the company cut prices on two GPT-5.6 models by as much as 80 percent. The user number surfaced in a blog post from CFO Sarah Friar called "Building Abundant Intelligence," and the timing wasn't subtle. Hours later, Chinese rival DeepSeek opened the public beta of its V4 Flash model at a fraction of OpenAI's rates.

Put those three events side by side and you get a clear picture of where the AI market sits right now: enormous reach, collapsing prices, and no profits.

Inside the GPT-5.6 Price Cuts

The reductions hit the two lower tiers of the GPT-5.6 family and left the flagship alone.

GPT-5.6 Luna, the fastest and most lightweight model in the lineup, took the deepest cut — 80 percent. API pricing fell to $0.20 per million input tokens and $1.20 per million output tokens, down from $1 and $6. GPT-5.6 Terra, the mid-tier option, dropped 20 percent to $2 per million input tokens and $12 per million output tokens. Sol, the most capable model in the family, kept its existing pricing.

 

Model

 

 

Input (per 1M tokens)

 

 

Output (per 1M tokens)

 

 

Change

 

 

GPT-5.6 Luna

 

 

$0.20 (previously $1)

 

 

$1.20 (previously $6)

 

 

−80%

 

 

GPT-5.6 Terra

 

 

$2

 

 

$12

 

 

−20%

 

 

GPT-5.6 Sol

 

 

Unchanged

 

 

Unchanged

 

 

None

 

Sam Altman framed the move on X as an effort to deliver the best price/intelligence tradeoff at every level, which reads less like a discount announcement and more like a positioning statement aimed at every tier of the market at once.

The Efficiency Argument Behind the Discounts

OpenAI's own explanation points to infrastructure rather than market pressure. The company credits GPT-5.6 Sol with helping optimize the production software that serves its models — work that cut end-to-end serving costs by 20 percent and pushed token-generation efficiency up by more than 15 percent, largely through improvements to speculative decoding.

That's a real technical story. It's also a convenient one. A 20 percent reduction in serving cost doesn't obviously fund an 80 percent reduction in list price, and the gap between those two numbers is where the competitive explanation lives.

Why the Repricing Happened So Fast

The GPT-5.6 series had only been generally available for three weeks when the cuts landed. That's an unusually short window, and analysts have tied the timing to competitive dynamics rather than any planned product roadmap. Models don't normally get repriced before most enterprise buyers have finished evaluating them.

DeepSeek V4 Flash Undercuts Even the New Floor

DeepSeek's V4 Flash entered public beta at roughly $0.14 per million input tokens and $0.28 per million output tokens. That undercuts Luna's freshly reduced pricing — and the output-token gap is the one that matters most, since output tokens dominate the bill for generation-heavy workloads.

Investor Michael Burry read the sequence as cause and effect, posting on X that OpenAI's cuts were preparation for the DeepSeek launch.

The strategic logic favors DeepSeek in a specific slice of the market. V4 Flash is positioned to absorb enterprise workloads where cost outweighs brand loyalty — high-volume classification, data extraction, and coding tasks where raw speed is the deciding factor. Those are exactly the jobs that don't need a frontier model, and exactly the jobs that generate the most tokens.

Chinese open-weight models have been gaining ground more broadly. They currently hold all five top positions in global usage rankings measured by routed inference volume. That metric has real limits — OpenAI and Anthropic still capture the majority of industry revenue through premium enterprise pricing, and first-party consumer traffic through ChatGPT and Gemini dwarfs router traffic entirely. But routed volume is a useful early signal of what developers pick when brand isn't part of the decision.

Anthropic Was the Other Trigger

The pricing gap with Anthropic flipped. Claude Sonnet 4.6, Anthropic's mid-tier model, sits at $3 per million input tokens and $15 per million output tokens — now meaningfully above Terra's revised rates.

The Wall Street Journal reported that OpenAI leadership spent weeks debating the reductions, driven by concern that Anthropic would cut Claude pricing first. That detail reframes the whole announcement. This wasn't a roadmap item that happened to ship; it was a defensive move deliberated internally for weeks and then triggered by the calendar of a competitor's launch.

Adding pressure: both OpenAI and Anthropic filed confidentially for public listings in June. Every pricing decision from here gets read as a signal about margin durability.

Enterprise Buyers Have Started Counting Tokens

The demand side changed too, and that shift is easy to miss underneath the headline numbers.

After a stretch where companies encouraged employees to use AI tools freely, many are now actively pulling spending back in. Amazon's engineering organization capped AI spending following cost overruns. OpenAI itself shipped hard spend limits for API customers on July 22 — a feature you only build when a meaningful share of your customer base is asking for it.

Arun Chandrasekaran, a vice president analyst at Gartner, told Business Insider that buyers are looking at a pivotal moment to extract more value from their AI partnerships. He also framed the price competition as an early test of whether frontier lab business models actually hold up, with public listings on the horizon.

Usage Is Deepening, Not Just Widening

The engagement data OpenAI shared is arguably more interesting than the billion-user headline, because it describes behavior rather than reach.

  • Users send roughly 50 percent more messages per day six months after signing up than they did at the start.
  • They use ChatGPT for about twice as many distinct types of tasks over that same period.
  • Agentic work through Codex now accounts for 99.8 percent of weekly output tokens across the company.

That last figure is the one worth sitting with. If nearly all output tokens come from agentic workloads, then consumer chat is essentially a rounding error in compute terms, and the economics of the business are being set by long-running automated tasks — the kind that consume enormous volumes of output tokens and are extremely sensitive to per-token pricing. It explains why an 80 percent cut on the cheapest, fastest model is a bigger strategic lever than it looks.

The Profitability Gap Underneath All of It

None of this scale has translated into a working income statement. OpenAI reported $13.07 billion in revenue for 2025 against a net loss of $38.5 billion, and analysts expect losses to continue through at least 2030.

That's roughly three dollars lost for every dollar earned — before the price cuts. Cutting a model's price by 80 percent works only if volume expands enough to compensate, or if serving costs keep falling faster than prices do. The efficiency gains OpenAI described suggest the second lever is real. Whether it moves fast enough to outrun both DeepSeek's pricing and $38.5 billion in annual losses is the open question, and it's the one that will define the next stretch of this market.