Microsoft has rolled out two new in-house AI systems, MAI-Image-2.5-Pro and MAI-Voice-2-Flash, and the numbers attached to them say more than the models themselves do. Announced Thursday through the company's AI division, both are already running in production — and in some workloads they're cutting GPU costs by as much as 89% compared with OpenAI's models. That's not a benchmark score from a lab. That's a line item on an infrastructure bill.

The two models join a family of seven MAI systems that Microsoft first showed off at its Build developer conference in June. The company frames the whole effort as a move toward long-term self-sufficiency, which is a polite way of describing what's actually happening: Microsoft is systematically swapping third-party AI from OpenAI and Anthropic out of its own product lineup and replacing it with technology it owns.

Where the MAI Models Are Already Running

This isn't a preview or a limited pilot. MAI models are live across Bing Image Creator, PowerPoint, OneDrive, Dynamics 365, and Azure.

Bing Image Creator has gone fully in-house — every image it generates now comes from MAI-Image-2.5, with no outside model in the pipeline. In PowerPoint, swapping to the same image model brought GPU costs down by up to 84% against OpenAI's GPT-Image-2. And in Dynamics 365 Contact Center, where customers include T-Mobile and EasyJet, MAI-Voice-2-Flash delivered the headline figure: GPU costs down as much as 89%.

The Copilot side of the business is shifting too. Reporting from Bloomberg indicated that Microsoft has started routing tens of thousands of Copilot prompts per week in Excel and Outlook to MAI models rather than to OpenAI or Anthropic. Mustafa Suleyman, who runs Microsoft AI, was direct with Bloomberg about the reasoning — the company spends heavily with Anthropic, and the objective is to shrink that spend and eventually get rid of it altogether.

The Pricing Signal

MAI-Voice-2-Flash carries a listed price of $15 per million characters. That's 32% cheaper than the model it replaces, and twice as fast. Two improvements pointing the same direction: Microsoft wants voice workloads that used to be expensive enough to ration to become cheap enough to run everywhere.

A Model Family Built Across Every Major Category

The MAI lineup now covers reasoning, coding, image generation, voice, and transcription — deliberately broad, because narrow coverage would leave gaps that still have to be filled by someone else's model.

 

Model

 

 

Where it runs

 

 

MAI-Image-2.5 / 2.5-Pro

 

 

Bing Image Creator, PowerPoint

 

 

MAI-Voice-2-Flash

 

 

Dynamics 365 Contact Center

 

 

MAI-Code-1-Flash

 

 

GitHub Copilot, Visual Studio Code

 

 

MAI-Thinking-1

 

 

Microsoft Foundry

 

MAI-Thinking-1 is the company's first reasoning model and is available through Microsoft Foundry. MAI-Code-1-Flash has been folded into GitHub Copilot and Visual Studio Code, putting Microsoft's own model directly into the developer tools that made Copilot a household name among engineers.

One claim underpins the credibility of the whole family: Microsoft says the models were trained from scratch on commercially licensed data, with no distillation from third-party systems. It's a pointed statement, and an important one. A model quietly distilled from a competitor's system would make "self-sufficiency" a fiction — you'd still be dependent, just less visibly.

The Shift Is Real, but It's Incremental

It's worth keeping the scale honest. Bloomberg's reporting noted that MAI models still handle only a small percentage of Microsoft's total AI traffic, and that complex tasks continue to get routed to OpenAI's frontier models. The partnership with OpenAI remains intact, backed by roughly $13 billion in investment.

So the picture isn't a clean break. It's more like a triage system: high-volume, well-defined work that runs constantly and costs a fortune in aggregate — image generation in PowerPoint, voice in a contact center — moves in-house first. The hard, open-ended reasoning tasks stay with the frontier models, at least for now.

That sequencing makes sense. Those repetitive workloads are exactly where per-token fees compound into real money, and they're also the easiest place for a good-enough model to be genuinely good enough.

Why Cost Pressure Is Driving This

The economics behind the pivot aren't unique to Microsoft. As AI usage scales across enterprise products, inference costs stop being a rounding error and start shaping product decisions. Every prompt a user fires off in Excel has a price attached, and multiplied across a customer base that size, the arithmetic gets uncomfortable fast.

By running its own models on its own Azure infrastructure — and on its custom Maia silicon — Microsoft sidesteps per-token payments to outside providers entirely. It owns the model, the data center, and increasingly the chips underneath. Each layer it controls is a layer nobody else can charge it for.

VentureBeat made the sharper observation here: the production cost data may end up mattering more than the model launches did. Plenty of companies announce models. Very few publish what those models cost to run against a named competitor in live products. That disclosure hands other enterprises a template — and a benchmark to hold their own vendors against.

Microsoft's own framing of the goal is straightforward enough: its products, running on its models, built for the people using them. The strategic version is less tidy. The company is trying to keep a $13 billion partnership productive while methodically reducing how much it needs that partnership to function.