Microsoft's MAI models go production-wide, undercutting OpenAI on Copilot GPU costs by up to 89%
MAI-Image-2.5-Pro and MAI-Voice-2-Flash now run at scale across Bing, PowerPoint, OneDrive, Dynamics 365 and Dragon Copilot, giving Microsoft a margin lever on its own Copilot stack.
Microsoft has flipped MAI-Image-2.5-Pro and MAI-Voice-2-Flash from preview into production across Bing, PowerPoint, OneDrive, Dynamics 365, Excel, GitHub Copilot and Azure, and the numbers it’s publishing alongside the rollout describe a cost restructuring more than a model launch. Per Microsoft AI’s Superintelligence team, MAI-Image-2.5 cuts GPU costs by up to 84% versus OpenAI’s GPT-Image-2 inside PowerPoint, and MAI-Voice-2-Flash trims up to 89% off inference costs versus the OpenAI voice model powering Dynamics 365 Contact Center.
That’s the news. The frame is that Microsoft’s largest single vendor relationship is quietly becoming optional inside its own product surface.
Bing Image Creator has migrated fully to MAI-Image-2.5, making it the first entirely in-house Microsoft consumer image tool. In OneDrive, the same model is driving a 26% increase in save rates, roughly 25% lower P95 latency, and 2.5 times greater efficiency under medium-utilization production workloads. Dragon Copilot, which serves 170,000 medical providers and processed 28 million patient encounters last quarter, now runs on MAI-Transcribe-1.5 across 58 languages, with internal evaluations reporting a 50% relative reduction in both transcription and language-identification error rates across most of them.
On the coding side, MAI-Code-1-Flash lands at 5B parameters and posts 51% on SWE Bench Pro, with Microsoft projecting roughly a 10x improvement in output tokens per dollar versus GPT-5.5 for fine-tuned deployments. Empowering.Cloud’s August enterprise update also flags MAI-Cyber-1-Flash as Microsoft’s first purpose-built security model, extending the family sideways into the SOC.
The strategic read is straightforward. Every point of Copilot gross margin Microsoft can pull back from OpenAI’s API meter flows directly to Azure’s operating leverage, and the company is now signaling extensions of the swap into Copilot Chat, Outlook and deeper into PowerPoint. Both flagship models are in public preview through Microsoft Foundry and the MAI Playground.
There’s a recognizable pattern here. Platform companies rarely stay tenants on infrastructure they can rebuild. AWS spent the 2010s replacing Intel silicon with Graviton for exactly this reason, and the framing today, in-house parts benchmarked against the incumbent with dollar figures attached, is the same playbook applied to a partner Microsoft is contractually entangled with. The MAI numbers aren’t a research announcement. They’re a negotiating position rendered as telemetry.
Sources
- https://microsoft.ai/news/introducing-mai-image-2-5-pro-and-mai-voice-2-flash/
- https://venturebeat.com/infrastructure/microsoft-launches-new-in-house-ai-models-it-says-cut-costs-up-to-89-versus-openai
- https://microsoft.ai/news/microsoft-build-2026-mai-keynote-transcript/
- https://aiweekly.co/alerts/microsoft-swaps-openai-image-models-out-of-powerpoint-and-bing
- https://empowering.cloud/microsoft-365-ai-workplace-update-august-2026/