Why it matters
  • Lead. Microsoft has unveiled seven proprietary AI models under its MAI family, marking its most ambitious push yet to reduce dependence on OpenAI and compete at the frontier of AI development with models it fully controls.
  • Fact. The flagship reasoning model, MAI-Thinking-1, matched Anthropic’s Claude Opus 4.6 on the SWE-Bench Pro coding benchmark and was preferred over Claude Sonnet 4.6 in independent blind evaluations — while a customised MAI model for enterprise use outperformed OpenAI’s GPT-5.5 on McKinsey benchmarks at roughly one-tenth the cost.
  • Stake. With roughly $18bn invested across OpenAI and Anthropic, Microsoft now faces a structural tension between its partnership obligations and its incentive to route traffic through cheaper, margin-expanding models it owns outright.

At its Build 2026 developer conference in San Francisco on 2 June, Microsoft unveiled seven in-house AI models under the MAI (Microsoft AI) brand — spanning reasoning, coding, image generation, transcription, and voice synthesis. The announcement, led by Microsoft AI chief executive Mustafa Suleiman, was framed explicitly as a bid for “long-term self-sufficiency” after years in which the company’s AI strategy has been almost entirely defined by its relationship with OpenAI. Microsoft chief executive Satya Nadella described the moment as a shift from “consuming a frontier model to fully participating at the frontier.”

The seven models and what they do

The centrepiece is MAI-Thinking-1, Microsoft’s first purpose-built reasoning model, trained from scratch without OpenAI data. It is a sparse Mixture-of-Experts architecture with 35 billion active parameters and approximately one trillion total parameters, paired with a 256,000-token context window. In independent blind evaluations conducted by Surge, it was preferred over Anthropic’s Claude Sonnet 4.6 and matched Claude Opus 4.6 on the SWE-Bench Pro software engineering benchmark. The model is available in private preview via Microsoft Foundry, with support for function calling and the widely-used Chat Completions API. The remaining six models cover specialist workloads: MAI-Code-1-Flash, a five-billion-parameter coding model already rolling out across GitHub Copilot plans and the Visual Studio Code model picker; MAI-Image-2.5 and its Flash variant, supporting text-to-image and image-to-image tasks, already embedded in PowerPoint with a gradual OneDrive release under way; MAI-Transcribe-1.5, which covers 43 languages and leads independent accuracy scores on the FLEURS benchmark; and MAI-Voice-2 and its Flash variant, which add natural speech generation across more than 15 languages with voice-adaptation capability.

The strategic and commercial calculus

The business logic behind the launch is stark. Microsoft has invested approximately $13bn in OpenAI and a further $5bn in Anthropic, reselling both companies’ models through Azure — an arrangement that carries significant external cost exposure. When enterprise clients at McKinsey benchmarked a customised MAI model against GPT-5.5, the Microsoft model delivered a higher win rate on quality while running at roughly one-tenth of the cost, according to Suleiman. That gap matters enormously at scale: every query routed through a proprietary MAI model rather than a licensed frontier model captures margin that would otherwise flow to a third party. Microsoft has also recently restructured its agreement with OpenAI, capping revenue-sharing payments and terminating exclusive marketing rights — moves that signal an accelerating strategic divergence even as the two companies remain formally allied.

For developers, the practical implications are immediate. MAI-Code-1-Flash is already available at no additional cost across all GitHub Copilot plans, and Microsoft claims it outperforms Anthropic’s Claude Haiku 4.5 by 16 percentage points on SWE-Bench Pro while using 60 per cent fewer tokens on complex tasks. MAI-Thinking-1 is accessible through third-party inference platforms including OpenRouter, Fireworks AI, and Baseten, in addition to Microsoft Foundry itself. The image and voice models are priced through Azure Speech and the Foundry Model Catalogue — MAI-Transcribe-1.5, for instance, starts at $0.36 per hour — making them directly competitive with equivalent offerings from Google and Amazon Web Services.

Suleiman characterised the broader ambition as building a “hill-climbing machine” — an internal capability that compounds over time rather than depending on the pace of external partners. Whether Microsoft can sustain frontier-level quality from its own research organisation, rather than merely offering efficient mid-tier models, remains to be tested. MAI-Thinking-1’s benchmark results are competitive but drawn from evaluations Microsoft commissioned and released itself; independent third-party validation at scale is still pending. What is not in doubt is the direction of travel: the world’s largest software company is investing heavily to ensure that the next wave of AI infrastructure runs, at least in part, on models it built itself.