Microsoft Replaces OpenAI and Anthropic Models with Self-Developed MAI in Copilot to Cut Costs
Updated: Jul 20
Microsoft is shifting Copilot workloads away from OpenAI and Anthropic models. It now routes more requests through its internal MAI models instead.
The move targets spending reduction. MAI models already run thousands of tasks each week inside Excel and Outlook.
Initial tests show the approach works for simple queries. More demanding tasks still route back to external providers for now.
MAI Models Enter Production Use
MAI models handle routine requests across productivity tools. They process data drawn from commercial sources under clean licensing terms.
Weekly volume sits at several thousand inferences inside selected products. That share remains small relative to total Copilot traffic.
The company released its first reasoning model, MAI-Thinking 1, during the Build conference. The model targets coding tasks and claims parity with certain frontier systems on internal metrics.
Independent checks place its performance closer to DeepSeek V3.2 on public leaderboards.
Cost Pressure Drives the Switch
Microsoft has spent heavily on external model access over the past two years. Executives have stated the goal is to cut and eventually eliminate payments to Anthropic.
The same message covers OpenAI volumes over time. Internal models carry lower per-query costs once scaled.
Pricing signals point to a usage-based future for Copilot. MAI would serve as the default layer. Third-party models would sit behind paid add-ons.
This structure lets Microsoft control variable costs while still offering premium options.
Performance Claims Face Scrutiny
MAI-Thinking 1 was positioned as competitive with Sonnet 4.6 and Opus 4.6 in coding. Public benchmarks show larger gaps than the announcement suggested.
The gap narrows only on narrow tasks. Broader reasoning and long-context work still favor the external models Microsoft aims to replace.
Users in early internal tests report faster responses on simple requests. They also note occasional drop-offs in accuracy when prompts grow complex.
The company has not released full side-by-side evaluations yet.
Data Sources Raise Separate Questions
Microsoft states MAI training uses commercial-grade licensed data. Public reports trace parts of that data back to Common Crawl crawls.
The distinction matters for compliance and output quality. Many downstream customers want explicit guarantees on training provenance.
Current documentation leaves room for interpretation on how much web-scraped material remains inside the final weights.
What to Watch Next
Three signals will show whether internal models can sustain the shift. First, quarterly earnings will reveal changes in external API spend. Second, public benchmark updates on MAI-Thinking 1 will appear after more users gain access. Third, any expansion of MAI coverage into additional Copilot surfaces will indicate scaling readiness.
Each outcome will either tighten or loosen the current cost-performance trade-off.



