Featured image source: Pexels (free to use)
Price wars in AI usually arrive dressed up as capability announcements. This one did not bother.
On September 3, Microsoft AI released MAI-Transcribe-2 and priced an hour of transcribed audio at 10 cents. Its own previous model, MAI-Transcribe-1, launched about five months earlier at 36 cents. That is a 72 percent cut against itself, in under half a year, on a product category that most people assumed had already been commoditized.
The claims attached to it are the aggressive part. Microsoft says the model is not only cheaper than what OpenAI, Google and ElevenLabs are selling, but more accurate and several times faster than all of them at the same time. Cheaper usually means worse at something. Microsoft is arguing it means worse at nothing.
The short version
- The price: $0.10 per audio hour, down from $0.36 for the previous model.
- The accuracy claim: first place on the FLEURS multilingual benchmark across 60 languages, at a 5.2 percent average word error rate.
- The speed claim: up to 10x faster than OpenAI’s GPT-Transcribe, 7x faster than ElevenLabs Scribe v2, 5x faster than Gemini 3.5 Transcribe.
- The catch: the 10 cent rate runs to the end of 2026, and Microsoft has not said what happens in January.
The number that actually moves things
Ten cents an hour sounds small in isolation. It is more useful to think in volume, because that is who this pricing is aimed at.
An hour of audio for a dime works out to about $1.67 per thousand minutes. A team processing a thousand hours of recordings a month, which is an ordinary amount for a mid-sized call center, a media archive or a compliance department, moves from a $360 monthly line item to a $100 one. Multiply that across an enterprise contract and the saving stops being a rounding error and starts being the reason someone schedules a migration meeting.
That is the part worth sitting with. Microsoft is not undercutting a competitor by a few percent to win a bake-off. It cut its own price by nearly three quarters, which is the behavior of a company that has decided the margin on speech recognition is not where the money is and would rather own the pipe.
The accuracy claim, and how to read it
Microsoft’s headline result is FLEURS, a multilingual speech benchmark. MAI-Transcribe-2 ranks first across 60 languages with an average word error rate of 5.2 percent, ahead of OpenAI’s GPT-Transcribe, Google’s Gemini 3.5 Transcribe, ElevenLabs Scribe v2 and Whisper V3-Large.
A 5.2 percent word error rate means roughly one word in twenty needs fixing. In practice that is the difference between a transcript you skim and correct, and a transcript you rewrite. For most professional uses, the useful threshold sits somewhere in that range, which is why small movements in this number matter more than they look.
The honest caveat is that these are Microsoft’s own published comparisons, run by the company selling the model. That is standard practice across the industry and it is also a reason to wait for independent replication before treating the ranking as settled. Benchmark leadership in speech recognition has historically been fragile, and real-world audio is messier than any evaluation set: overlapping speakers, bad microphones, background noise, and accents that are underrepresented in training data.
Where it sits against the field
| Model | Vendor | Microsoft’s speed comparison |
|---|---|---|
| MAI-Transcribe-2 | Microsoft AI | Baseline, $0.10 per audio hour |
| GPT-Transcribe | OpenAI | Up to 10x slower |
| Scribe v2 | ElevenLabs | 7x slower |
| Gemini 3.5 Transcribe | 5x slower | |
| Whisper V3-Large | Open weights | Beaten on FLEURS accuracy |
ElevenLabs is the company with the most to lose here. Its transcription business is a product line rather than a loss leader attached to a cloud platform, which is a harder position to defend when the competition is willing to price near cost and make the money somewhere else.
The calendar problem
Here is the detail that should shape how anyone actually uses this.
The 10 cent rate is a limited-time price that runs through the end of 2026. Microsoft has not published what the rate becomes on January 1, 2027. That is an unusual thing to leave open on an infrastructure product, because infrastructure pricing is exactly what companies build annual budgets and multi-year contracts around.
Before you migrate a pipeline to it
- Model the January price, not the September one. Run your numbers at $0.36 as well as $0.10 and see whether the move still makes sense.
- Keep the integration swappable. Speech-to-text is one of the easiest AI services to abstract behind an interface. Do that now rather than later.
- Test on your worst audio, not your best. Benchmark sets are clean. Your call recordings are not.
- Check language coverage against your actual traffic. An average across 60 languages says very little about the three you care about.
None of that is a reason to avoid the model. It is a reason to treat the current price as promotional, which is what Microsoft has said it is.
Why this keeps happening in AI right now
The wider pattern is that the price of a specific AI capability collapses roughly as soon as more than two credible vendors can deliver it. Speech recognition has now crossed that line. Frontier reasoning models have not, which is why OpenAI can charge $50 per million output tokens for its newest model while transcription races toward the floor.
The other half of the pattern is consolidation. The companies that can afford to sell a capability at cost are the ones that own the surrounding platform, which is the same logic that drove Nvidia’s $13 billion move on Hugging Face. Microsoft does not need speech recognition to be profitable. It needs speech recognition to run on Azure.
It is also worth noting who built it. This came out of Microsoft AI, the division under Mustafa Suleyman, rather than through the OpenAI partnership. Microsoft shipping a model that directly undercuts OpenAI on price, speed and accuracy is a fairly clear statement about how much of its own stack it intends to own, and it fits a run of in-house tooling from the company this year, alongside experiments like the collaborative AI development tools rivals have been shipping.
Bottom line
If you transcribe audio at any volume, this is the cheapest credible option on the market today and it is not close. The accuracy and speed claims come from Microsoft and deserve independent verification, but the price is a fact rather than a benchmark, and it is the part that changes budgets.
Just build with the assumption that 10 cents is an introductory number. The model will still be there in January. The price on it might not be.

