News
Microsoft releases MAI-Transcribe-2, a speech recognition model cheaper and faster than rivals

Microsoft AI has released MAI-Transcribe-2, a speech recognition model supporting 60 languages that the company markets as the most accurate, fastest and cheapest on the market. The price dropped 72 percent from the previous version.
Contents
Microsoft AI announced the launch of MAI-Transcribe-2 on September 3, 2026, a new version of its in-house speech recognition model. The company says the model beats out competing systems from OpenAI, Google and ElevenLabs on both accuracy and processing speed, while costing significantly less than the previous generation.
This is already the model's third release in under six months. Microsoft launched MAI-Transcribe-1 in April 2026 with support for 25 languages at $0.36 per hour of recording. In June, language support grew to 43 with the MAI-Transcribe-1.5 release. The September launch of MAI-Transcribe-2 adds another 17 languages while cutting the price by nearly fourfold.
Benchmark results
Microsoft bases its performance claims on two independent rankings. In the public FLEURS benchmark, which tests transcription across multiple languages, MAI-Transcribe-2 ranked first among the 60 languages evaluated, with an average word error rate of 5.2 percent. On the Artificial Analysis leaderboard, which maps the tradeoff frontier between accuracy and latency, the model placed second by word error rate.
On speed, Microsoft compared its model directly against three major competitors when processing long recordings. MAI-Transcribe-2 is said to process material ten times faster than OpenAI's GPT-Transcribe, seven times faster than ElevenLabs' Scribe v2, and five times faster than Google's Gemini 3.5 Transcribe.
The math for large-scale transcription users
The price cut translates directly into savings for businesses processing large volumes of recordings. A company transcribing 100,000 hours of audio a year would pay $10,000 under the new rate, compared to $36,000 under MAI-Transcribe-1 pricing. That kind of scale is typical of customer service centers, newsrooms processing video recordings, and companies offering automatic captioning.
Use cases and operating modes
The model is available through Azure AI Foundry and, according to Microsoft, is meant to power video captions, meeting transcription, accessibility tools for people with disabilities, call center conversation analysis, content creation workflows, and voice agents. MAI-Transcribe-2 offers two output modes: verbatim, which preserves every filler word and repetition, and cleaned, which strips out stutters and slips of the tongue from the final text.
A keyword bias feature lets users pre-specify industry terms, names or proper nouns likely to appear frequently in a given recording, aimed at improving recognition of specialized medical, legal or technical vocabulary.
Position in the transcription market race
The speech recognition model market became an arena of intense price and technology competition among the biggest AI players in 2026. OpenAI, Google and ElevenLabs regularly update their own transcription models, and Microsoft frames this launch as part of a broader strategy to build out its own MAI models in-house rather than relying solely on technology from partners like OpenAI.
For Polish companies using cloud transcription services, especially those processing recordings in multiple languages, the lower price and wider language support represent a real opportunity to cut operating costs without sacrificing speech recognition quality.