Microsoft AI released MAI-Transcribe-2 on Thursday, tossing a direct challenge at competitors by combining frontier-level transcription speed with a disruptive price tag of 10 cents per audio hour.
The model marks a rapid push by Microsoft to deliver dedicated speech recognition models directly through Microsoft Foundry and MAI Playground, bypassing outside labs.
The limited-time $0.10-per-hour rate is about 72% lower than the $0.36-per-hour price listed for MAI-Transcribe-1 and 1.5. The lower rate could substantially reduce costs for businesses processing call-center recordings, meetings and compliance audio at scale.
Microsoft is also including several enterprise-oriented features with the model:
- Speaker diarization and timestamps: Provides speaker-labeled segments alongside word-level timestamps for fast search and navigation.
- Configurable formatting: Offers a "verbatim" style to keep filler words for audits and a "clean" style to generate readable meeting notes.
- Broad audio support: Built-in automatic language detection and code-switching handle natural multilingual shifts such as Hinglish and Spanglish across 60 languages.
- Domain adaptation: Keyword biasing allows teams to guide the model on industry-specific acronyms, jargon, and proper names.
Benchmark performance and independent testing
Microsoft claims MAI-Transcribe-2 leads the public FLEURS multilingual benchmark across 60 languages with an average Word Error Rate (WER) of 5.2%.
Artificial Analysis currently ranks the model second overall for accuracy, behind Alibaba’s streaming Fun-Realtime-ASR preview. In its non-streaming tests, MAI-Transcribe-2 recorded a 2.0% WER and a median speed factor of 410.7x real time. Microsoft says the model ran roughly 10 times faster than OpenAI’s GPT-Transcribe, 7.5 times faster than ElevenLabs’ Scribe v2 and 4.6 times faster than Google’s Gemini 3.5 Transcribe in the independent tests.
Microsoft expands its in-house AI portfolio
Unlike a general-purpose multimodal model, MAI-Transcribe-2 is designed specifically for speech recognition. That narrower focus may help Microsoft deliver faster batch processing at a lower cost for transcription-heavy workloads.
The release also expands Microsoft’s portfolio of internally developed AI models, although the company has not said whether MAI-Transcribe-2 will replace existing transcription technology across Teams, Nuance or its other products.
What eWeek found: How Microsoft compares
A closer look at benchmark data shows that while MAI-Transcribe-2 posts elite composite marks, performance varies across individual languages.
Model | Artificial Analysis WER | Median speed factor | Estimated price per hour |
Microsoft MAI-Transcribe-2 | 2.0% | 410.7x | $0.10 limited-time rate |
ElevenLabs Scribe v2 | 2.2% | 54.7x | About $0.22 |
Google Gemini 3.5 Transcribe | 2.6% | 89.9x | About $0.30 |
OpenAI GPT-Transcribe | 3.3% | 40.0x | About $0.27 |
Microsoft MAI-Transcribe-1.5 | 2.4% | 190.3x | $0.36 |
Artificial Analysis normalizes pricing based on the cost of processing 1,000 minutes of audio. Its speed figures are rolling median results and may change as new tests are completed.
MAI-Transcribe-2 outclasses competitors in Chinese (4.5% WER) and French (2.8% WER). However, Google's Gemini 3.1 Pro beats it on languages like Armenian (6.1% vs. 12.6%) and Danish (5.3% vs. 5.4%), while Scribe v2 takes Afrikaans (10.6% vs. 13.1%).
Production trade-offs and enterprise risks
Engineering teams should note that Microsoft Learn classifies MAI-Transcribe as a public preview without a formal service-level agreement and does not recommend it for production workloads. The service currently accepts WAV, MP3 and FLAC files up to 300 MB or two hours in length.
The $0.10-per-hour offer is also scheduled to expire at the end of 2026, and Microsoft has not disclosed what the model will cost afterward. For enterprise teams, its price, speed and multilingual support make it worth testing against representative recordings, but the preview status and temporary pricing favor a limited evaluation over an immediate migration of mission-critical workloads.
Read more: Mistral’s Voxtral Transcribe 2 offers another approach to enterprise speech recognition, combining batch transcription with an open-weight real-time model.


