Microsoft Takes on OpenAI and Google With 10-Cent AI Transcription

Blue waves of light in black background.
Sep 4, 2026
3 minute read
eWeek content and product recommendations are editorially independent. We may make money when you click on links to our partners. Learn More

Microsoft AI released MAI-Transcribe-2 on Thursday, tossing a direct challenge at competitors by combining frontier-level transcription speed with a disruptive price tag of 10 cents per audio hour.

The model marks a rapid push by Microsoft to deliver dedicated speech recognition models directly through Microsoft Foundry and MAI Playground, bypassing outside labs.

The limited-time $0.10-per-hour rate is about 72% lower than the $0.36-per-hour price listed for MAI-Transcribe-1 and 1.5. The lower rate could substantially reduce costs for businesses processing call-center recordings, meetings and compliance audio at scale.

Microsoft is also including several enterprise-oriented features with the model:

  • Speaker diarization and timestamps: Provides speaker-labeled segments alongside word-level timestamps for fast search and navigation.
  • Configurable formatting: Offers a "verbatim" style to keep filler words for audits and a "clean" style to generate readable meeting notes.
  • Broad audio support: Built-in automatic language detection and code-switching handle natural multilingual shifts such as Hinglish and Spanglish across 60 languages.
  • Domain adaptation: Keyword biasing allows teams to guide the model on industry-specific acronyms, jargon, and proper names.

Benchmark performance and independent testing

Microsoft claims MAI-Transcribe-2 leads the public FLEURS multilingual benchmark across 60 languages with an average Word Error Rate (WER) of 5.2%.

Artificial Analysis currently ranks the model second overall for accuracy, behind Alibaba’s streaming Fun-Realtime-ASR preview. In its non-streaming tests, MAI-Transcribe-2 recorded a 2.0% WER and a median speed factor of 410.7x real time. Microsoft says the model ran roughly 10 times faster than OpenAI’s GPT-Transcribe, 7.5 times faster than ElevenLabs’ Scribe v2 and 4.6 times faster than Google’s Gemini 3.5 Transcribe in the independent tests.

Microsoft expands its in-house AI portfolio

Unlike a general-purpose multimodal model, MAI-Transcribe-2 is designed specifically for speech recognition. That narrower focus may help Microsoft deliver faster batch processing at a lower cost for transcription-heavy workloads.

The release also expands Microsoft’s portfolio of internally developed AI models, although the company has not said whether MAI-Transcribe-2 will replace existing transcription technology across Teams, Nuance or its other products.

Advertisement

What eWeek found: How Microsoft compares

A closer look at benchmark data shows that while MAI-Transcribe-2 posts elite composite marks, performance varies across individual languages.

Model

Artificial Analysis WER

Median speed factor

Estimated price per hour

Microsoft MAI-Transcribe-2

2.0%

410.7x

$0.10 limited-time rate

ElevenLabs Scribe v2

2.2%

54.7x

About $0.22

Google Gemini 3.5 Transcribe

2.6%

89.9x

About $0.30

OpenAI GPT-Transcribe

3.3%

40.0x

About $0.27

Microsoft MAI-Transcribe-1.5

2.4%

190.3x

$0.36


Artificial Analysis normalizes pricing based on the cost of processing 1,000 minutes of audio. Its speed figures are rolling median results and may change as new tests are completed.

MAI-Transcribe-2 outclasses competitors in Chinese (4.5% WER) and French (2.8% WER). However, Google's Gemini 3.1 Pro beats it on languages like Armenian (6.1% vs. 12.6%) and Danish (5.3% vs. 5.4%), while Scribe v2 takes Afrikaans (10.6% vs. 13.1%).

Production trade-offs and enterprise risks

Engineering teams should note that Microsoft Learn classifies MAI-Transcribe as a public preview without a formal service-level agreement and does not recommend it for production workloads. The service currently accepts WAV, MP3 and FLAC files up to 300 MB or two hours in length.

The $0.10-per-hour offer is also scheduled to expire at the end of 2026, and Microsoft has not disclosed what the model will cost afterward. For enterprise teams, its price, speed and multilingual support make it worth testing against representative recordings, but the preview status and temporary pricing favor a limited evaluation over an immediate migration of mission-critical workloads.

Read more: Mistral’s Voxtral Transcribe 2 offers another approach to enterprise speech recognition, combining batch transcription with an open-weight real-time model.

Aminu Abdullahi

Aminu Abdullahi is a B2C and B2B technology and finance writer with more than six years of experience covering enterprise IT, cybersecurity, cloud computing, artificial intelligence, fintech, business software, and emerging technologies. His work has appeared in publications including TechRepublic, eWEEK, Channel Insider, Geekflare, Enterprise Networking Planet, eSecurity Planet, CIO Insight, and Webopedia. With a technical background in computer science, he specializes in translating complex technology topics into clear, accessible content for business leaders and decision-makers.

eWeek Logo

eWeek has the latest technology news and analysis, buying guides, and product reviews for IT professionals and technology buyers. The site's focus is on innovative solutions and covering in-depth technical content. eWeek stays on the cutting edge of technology news and IT trends through interviews and expert analysis. Gain insight from top innovators and thought leaders in the fields of IT, business, enterprise software, startups, and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.