Release2026-09-03

Microsoft released MAI-Transcribe-2, a speech-to-text model that turns audio into written transcripts. It covers 60 languages and detects the spoken language automatically. The model labels who is speaking and timestamps each word. Both features were absent from the previous version. It also accepts a list of keywords so industry-specific terms are spelled correctly. Microsoft says it handles background noise, poor recordings and speakers switching between languages. The company reports the lowest word error rate among leading speech models on the FLEURS benchmark, ahead of Whisper, GPT-Transcribe, ScribeV2 and Gemini. Pricing is cut sharply against the earlier version, described as limited-time. The model is available through Microsoft Foundry and the MAI playground.

What changed

MAI-Transcribe-1.5 covered 43 languages, cost $0.36 per hour, and lacked speaker labelling and word timestamps.

What it unlocks

Transcribing meetings or calls with speaker labels, word-level timestamps and custom vocabulary from one model.

  • $0.10 per hour of audio, limited time
  • 3.4% average word error rate
  • 60 languages
  • 1hr audio transcribed in 10 sec

What you need to act on it

  • Azure AI Foundry or MAI Playground access

Send this to someone who needs it

Shares the story and its sources. Nothing about you.

What does this mean for your job?

This is the story as everyone gets it. Once a week we send you the version written for your role — what changed, why it matters for the work you actually do, and one thing to try. Free while we tune it.