Release2026-09-03

Microsoft released MAI-Transcribe-2, a speech-to-text model. It labels who is speaking and gives timing for every word. Transcription style can be set to verbatim or cleaned of filler words. Keyword biasing helps it catch names, jargon and abbreviations. It detects the spoken language automatically and handles mixed-language speech such as Hinglish. Microsoft says it covers 60 languages and holds up in noisy recordings. The launch price is a limited-time offer that runs until the end of the year. Microsoft cites the FLEURS benchmark and Artificial Analysis tests for its accuracy and speed claims. It ranks second on the Artificial Analysis word-error-rate leaderboard. The model can be tried through Microsoft Foundry, MAI Playground and OpenRouter.

What changed

Earlier MAI transcription models lacked speaker labelling, word-level timing and style controls.

What it unlocks

Transcribing long, noisy, multilingual audio with speaker labels and per-word timing from one model.

  • $0.10 per hour of audio
  • 5.2% average word-error rate
  • 60 languages
  • 10x faster than GPT-Transcribe

What you need to act on it

  • access via Microsoft Foundry, MAI Playground or OpenRouter

Send this to someone who needs it

Shares the story and its sources. Nothing about you.

What does this mean for your job?

This is the story as everyone gets it. Once a week we send you the version written for your role — what changed, why it matters for the work you actually do, and one thing to try. Free while we tune it.