Release2026-09-03

Microsoft AI released MAI-Transcribe-2, a speech-to-text model, on Thursday. The launch price is 10 cents per hour of audio, called an early-bird rate. Microsoft has not named an end date or a standard price. The model covers 60 languages and handles noisy, overlapping real-world audio. It labels who is speaking, timestamps each word and accepts custom word lists. A verbatim mode keeps filler words for legal and compliance use. It also follows conversations that switch language mid-sentence. Microsoft claims first place on the FLEURS multilingual benchmark and second on Artificial Analysis. It says the model runs five to ten times faster than rivals from OpenAI, Google and ElevenLabs. The announcement says nothing about real-time transcription, speaker-labelling accuracy or data retention.

What changed

Microsoft's previous speech model covered 43 languages at $0.36 per hour.

What it unlocks

Bulk transcription with speaker labels, word timestamps and mixed-language handling at a dime an hour.

  • $0.10 per hour of audio, launch price
  • down from $0.36 per hour
  • 60 languages, up from 43
  • 5.2% average word error rate

What you need to act on it

  • access via Microsoft Foundry or MAI Playground

Send this to someone who needs it

Shares the story and its sources. Nothing about you.

What does this mean for your job?

This is the story as everyone gets it. Once a week we send you the version written for your role — what changed, why it matters for the work you actually do, and one thing to try. Free while we tune it.