Microsoft released MAI-Transcribe-2, a speech-to-text model that turns audio into written transcripts. It covers 60 languages and detects the spoken language automatically. The model labels who is speaking and timestamps each word. Both features were absent from the previous version. It also accepts a list of keywords so industry-specific terms are spelled correctly. Microsoft says it handles background noise, poor recordings and speakers switching between languages. The company reports the lowest word error rate among leading speech models on the FLEURS benchmark, ahead of Whisper, GPT-Transcribe, ScribeV2 and Gemini. Pricing is cut sharply against the earlier version, described as limited-time. The model is available through Microsoft Foundry and the MAI playground.
What changed
MAI-Transcribe-1.5 covered 43 languages, cost $0.36 per hour, and lacked speaker labelling and word timestamps.
What it unlocks
Transcribing meetings or calls with speaker labels, word-level timestamps and custom vocabulary from one model.
- $0.10 per hour of audio, limited time
- 3.4% average word error rate
- 60 languages
- 1hr audio transcribed in 10 sec
What you need to act on it
- Azure AI Foundry or MAI Playground access
Sources