Microsoft released MAI-Transcribe-2, a speech-to-text model. It labels who is speaking and gives timing for every word. Transcription style can be set to verbatim or cleaned of filler words. Keyword biasing helps it catch names, jargon and abbreviations. It detects the spoken language automatically and handles mixed-language speech such as Hinglish. Microsoft says it covers 60 languages and holds up in noisy recordings. The launch price is a limited-time offer that runs until the end of the year. Microsoft cites the FLEURS benchmark and Artificial Analysis tests for its accuracy and speed claims. It ranks second on the Artificial Analysis word-error-rate leaderboard. The model can be tried through Microsoft Foundry, MAI Playground and OpenRouter.
What changed
Earlier MAI transcription models lacked speaker labelling, word-level timing and style controls.
What it unlocks
Transcribing long, noisy, multilingual audio with speaker labels and per-word timing from one model.
- $0.10 per hour of audio
- 5.2% average word-error rate
- 60 languages
- 10x faster than GPT-Transcribe
What you need to act on it
- access via Microsoft Foundry, MAI Playground or OpenRouter
Sources