Google released Gemini 3.5 Transcribe, a speech-to-text model that cleans up filler words and self-corrections and formats the result as it goes. Developers can reach it through two interfaces, one for live streaming audio and one for recorded files with speaker labels and timestamps. It is in public preview, and consumer access is limited to the Gemini app on macOS in English, Rambler on Android in selected countries, and Chrome later.
What changed
Google's previous transcription model, Chirp 3, had higher error rates and slower response.
What it unlocks
Building live captioning, voice agents and meeting or call transcription with speaker labels and word-level timestamps, plus custom vocabulary for specialist jargon.
- 4.0% word error rate, streaming
- 2.6% word error rate, non-streaming
- 70% faster time to final text
- over 85 languages
What you need to act on it
- Gemini API access via Google AI Studio or Gemini Enterprise Agent Platform
- public preview
- macOS for the Gemini app dictation features
- select countries and languages for Rambler on Android
Sources