Release2026-08-26

Google released Gemini 3.5 Transcribe, a speech-to-text model that cleans up filler words and self-corrections and formats the result as it goes. Developers can reach it through two interfaces, one for live streaming audio and one for recorded files with speaker labels and timestamps. It is in public preview, and consumer access is limited to the Gemini app on macOS in English, Rambler on Android in selected countries, and Chrome later.

What changed

Google's previous transcription model, Chirp 3, had higher error rates and slower response.

What it unlocks

Building live captioning, voice agents and meeting or call transcription with speaker labels and word-level timestamps, plus custom vocabulary for specialist jargon.

  • 4.0% word error rate, streaming
  • 2.6% word error rate, non-streaming
  • 70% faster time to final text
  • over 85 languages

What you need to act on it

  • Gemini API access via Google AI Studio or Gemini Enterprise Agent Platform
  • public preview
  • macOS for the Gemini app dictation features
  • select countries and languages for Rambler on Android

Sources

Send this to someone who needs it

Shares the story and its sources. Nothing about you.

What does this mean for your job?

This is the story as everyone gets it. Once a week we send you the version written for your role — what changed, why it matters for the work you actually do, and one thing to try. Free while we tune it.