Google has launched Gemini 3.5 Transcribe, a new speech-to-text model designed to turn natural speech into clean, formatted text while handling background noise, specialist terminology and corrections made while speaking.
The model is aimed at voice agents, meeting transcription, call analysis and everyday dictation. Developers can access it through the Gemini API, while Google is also integrating the technology into products including Gboard, the Gemini app and Antigravity.
Key facts
- Gemini 3.5 Transcribe supports more than 85 languages.
- It can automatically remove filler words and recognise self-corrections.
- Custom vocabulary can be supplied for specialist terminology and names.
- Recorded audio can include speaker attribution and word-level timestamps.
- Google reports word error rates of 4.0% for streaming and 2.6% for non-streaming transcription.
- Google says final transcription latency is 70% better than its previous Chirp 3 model.
Our take
Speech is becoming a much more practical way of interacting with generative AI. Better transcription means people can dictate instructions, capture meetings and work with AI assistants without first turning everything into carefully structured text.
For New Zealand organisations, the interesting applications include professional services, interviews, investigations, meetings and field work. Accuracy still needs to be tested against New Zealand accents, Māori words, names and specialist terminology before relying on automated transcripts for important records.
Sources

