Google has introduced Gemini 3.5 Transcribe, a new AI model focused on enhancing speech-to-text capabilities. This model aims to produce more polished transcriptions by automatically editing out verbal fillers such as "ums" and "uhs," as well as self-corrections.
According to Google, Gemini 3.5 Transcribe offers significant performance improvements over its predecessor, Chirp 3. The company states the new model is approximately 70% faster in converting voice to final transcribed text. Additionally, the live-speech error rate has been reduced to 5.5 percent, an improvement from Chirp 3's 7.32 percent.
The AI model is also designed to better understand user intent, capable of removing hesitations and correcting spoken errors in real-time. It can also incorporate custom vocabulary for specialized jargon. The technology supports 85 languages and can process pre-recorded audio featuring up to three speakers.
While the model offers enhanced accuracy and speed, a potential drawback is the AI's alteration of the original spoken words. This may not be suitable for all situations where verbatim transcription is required.