All NewsEducationTV
Equities & FundsCrypto & Digital AssetsAI & TechnologyBusiness & CorporateUS Politics & PolicyGeopolitics & Global RiskMacro, Rates & FXCommodities & EnergyEuropean Politics & MarketsAsia-PacificReal Estate & Property
All NewsHome
← Back to AI & Technology

Google Unveils Gemini 3.5 Transcribe for AI-Powered Speech-to-Text

Created at 26 Aug · 7:26 PM1 source↑ Market-relevant
IN SHORT

Google has announced Gemini 3.5 Transcribe, an AI model designed to improve voice input by editing out hesitations and corrections for more polished text output. The model is reportedly 70% faster and more accurate than its predecessor, handling 85 languages and up to three speakers in pre-recorded audio.

Key Numbers

70 percentspeed improvement over Chirp 3
5.5 percentlive-speech error rate
7.32 percentChirp 3 error rate
85languages supported
3maximum speakers in pre-recorded audio

Who's Involved

Google
Announced Gemini 3.5 Transcribe AI model
Chirp 3
Previous voice-to-text engine
Google Unveils Gemini 3.5 Transcribe for AI-Powered Speech-to-Text

↳ Why This Matters

This advancement in AI-powered speech-to-text technology promises to make voice input more efficient and user-friendly across Google's ecosystem, potentially improving productivity for users who rely on voice for communication and content creation.

Key facts

  • Google has launched Gemini 3.5 Transcribe, an AI model for speech-to-text.
  • The model is designed to remove verbal hesitations like "ums" and "uhs" and correct self-corrections.
  • It boasts a 70% speed improvement over its predecessor, Chirp 3.
  • The live-speech error rate has been reduced to 5.5 percent.
  • Gemini 3.5 Transcribe supports 85 languages and can handle up to three speakers in pre-recorded audio.
  • Google has introduced Gemini 3.5 Transcribe, a new AI model focused on enhancing speech-to-text capabilities. This model aims to produce more polished transcriptions by automatically editing out verbal fillers such as "ums" and "uhs," as well as self-corrections.

    According to Google, Gemini 3.5 Transcribe offers significant performance improvements over its predecessor, Chirp 3. The company states the new model is approximately 70% faster in converting voice to final transcribed text. Additionally, the live-speech error rate has been reduced to 5.5 percent, an improvement from Chirp 3's 7.32 percent.

    The AI model is also designed to better understand user intent, capable of removing hesitations and correcting spoken errors in real-time. It can also incorporate custom vocabulary for specialized jargon. The technology supports 85 languages and can process pre-recorded audio featuring up to three speakers.

    While the model offers enhanced accuracy and speed, a potential drawback is the AI's alteration of the original spoken words. This may not be suitable for all situations where verbatim transcription is required.

    Frequently asked questions

    Gemini 3.5 Transcribe is a new AI model from Google designed to convert speech to text, with the ability to edit out verbal fillers and corrections for a cleaner output.

    It is reportedly 70% faster and has a lower live-speech error rate (5.5% compared to Chirp 3's 7.32%).

    Gemini 3.5 Transcribe supports 85 languages.

    Yes, it can process pre-recorded audio with up to three speakers.

    What Happens Next

    01Gemini 3.5 Transcribe will be integrated throughout the Google ecosystem.

    How It Developed

    Google announced Gemini 3.5 Transcribe, an AI model for speech-to-text.
    The model aims to streamline voice input by editing out hesitations and corrections.
    Gemini 3.5 Transcribe is reportedly 70% faster than the previous engine, Chirp 3.
    The live-speech error rate has dropped to 5.5 percent from 7.32 percent.
    The AI model can remove "ums" and "uhs" and edit corrections on the fly.
    It supports 85 languages and can process audio with up to three speakers.
    The model already powers the Gboard "Rambler" feature on Pixel phones.

    Sources

    T1
    Google announces Gemini 3.5 Transcribe for AI-powered speech-to-textvar abtest_2169206 = new ABTest(2169206, 'impression');Ars Technica

    Related Stories

    IBM Releases Granite 4.2 LLM Models Focused on Reasoning and Local Deployment
    26 Aug · 11:21 AM
    Alibaba Previews Qwen 4 Architecture with Qwen 3.8-Flash-Next Model
    25 Aug · 9:21 PM
    ChatGPT for Teachers expands to over 300,000 educators
    26 Aug · 5:21 PM
    Researchers Shrink AI Model, Making It Smarter Than Original
    25 Aug · 8:51 PM
    China's Moonshot AI in talks with Microsoft, Amazon, Google over Kimi K3 revenue sharing
    26 Aug · 7:56 AM