Google unveils Gemini 3.5 Transcribe for speech-to-text in 85+ languages

By The Desk
Tweet image from @Nairametrics

Google has launched Gemini 3.5 Transcribe, a speech-to-text model built for more accurate, natural audio transcription across over 85 languages.

Share

Google has unveiled Gemini 3.5 Transcribe, a new speech-to-text model designed to deliver more accurate and natural audio transcription across more than 85 languages, according to a report published by Nairametrics on X.

The announcement marks another expansion of Google's Gemini model family, which has grown beyond generative text and multimodal capabilities into specialized audio infrastructure. With Gemini 3.5 Transcribe, Google is targeting a persistent weakness in automated transcription: the gap between recognizing words and rendering them into coherent, natural-sounding text. The company's framing — "more accurate and natural audio transcription" — points to improvements in both raw word recognition and the fluency of the final output, a distinction that matters for real-world use cases such as meeting notes, captions, and voice-to-text applications.

Speech-to-text accuracy has long varied widely across languages, accents, and recording conditions. Models trained primarily on English-language data often degrade sharply when applied to lower-resource or less widely spoken languages. A model supporting 85+ languages signals an investment in breadth, though Google has not yet published the specific languages included, nor any comparative benchmarks against its existing speech services or competing products such as OpenAI's Whisper.

The announcement was shared on X by Nairametrics, a Nigerian financial news publication, with a post that linked to further coverage. As of the time of reporting, the post had drawn limited engagement — nine likes, one retweet, and one reply — suggesting the news has not yet broken widely across developer and AI communities. The original post did not include additional technical specifications, sample transcriptions, or statements from Google executives.

Google has not stated whether Gemini 3.5 Transcribe will be offered as a standalone API, integrated into Google Cloud's existing speech-to-text products, or embedded within consumer-facing tools like Google Meet, Recorder, or YouTube's automatic captioning system. The absence of pricing details and an availability timeline leaves open questions about how quickly developers can begin testing the model.

The launch also raises a competitive question. OpenAI has established Whisper as a widely used open-source transcription model, and specialized providers such as AssemblyAI and Deepgram have built businesses around high-accuracy speech recognition. Google's move with Gemini-branded transcription suggests the company intends to compete on both accuracy and language coverage while leveraging the Gemini name to signal a generational improvement over its older speech models.

What comes next depends on Google's follow-up disclosures. Developers, accessibility advocates, and enterprise customers will be watching for the publication of benchmark results across supported languages, details on real-time versus batch transcription modes, and whether the model supports punctuation, speaker diarization, and timestamping out of the box. Google has not indicated when those details will be released.

Share this article

Help others discover this story

https://www.techblit.com/google-unveils-gemini-35-transcribe-for-speech-to-text-in-85-languages