Google Deploys Gemini 3.5 for Multilingual Audio Tools
Google is expanding its AI language capabilities with Gemini 3.5 models and open-source datasets, aiming to bridge the digital divide for underrepresented languages and offline users.

Google has introduced several AI language updates, shifting from traditional text translation to native audio intelligence. The company's Gemini 3.5 Live Translate now powers real-time spoken translation across 70 languages and more than 2,000 language pairs, capturing emotional cues and code-switching. Alongside it, Gemini 3.5 Transcribe serves as a precise speech-to-text model that filters noise and jargon, powering features like Rambler on Android Gboard. To support its broader 1,000 Languages Initiative, Google trained its Universal Speech Model on 12 million hours of audio, leveraging cross-lingual transfer learning to assist under-resourced languages.
To gather data for underrepresented languages, Google has established grassroots partnerships. The WAXAL dataset covers 27 Sub-Saharan African languages spoken by over 100 million people across more than 26 countries. In India, Project Vaani has mapped over 30,000 hours of speech across 109 languages from 155,000 speakers. Additionally, the Amplify Initiative gathered 15,000 multimodal data points with help from 1,600 local experts and 20 universities. Developers can explore these resources through Language Explorer, an interactive tool visualizing LinguaMeta, which maps over 7,000 languages.
For environments with limited internet, Google released TranslateGemma, a family of lightweight open models trained across 55 languages that run directly on-device. For users with basic feature phones, Google supports the Ask Viamo Anything voice assistant, which has answered over 2 million questions in Rwanda using Gemini. Google is also addressing accessibility with Sign Language-to-Text, trained on over 50 sign languages, which enables sign-to-text dictation on Pixel 11 starting with American Sign Language.
For AI practitioners and developers, these releases provide critical open-source datasets and lightweight models to build highly localized applications. By utilizing TranslateGemma and the LinguaMeta repository, developers can deploy speech and translation tools that function offline and respect regional nuances, such as the Maori pronunciations integrated into Google Maps. This shift allows engineers to move away from rigid text-to-speech pipelines and build applications that naturally process real-world human communication.
This is our own summary of reporting by Google AI Blog



