Google AI Releases Gemini 3.5 Live Translate Supporting Over 70 Languages
Updated: Jul 20
Google AI released Gemini 3.5 Live Translate. The model handles live audio streams across more than 70 languages. It keeps original tone, rhythm and pitch during translation.
The update targets real conversations. Grab, the Southeast Asian super app, already tests the tool for drivers and riders who speak different languages. Those users start over 10 million voice calls each month.
Direct audio processing changes speed
Gemini 3.5 Live Translate ingests raw audio instead of first converting speech to text. This step removes one conversion layer. Latency drops to near real time.
Developers reach the model through the Gemini Live API. They can pair it with LiveKit, Fishjam, Pipecat or Vision Agents. LiveKit already runs virtual meeting rooms that translate speech on the fly.
Grab shows the first large scale test
Grab logs more than 10 million monthly voice calls between drivers and passengers. Many cross language lines. The company now explores Gemini 3.5 Live Translate to reduce miscommunication during rides.
Early tests focus on common phrases about directions, prices and traffic. The model must preserve urgency and politeness while switching languages.
Software partners push streaming limits
Software Mansion combined the model with the MoQ protocol. The combination cuts buffering during live streams. VisionAgents AI showed instant language switches inside a single conversation.
These integrations demonstrate that the audio pipeline can handle variable network conditions without dropping words.
The main pressure lands on earlier translation services
Existing cloud translation products rely on separate speech recognition then machine translation steps. Gemini 3.5 Live Translate skips that sequence. Rivals must now decide whether to rebuild their stack or lose ground on latency.
Microsoft and Amazon offer similar products. Neither has announced a single model that processes raw audio end to end for over 70 languages at the same claimed speed.
Limits remain around context and accents
The company states the model works across more than 70 languages. It does not yet detail performance on rare dialects or strong regional accents. Independent tests have not appeared.
Grab will run longer pilots before wider rollout. Any drop in accuracy during noisy rides could slow adoption.
What to watch next
Three signals will show whether the approach holds. First, Grab will publish internal accuracy numbers after the pilot ends. Second, LiveKit will release call volume data from its multi language rooms. Third, Google will update the Gemini Live API changelog with supported language additions or removals.
Each data point will either reinforce the latency claim or expose gaps that competitors can attack.



