Fluid, natural voice translation with Gemini 3.5 Live Translate
Google is rolling out Gemini 3.5 Live Translate, a speech-to-speech model for 70+ languages that aims for natural, near real-time translation.
Intelligence analysis by GPT-5.4 Mini

Google says its new audio model, Gemini 3.5 Live Translate, translates speech continuously rather than waiting for turn boundaries, so conversations stay closer to the speaker’s pace. The company is rolling it out across Gemini Live API, Google Meet, and the Google Translate app.
Google made a talking translator that tries to keep up like a live interpreter in a headset, instead of stopping and starting after every sentence. It works in many languages and is meant to sound more natural, like a person who can whisper the meaning right away.
Analysis
What Google announced
Google introduced Gemini 3.5 Live Translate, its latest audio model for live speech-to-speech translation. The company says it automatically detects 70+ languages and produces translated speech that preserves intonation, pacing, and pitch.
What is different
Unlike turn-by-turn translation systems that wait for a speaker to finish, this model processes speech as it is streamed. Google says that lets it stay only a few seconds behind the speaker while avoiding awkward pauses. It is also designed to handle multilingual input without manual setup and to work in noisy, unpredictable environments.
Where it is coming first
The rollout starts today in three places: public preview for developers through the Gemini Live API and Google AI Studio, private preview for enterprises in Google Meet starting this month, and the Google Translate app on Android and iOS. In Meet, Google says the update expands speech translation from five languages to 70+ and enables more than 2,000 language combinations in a meeting.
Product details and partners
Google also says Android users will get a new listening mode that plays translations through the phone’s earpiece, which may help in situations where headphones are not available. The company says all generated audio is watermarked with SynthID so AI audio remains detectable. It also points to ecosystem partners such as Agora, Fishjam, LiveKit, Pipecat, Vision Agents, and Grab, which is testing the model for near real-time communication between drivers and travelers.
The article frames this as both a consumer feature and a developer platform capability, with Google emphasizing live interpretation for meetings, lessons, broadcasts, and similar use cases.
Key points
- Google launched Gemini 3.5 Live Translate for near real-time speech-to-speech translation in 70+ languages.
- The model translates continuously instead of waiting for speakers to finish, aiming to reduce pauses and latency.
- Google says the system preserves tone, pacing, and pitch and can handle noisy, multilingual environments.
- The feature is rolling out to the Gemini Live API, Google Meet, and the Google Translate app on Android and iOS.
- Google says all generated audio is watermarked with SynthID to keep AI audio detectable.
If the model works as described, it could make cross-language conversations feel much less clunky in meetings, travel, and classroom settings. The wider language coverage and faster response time could make live translation useful to more people in more places.
Real-time speech translation still has to balance speed, accuracy, and naturalness, so mistakes or delays could reduce trust in important conversations. The rollout is also staged through previews, which suggests the experience may still need tuning before it feels consistently reliable at scale.



