2026年6月9日
|
5分钟阅读
Gemini 3.5 Live Translate 是我们最新的音频模型,可在超过70种语言中实现近乎实时的语音到语音翻译。
Anuda Weerasinghe
产品经理
Tony Lu
高级首席软件工程师

收听文章
[[时长]]分钟
此内容由 Google AI 生成。生成式 AI 为实验性技术。
二十年前,Google 翻译作为我们开创性的机器学习实验之一起步,旨在将语言科学转化为人类连接的魔力。如今,该实验已取得长足进展,每月为数十亿用户翻译超过一万亿个单词。
今天,我们迈出了下一步,发布了 Gemini 3.5 Live Translate,这是我们用于实时语音到语音翻译的最新音频模型。
该模型可自动检测70多种语言,并生成流畅、自然的翻译语音,同时保留说话者的语调、语速和音高。与需要等待说话者说完才响应的逐轮系统不同,3.5 Live Translate 持续生成语音,在等待上下文以提高质量与立即翻译以保持与说话者同步之间取得平衡。它提供流畅的音频,没有尴尬的停顿,在整个会话过程中仅比说话者延迟几秒。
Gemini 3.5 Live Translate 从今天起开始在 Google 产品中推出:
- 面向开发者,通过 Gemini Live API 和 Google AI Studio 提供公开预览
- 面向企业,从本月起在 Google Meet 中提供私有预览
- 面向所有人,通过 Android 和 iOS 上的 Google 翻译提供
使用 3.5 Live Translate 进行构建
Gemini 3.5 Live Translate 在语音流式传输时进行处理,实现跨语言更无缝的连接。该模型处理多语言输入,无需手动配置设置。同时,其噪声鲁棒性确保应用能够处理嘈杂、不可预测的环境。您可以使用其功能来促进多语言通话、会议、课程、广播等的实时口译。
观看 Gemini Live API 的实际应用,实现配音和同步多语言翻译。深入了解 demo 或 Gemini Cookbook 中的更多示例代码。
通过利用 Gemini Live API,像 Agora、Fishjam、LiveKit、Pipecat 和 Vision Agents 这样的开发者平台使开发者能够轻松构建和部署语音翻译应用。这些集成处理复杂的实时媒体流基础设施,因此开发者可以专注于用户体验。
我们的合作伙伴 Grab 正在测试该模型,以实现司机和乘客在接送点之间近乎实时的多语言交流。这些用户每月通过 Grab 进行超过1000万次语音通话。
了解 Grab 如何测试 3.5 Live Translate 以改变用户之间的沟通。
阅读早期评价
除了 Grab,CJ ENM、LiveKit 等公司也对 3.5 Live Translate 给予了积极反馈,称赞其令人印象深刻的翻译质量、准确性和低延迟:
在视频会议中体验 3.5 Live Translate
Google Meet 中的语音翻译将很快使用 3.5 Live Translate,通过以下方式改善体验:
- 提供70多种语言,相比之前仅限五种语言有所改进,
- 支持在一次会议中跨越2000多种语言组合进行对话,从之前仅支持与英语互译的状态扩展,
- 更新界面以提供对语音翻译的即时访问。
我们从本月开始向选定的商业 Google Workspace 客户以私有预览形式推出此更新,随后将在今年晚些时候更广泛地推出。
Google Meet 参与者使用语音翻译在英语、普通话和瑞典语之间进行交流。
在 Android 或 iOS 上的 Google 翻译应用中获取 3.5 Live Translate
该模型也正在全球范围内的 Google 翻译应用中推出,适用于 Android 和 iOS。使用实时翻译功能时,只需连接任何一副耳机,即可体验更无缝的翻译,该翻译能反映说话者的语调,覆盖70多种语言。
对于 Android 用户,我们还开始推出带有 3.5 Live Translate 的新“聆听模式”,让您可以直接通过手机听筒听到翻译。只需像普通通话一样将手机贴近耳朵,翻译后的音频就会直接传输给您。当您想快速听到翻译而不想让别人听到,且手边没有耳机时,这种新体验会很有帮助。
使用新的聆听模式,用户可以直接通过手机听筒听到西班牙语导游讲解的近乎实时的英语翻译。
使用 SynthID 添加水印
我们模型生成的所有音频都使用 SynthID 添加了水印。这种不可察觉的水印直接嵌入音频输出中,确保 AI 生成的内容保持可检测性,以帮助防止错误信息。有关我们安全和责任方法的详细信息,请查看模型卡片。
Jun 09, 2026
|
5 min read
Gemini 3.5 Live Translate is our latest audio model, delivering near real-time speech-to-speech translation in over 70 languages.
Anuda Weerasinghe
Product Manager
Tony Lu
Senior Staff Software Engineer

Listen to article
[[duration]] minutes
This content is generated by Google AI. Generative AI is experimental
Twenty years ago, translation at Google began as one of our pioneering machine learning experiments to turn the science of language into the magic of human connection. That experiment has come a long way with over a trillion words being translated for billions of users across our products every month.
Today, we’re taking our next step with the release of Gemini 3.5 Live Translate, our latest audio model for live speech-to-speech translation.
The model automatically detects 70+ languages and generates smooth, natural-sounding translated speech that preserves the speakers' intonation, pacing and pitch. Unlike turn by turn systems that wait for the speaker to finish speaking before responding, 3.5 Live Translate generates speech continuously, balancing the trade-off between waiting for context to improve quality and translating immediately to stay in sync with the speaker. It delivers fluid audio without awkward pauses and stays just a few seconds behind the speaker throughout the session.
Gemini 3.5 Live Translate is rolling out starting today across Google products:
- For developers in public preview via the Gemini Live API and Google AI Studio
- For enterprises in private preview starting this month in Google Meet
- For everyone via Google Translate on Android and iOS
Build with 3.5 Live Translate
Gemini 3.5 Live Translate processes speech as it’s streamed, enabling a more seamless connection across languages. The model handles multilingual inputs without the need to manually configure settings. At the same time, its noise robustness ensures applications can handle loud, unpredictable environments. You can use its capabilities to help facilitate live interpretation for multilingual calls, meetings, lessons, broadcasts and more.
Watch the Gemini Live API in action, enabling dubbing and simultaneous multi-language translation. Dive into the demo or more example code in the Gemini Cookbook.
By utilizing the Gemini Live API, developer platforms like Agora, Fishjam, LiveKit, Pipecat, and Vision Agents enable developers to build and deploy voice translation apps with ease. These integrations handle the complex real-time media streaming infrastructure, so developers can focus on the user experience.
Our partners at Grab are testing the model to enable multilingual communication in near real-time between drivers and travelers at pickups. These users make over 10 million voice calls per month through Grab.
See how Grab has been testing 3.5 Live Translate to transform communication between users.
Read the early reviews
In addition to Grab, companies like CJ ENM, LiveKit and others have shared positive feedback on 3.5 Live Translate highlighting its impressive translation quality, accuracy and low latency:
Experience 3.5 Live Translate in your video meetings
Speech translation in Google Meet will soon use 3.5 Live Translate, improving the experience by:
- Offering 70+ languages, an improvement from the previous limit of just five languages,
- Enabling conversations across over 2000+ language combinations in one meeting, expanding from the previous state of only translating to and from English,
- Updating the interface to provide instant access to speech translation.
We’re launching this update in private preview for select business Google Workspace customers starting this month, followed by a broader rollout later this year.
Google Meet participants use speech translation to communicate across English, Mandarin, and Swedish.
Get 3.5 Live Translate in the Google Translate app on Android or iOS
The model is also rolling out on the Google Translate app globally, on both Android and iOS. When using the Live translate feature, simply connect any pair of headphones to experience a more seamless translation that mirrors the speaker’s tone across 70+ languages.
For Android users, we’re also starting to roll out a new ‘listening mode’ with 3.5 Live Translate that lets you hear translations directly through your phone’s earpiece. Simply hold your phone to your ear just like a regular call, and the translated audio streams straight to you. This new experience can be helpful in situations where you want to quickly hear translations without others hearing, and you don’t have your headphones handy.
Using the new listening mode, users can hear a near real-time English translation of a guided tour in Spanish directly through their phone's earpiece.
Watermarked with SynthID
All audio generated by our models is watermarked with SynthID. This imperceptible watermark is woven directly into the audio output, ensuring AI-generated content remains detectable to help prevent misinformation. For details on our approach to safety and responsibility, review the model card.
本文内容采集自官方网站,排版和翻译可能与原页面存在差异。
阅读官方全文