我们正在推出 GPT‑Live,新一代语音模型,让与 AI 的对话更像真实交流。
GPT‑Live 基于全双工架构构建,意味着它能同时听和说。在对话中,GPT‑Live 可以用“嗯哼”或“对”这样的短语表示它在认真听,进行快速来回交流,或者在你需要思考时保持安静。结果就是,一种令人耳目一新的、易于交谈的语音体验。
GPT‑Live 也是我们迄今为止最智能的语音模型。对于需要网络搜索、更深层推理或更复杂工作的问题,它会在后台委托给最新的前沿模型,并在准备好时将结果带回对话中。在工作时,GPT‑Live 可以继续与你交谈,保持对话流畅。上线时,GPT‑Live 将在后台使用 GPT‑5.5。随着我们发布新的前沿模型,我们将持续更新 GPT‑Live 所使用的模型。
这些进步驱动了全新的 ChatGPT 语音体验,使其更智能、更自然易用。随着时间的推移,我们相信这项研究还将解锁使用语音进行日益复杂、更长时间运行和更具代理性的工作的能力。
我们今天开始向全球 ChatGPT 用户推出两个版本的 GPT‑Live——GPT‑Live‑1 和 GPT‑Live‑1 mini。我们还计划很快将它们引入 API,开发者和企业可以使用此表格注册以接收通知。
我们的愿景是实现真正自然的人机交互:一个与 AI 协作感觉像与另一个人合作一样流畅和响应迅速的世界,同时推理和复杂任务执行在后台无缝进行。
旧一代的语音 AI 系统让我们更接近这一愿景,但伴随着重要的权衡。
级联语音系统依赖一系列模型依次处理每一轮对话。最初的 ChatGPT Voice 将三个模型串联在一起:一个语音转文本模型转录你的语音,一个大语言模型生成回复,一个文本转语音模型将其转换回语音。这种方法使我们首次能够与前沿 AI 模型对话,但复杂性带来了代价:信息可能在模型间丢失,且响应缓慢而僵硬。
级联语音系统
响应缓慢而僵硬,长时间停顿
转录文本
与标准语音模式(使用 GPT-5.5 Instant)的示例对话
像 ChatGPT 高级语音模式 这样的基于轮次的语音模型在单个模型内处理和生成音频,减少了延迟,使对话更流畅——但它们仍然通过离散的轮次运作。模型必须等待用户停止说话才能响应,导致僵硬的来回交流。此外,由于轮次检测基于静默,即使是短暂的停顿或背景噪音也可能被误认为是轮次结束——导致模型在不自然的时间点打断。
基于轮次的语音模型
响应稍快且更流畅,但与模型的来回交流仍然感觉僵硬
转录文本
与 ChatGPT 高级语音模式的示例对话
GPT‑Live 通过两项架构变化解决了这些限制。
首先,我们使用全双工架构构建了 GPT‑Live,以实现持续交互。GPT‑Live 不是处理一系列独立消息,而是在生成输出的同时持续处理输入。因此,模型可以每秒多次做出交互决策:是否说话、继续倾听、暂停、打断或调用工具。
这使得模型能够进行更自然的来回交流,保持更好的时间感,甚至执行实时翻译。
持续交互
快速、自然、富有表现力的响应和更积极的倾听
转录文本
与 GPT-Live-1(使用 GPT-5.5 Instant)的示例对话
其次,我们将处理持续交互的 GPT‑Live 与更深层的工作解耦。当问题需要搜索、推理或更强大的代理能力时,GPT‑Live 可以将任务委托给另一个模型(如 GPT‑5.5)。这使得它即使在后台处理多个任务时,也能保持对话继续。
这一架构变化还允许 GPT‑Live 持续使用最新的模型和代理,将前沿智能与自然交互相结合。
委托进行更深层工作
GPT-Live 提供快速、自然的响应,而 GPT-5.5 在后台处理搜索
转录文本
与 GPT-Live-1(使用 GPT-5.5 Instant)的示例对话
我们构建了新的人工评估来衡量愉悦度和对话流畅度。在这些头对头比较中,在匹配的 5-10 分钟对话中,GPT‑Live‑1 和 GPT‑Live‑1 mini 在总体偏好、轮次切换、打断、对话流畅度以及每次交互的自然感方面,都明显优于高级语音模式。
GPQA:GPT‑Live‑1 在 GPQA 上大幅优于高级语音模式,该测试评估生物学、化学和物理学领域的专家级科学推理能力。
BrowseComp:GPT‑Live‑1 在 BrowseComp 上相比高级语音模式有显著提升,该测试评估代理性网络搜索和查找难以定位信息的能力。
τ³-Voice Telecom(内部变体)**:GPT‑Live‑1 在 τ³-Voice Telecom 上优于高级语音模式,该测试评估语音代理在真实的多轮电信支持任务中的表现。
GPT‑Live‑1(即时版)和 GPT‑Live‑1 mini 在后台使用 GPT‑5.5 Instant 模型,而 GPT‑Live‑1 Medium 和 GPT‑Live‑1 High 使用 GPT‑5.5 Thinking 模型,并采用中等和高推理努力。 ** 我们为此评估使用了由最新推理模型驱动的定制用户模型。
每周,超过 1.5 亿人使用语音和听写等功能与 ChatGPT 交谈。他们用它来获取免提的日常帮助、练习语言、讲睡前故事,或在通勤时聊天。
从今天开始,当你点击语音按钮与 ChatGPT 交谈时,你将获得由 GPT‑Live 驱动的改进体验——更自然的对话、更智能的回答、更好的倾听和视觉响应。
现在与 ChatGPT 交谈应该感觉更像真正的对话。你可以用问题打断,暂停整理思绪,或要求 ChatGPT 放慢速度。它会自然地用“嗯哼”或“明白了”这样的短语确认你在说什么,让你知道它在跟上。我们还为 GPT‑Live 重新制作了 ChatGPT 中的九种不同声音。
ChatGPT 语音现已搭载我们最新的前沿模型,在您需要时提供更智能的回答。您还可以根据需求选择推理级别:即时模式可快速响应,中或高模式则让 ChatGPT 花更多时间思考。
当您需要思考片刻时,ChatGPT 语音现在会耐心等待,而非贸然插话打断。若您要求它保持安静倾听,它也会照做。面对过往车辆或附近交谈等背景噪音时,ChatGPT 能更专注于您的声音而不受干扰。
有些答案通过可视化呈现会更实用。在对话过程中,ChatGPT 现在可针对天气、股票、体育等主题展示丰富的视觉卡片。语音功能仍支持搜索、记忆、图像及文件上传。
这使得 ChatGPT 语音体验在日常使用中更自然、更强大、更实用。
GPT‑Live 默认以安全为设计核心。它在继承最新模型安全改进的基础上,针对关键风险领域增加了专项安全训练,并专门为语音场景设计了新型防护措施。
为更真实反映用户在实际场景中的语音使用方式,我们首先将安全测试扩展至全新的原生音频评估体系。同时借鉴高级语音模式的经验,创建了基于生成音频的合成评估,以更聚焦关键安全领域,包括自残、精神病与躁狂、对AI的情感依赖、暴力及色情内容。内部专家还对模型进行了语音特有风险的红队测试。
测试结果显示,GPT‑Live 在几乎所有评估领域的表现均与高级语音模式相当或更优。更多测试与防护详情可参阅 GPT‑Live 系统卡(在新窗口打开)。
由于语音对话实时进行,我们还构建了可在模型说话时即时响应的防护机制。当系统检测到潜在不安全输出时,可引导模型生成更安全的回应,显示额外安全提示或资源,或在高风险情况下终止语音对话。针对涉及自残的对话,我们调整了 ChatGPT 的语音支持流程,包括提供经专家验证的危机热线支持。
我们设计了额外保护措施以支持青少年用户,并在模型中直接训练了适龄行为模式以降低不当回应风险。家长可通过家长控制功能选择是否允许青少年使用 ChatGPT 语音,当涉及潜在自残或自杀意图等高危情况时,关联家长可能收到通知。
我们还将推出针对情感依赖的长期测量与发布后监测,以持续改进认知并完善防护措施。基于此前对情感使用与心理健康的研究,这将帮助我们识别新兴模式,优化系统在情感敏感互动中的回应方式。
最后,GPT‑Live 专为对话设计,不涉及声音模仿。它使用 ChatGPT 中预设的语音集,并配备防护措施防止模仿真人声音。
我们致力于保障安全与福祉,并将根据实际使用经验持续强化这些防护措施。
GPT‑Live 现已面向全球 iOS、Android 及 ChatGPT.com(在新窗口打开) 的 ChatGPT 用户逐步推出。GPT‑Live‑1 将成为 Go、Plus 及 Pro 用户 ChatGPT 语音的默认模型,GPT‑Live‑1 mini 则成为免费用户的默认模型。更多可用性详情请参阅我们的帮助中心(在新窗口打开)。
我们已针对 ChatGPT 中最常用的语言优化了 GPT‑Live。对于某些语言,模型可能存在非母语口音或流利度不足的问题。我们正积极致力于改善多语言体验。
发布初期,GPT‑Live 暂不支持 ChatGPT 中的视频语音或屏幕共享功能,但我们正努力尽快引入这些能力。您仍可访问包含标准语音模式和高级语音模式在内的 ChatGPT 语音旧版本,这些功能在这些版本中可用。
We’re launching GPT‑Live, a new generation of voice models that make talking with AI feel much more like having a real conversation.
GPT‑Live is built on a full-duplex architecture, meaning it can listen and speak at the same time. During conversations, GPT‑Live can show it’s paying attention with phrases like “mhmm” or “yeah”, engage in quick back-and-forth, or just stay quiet when you need a moment to think. The result is a voice experience that is refreshingly easy to talk to.
GPT‑Live is also our smartest voice model yet. For questions that require web search, deeper reasoning, or more complex work, it delegates to our latest frontier model behind the scenes and brings the result back into the conversation when it’s ready. While it works, GPT‑Live can keep talking with you and maintain the flow of conversation. At launch, GPT‑Live will use GPT‑5.5 in the background. As we release new frontier models, we’ll continuously update the model used by GPT‑Live.
These advances power a new ChatGPT Voice experience that is more intelligent and natural to use. Over time, we believe this research will also unlock the ability to use voice for increasingly complex, longer-running, and more agentic work.
We’re beginning to roll out two versions of GPT‑Live – GPT‑Live‑1 and GPT‑Live‑1 mini – to ChatGPT users globally today. We also plan to bring them to the API soon, and developers and enterprises can sign up to be notified using this form.
Our vision is to enable truly natural human–AI interaction: a world where collaborating with AI feels as fluid and responsive as working with another person, while reasoning and complex task execution happen seamlessly in the background.
Older generations of voice AI systems brought us closer to that vision, but with important tradeoffs.
Cascaded voice systems rely on a series of models acting one after another to process each turn. The original ChatGPT Voice chained three models together: a speech-to-text model to transcribe your speech, a large language model to produce a response, and a text-to-speech model to convert it back into speech. This approach enabled us to talk to frontier AI models for the first time, but the complexity came at a cost: information could be lost across models, and responses were slow and stilted.
Cascaded voice system
Slow and stilted responses, long pauses
Transcript
Example conversation with Standard Voice Mode, using GPT-5.5 Instant
Turn-based voice models like ChatGPT Advanced Voice Mode processed and generated audio within a single model, reducing latency and making conversations smoother — but they still operated through discrete turns. The model had to wait for the user to stop speaking before responding, resulting in rigid back-and-forth. In addition, because turn detection is based on silence, even a brief pause or background noise could be mistaken for the end of turn — causing the model to interrupt at unnatural times.
Turn-based voice model
Slightly faster and smoother responses, but back-and-forth with the model still feels rigid
Transcript
Example conversation with ChatGPT Advanced Voice Mode
GPT‑Live addresses these limitations through two architectural changes.
First, we built GPT‑Live for continuous interaction using a full-duplex architecture**.**Instead of processing a sequence of separate messages, GPT‑Live continuously processes input while generating output. The model can therefore make interaction decisions many times per second: whether to speak, continue listening, pause, interrupt, or invoke a tool.
This allows the model to engage in more natural back-and-forth, maintain a better sense of time, and even perform live translation.
Continuous interaction
Fast, natural, expressive responses and more active listening
Transcript
Example conversation with GPT-Live-1, using GPT-5.5 Instant
Second, we decoupled GPT‑Live — which handles continuous interaction — from deeper work. When a question requires search, reasoning, or more agentic capabilities, GPT‑Live can delegate the task to another model like GPT‑5.5. This allows it to keep the conversation going, even as it handles multiple tasks in the background.
This architectural change also allows GPT‑Live to continuously use the latest models and agents, combining frontier intelligence with natural interaction.
Delegation for deeper work
GPT-Live provides fast, natural responses, while GPT-5.5 handles search in the background
Transcript
Example conversation with GPT-Live-1, using GPT-5.5 Instant
We built new human evaluations to measure pleasantness and the flow of conversation. In these head-to-head comparisons, GPT‑Live‑1 and GPT‑Live‑1 mini are strongly preferred over Advanced Voice Mode in matched 5–10 minute conversations that measure overall preference, turn-taking, interruptions, conversational flow, and how natural each interaction felt.
GPQA: GPT‑Live‑1 substantially outperforms Advanced Voice Mode on GPQA, which tests expert-level scientific reasoning across biology, chemistry, and physics.
BrowseComp: GPT‑Live‑1 shows strong gains over Advanced Voice Mode on BrowseComp, which tests agentic web search and the ability to find difficult-to-locate information.
τ³-Voice Telecom (internal variant)**: GPT‑Live‑1 outperforms Advanced Voice Mode on τ³-Voice Telecom, which tests voice agents on realistic, multi-turn telecom support tasks.
GPT‑Live‑1 (instant) and GPT‑Live‑1 mini use the GPT‑5.5 Instant model in the background, while GPT‑Live‑1 Medium and GPT‑Live‑1 High use the GPT‑5.5 Thinking model with medium and high reasoning effort.
** We used a customized user model, powered by our latest reasoning models, for this eval.
Each week, more than 150 million people talk to ChatGPT using features like Voice and Dictation. They use it to get hands-free everyday help, to practice languages, tell bedtime stories, or just chat during their commute.
Starting today, when you tap the Voice button to talk with ChatGPT, you’ll get an improved experience powered by GPT‑Live—with more natural conversations, smarter answers, better listening, and visual responses.
Talking with ChatGPT should now feel much more like a real conversation. You can interrupt with a question, pause to gather your thoughts, or ask ChatGPT to slow down. It naturally acknowledges what you’re saying with phrases like “mhmm” or “got it,” so you know it’s following along. We’ve also remastered the nine distinct voices in ChatGPT for GPT‑Live.
ChatGPT Voice can now draw on our latest frontier models, giving you smarter answers when you need them. You can also choose the level of reasoning that fits your needs: Instant for fast responses, or Medium and High when you want ChatGPT to spend more time thinking.
If you take a moment to think, ChatGPT Voice now waits instead of jumping in and interrupting. If you ask it to stay quiet and listen, it will. And when there’s background noise, like passing traffic or nearby conversations, ChatGPT is better at focusing on your voice instead of getting distracted.
Some answers are more useful when you can see them. While you’re talking, ChatGPT can now show rich visual cards for topics like weather, stocks, sports, and more. Voice also continues to support search, memory, images, and file uploads.
The result is a ChatGPT Voice experience that feels more natural, more capable, and more useful in everyday life.
GPT‑Live was designed to be safe by default. It builds on the safety advances from our latest models while adding dedicated safety training across key risk areas and new safeguards designed specifically for voice.
To better reflect how people use voice in real-life settings, we began by expanding our safety testing to include new audio-native evaluations. We also created synthetic evaluations that use generated audio to focus more intensively on key safety areas, drawing on what we learned from Advanced Voice Mode. Those areas include self-harm, psychosis and mania, emotional reliance on AI, violence, and sexual content. Internal experts also red-teamed the model for risks unique to voice.
In our testing, GPT‑Live performed comparably to or better than Advanced Voice Mode across nearly all of the areas we evaluated. You can read more about our testing and safeguards in the GPT‑Live system card(opens in a new window).
Because voice conversations unfold in real time, we also built safeguards that can act while the model is speaking. When the system detects potentially unsafe output, it can steer the model toward a safer response, surface additional safety messaging or resources, or end the voice conversation in higher-risk cases. For conversations involving self-harm, we adapted ChatGPT’s support flows for voice, including offering expert-vetted crisis helpline support.
We designed additional protections to support teen users, and trained age-appropriate behavior directly into the model to reduce the risk of inappropriate responses. Parents can choose whether their teen can use ChatGPT Voice through Parental Controls, and linked parents may be notified in higher-risk situations involving signs of potential self-harm or suicidal intent.
We’re also rolling out longer-term measurement and post-launch monitoring focused on emotional reliance to continue improving our understanding and refining safeguards. Building on our previous research into affective use and emotional well-being, this will help us identify emerging patterns and improve how the system responds in emotionally sensitive interactions.
Finally, GPT‑Live is designed for conversation, not voice impersonation. It uses a set of predefined voices in ChatGPT, with safeguards to prevent it from imitating a real person’s voice.
We’re committed to supporting safety and well-being and will keep strengthening these protections as we learn from real-world use.
GPT‑Live is rolling out now to ChatGPT users globally across iOS, Android, and ChatGPT.com(opens in a new window). GPT‑Live‑1 will become the default model powering ChatGPT Voice for Go, Plus, and Pro users, and GPT‑Live‑1 mini will become the default for Free users. More availability details can be found in our Help Center(opens in a new window).
We’ve optimized GPT‑Live for some of the most popular languages in ChatGPT. For certain languages, the model may have a non-native accent or gaps in fluency. We’re actively working to improve the experience across languages.
At launch, GPT‑Live will not support voice with video or screen sharing in ChatGPT, but we’re working to introduce these capabilities soon. You can still access legacy versions of ChatGPT Voice, including Standard and Advanced Voice Mode, where these features are available.
本文内容采集自官方网站,排版和翻译可能与原页面存在差异。
阅读官方全文