字号 ·· | 护眼
theregister

谷歌Gemini可以呈现卡通或逼真的虚拟形象,为生成式对话进行唇形同步Google Gemini can present cartoon or lifelike avatars to lip-sync generative chatter

点「原文对照」整页切到原文,或双击某段只看那段的原文。

谷歌为其实时对话模型 Gemini 3.8 Live 增加了一项名为“Live Avatar”的功能,该功能能够生成动画角色和逼真的数字人,为 Gemini 机器生成的语音进行对口型配音。谷歌 DeepMind 研究科学家张硕茵(Shuo-yiin Chang)和 Gemini 软件工程师郑承杰(CJ Zheng)在一篇博文中解释道:“对话本质上是多模态的:我们通过倾听、注视、说话以及使用面部表情来进行交流。Live Avatar 将这些能力引入了企业智能体。通过同时处理视觉和音频输入,它能生成丰富的对话,从而带来更全面的体验。”

Google has augmented its live dialogue model Gemini 3.8 Live with a feature called "Live Avatar" that's capable of conjuring animated characters and realistic human simulacra to lip-sync Gemini's machine-generated speech. "Conversation is inherently multimodal: we listen, look, speak, and use facial expressions to communicate," explained Shuo-yiin Chang, research scientist at Google DeepMind, and CJ Zheng, Gemini software engineer, in a blog post. "Live Avatar brings these capabilities to enterprise agents. By processing visual and audio inputs simultaneously, it generates enriching conversations for a more comprehensive experience." Five years ago, DeepMind warned about various risks associated with the use of large language models, among them human interaction harms like anthropomorphising AI models – the attribution of human characteristics to software-based interactions. Anthropomorphism here applies broadly to visual human imitation as well as human-like speech and text. "Anthropomorphizing [language models] may inflate users' estimates of the conversational agent’s competencies," noted the more than 20 authors of DeepMind's paper [PDF], "Ethical and social risks of harm from Language Models." "For example, users may falsely infer that a conversational agent that appears human-like in language also displays other human-like characteristics, such as holding a coherent identity over time, or being capable of empathy, perspective-taking, and rational reasoning. As a result, they may place undue confidence, trust, or expectations in these agents." That was in 2021. In August 2025, the risks posed by humanizing AI reached the courtroom when the parents of Adam Raine sued OpenAI, alleging [PDF] that ChatGPT contributed to their son's suicide. Among other things, the complaint lays blame on "anthropomorphic mannerisms calibrated to convey human-like empathy." By December 2025, Google boffins were still grappling with the need to understand the impact of humanizing AI. In a paper [PDF] titled "How Tech Workers Contend with Hazards of Humanlikeness in Generative AI," Mark Díaz, Renee Shelby, Eric Corbett, and Andrew Smart from Google Research reported that tech workers participating in focus groups expressed similar reservations to those cited in the company's prior work. "Tech workers expressed significant concern that humanlikeness fosters a false sense of reliability and trust shaping their concerns as both users and developers, linking to fluid, natural language and tone, which can obscure errors," they wrote. So even as Google AI researchers highlight the need to understand how the "humanlikeness" of AI relates to perceived hazards and responsible AI development, Google Gemini Enterprise is moving straight to deployment. Live Avatars can be drawn from a pre-made library of characters or a custom avatar can be deployed if enterprise allow-listing is enabled. Noting how Live Avatar can make tool calls and fetch data in the background, as well as converse in 97 languages, making it suitable for scenarios like handling hotel guest check-ins, Chang and Zheng insist that Google created these Gemini-generated personas with trust and safety in mind. "We built Live Avatar with strict safeguards designed to respect identity, and keep AI-generated content transparent," they said. ®

五年前,DeepMind 曾对使用大语言模型相关的各种风险发出过警告,其中包括人类交互带来的危害,例如对 AI 模型进行拟人化——即把人类特征赋予基于软件的交互。这里的拟人化既广义地适用于视觉上的拟人模仿,也适用于类人的语音和文本。

Google has augmented its live dialogue model Gemini 3.8 Live with a feature called "Live Avatar" that's capable of conjuring animated characters and realistic human simulacra to lip-sync Gemini's machine-generated speech. "Conversation is inherently multimodal: we listen, look, speak, and use facial expressions to communicate," explained Shuo-yiin Chang, research scientist at Google DeepMind, and CJ Zheng, Gemini software engineer, in a blog post. "Live Avatar brings these capabilities to enterprise agents. By processing visual and audio inputs simultaneously, it generates enriching conversations for a more comprehensive experience." Five years ago, DeepMind warned about various risks associated with the use of large language models, among them human interaction harms like anthropomorphising AI models – the attribution of human characteristics to software-based interactions. Anthropomorphism here applies broadly to visual human imitation as well as human-like speech and text. "Anthropomorphizing [language models] may inflate users' estimates of the conversational agent’s competencies," noted the more than 20 authors of DeepMind's paper [PDF], "Ethical and social risks of harm from Language Models." "For example, users may falsely infer that a conversational agent that appears human-like in language also displays other human-like characteristics, such as holding a coherent identity over time, or being capable of empathy, perspective-taking, and rational reasoning. As a result, they may place undue confidence, trust, or expectations in these agents." That was in 2021. In August 2025, the risks posed by humanizing AI reached the courtroom when the parents of Adam Raine sued OpenAI, alleging [PDF] that ChatGPT contributed to their son's suicide. Among other things, the complaint lays blame on "anthropomorphic mannerisms calibrated to convey human-like empathy." By December 2025, Google boffins were still grappling with the need to understand the impact of humanizing AI. In a paper [PDF] titled "How Tech Workers Contend with Hazards of Humanlikeness in Generative AI," Mark Díaz, Renee Shelby, Eric Corbett, and Andrew Smart from Google Research reported that tech workers participating in focus groups expressed similar reservations to those cited in the company's prior work. "Tech workers expressed significant concern that humanlikeness fosters a false sense of reliability and trust shaping their concerns as both users and developers, linking to fluid, natural language and tone, which can obscure errors," they wrote. So even as Google AI researchers highlight the need to understand how the "humanlikeness" of AI relates to perceived hazards and responsible AI development, Google Gemini Enterprise is moving straight to deployment. Live Avatars can be drawn from a pre-made library of characters or a custom avatar can be deployed if enterprise allow-listing is enabled. Noting how Live Avatar can make tool calls and fetch data in the background, as well as converse in 97 languages, making it suitable for scenarios like handling hotel guest check-ins, Chang and Zheng insist that Google created these Gemini-generated personas with trust and safety in mind. "We built Live Avatar with strict safeguards designed to respect identity, and keep AI-generated content transparent," they said. ®

DeepMind 论文《语言模型的伦理与社会危害风险》(PDF)的二十多位作者指出:“对[语言模型]进行拟人化可能会夸大用户对对话智能体能力的估计。例如,用户可能会错误地推断,一个在语言上表现得像人类的对话智能体,同时也展现出其他人类特征,比如拥有连贯的长期身份,或者具备同理心、换位思考和理性推理的能力。因此,他们可能会对这些智能体产生过度的信心、信任或期望。”

Google has augmented its live dialogue model Gemini 3.8 Live with a feature called "Live Avatar" that's capable of conjuring animated characters and realistic human simulacra to lip-sync Gemini's machine-generated speech. "Conversation is inherently multimodal: we listen, look, speak, and use facial expressions to communicate," explained Shuo-yiin Chang, research scientist at Google DeepMind, and CJ Zheng, Gemini software engineer, in a blog post. "Live Avatar brings these capabilities to enterprise agents. By processing visual and audio inputs simultaneously, it generates enriching conversations for a more comprehensive experience." Five years ago, DeepMind warned about various risks associated with the use of large language models, among them human interaction harms like anthropomorphising AI models – the attribution of human characteristics to software-based interactions. Anthropomorphism here applies broadly to visual human imitation as well as human-like speech and text. "Anthropomorphizing [language models] may inflate users' estimates of the conversational agent’s competencies," noted the more than 20 authors of DeepMind's paper [PDF], "Ethical and social risks of harm from Language Models." "For example, users may falsely infer that a conversational agent that appears human-like in language also displays other human-like characteristics, such as holding a coherent identity over time, or being capable of empathy, perspective-taking, and rational reasoning. As a result, they may place undue confidence, trust, or expectations in these agents." That was in 2021. In August 2025, the risks posed by humanizing AI reached the courtroom when the parents of Adam Raine sued OpenAI, alleging [PDF] that ChatGPT contributed to their son's suicide. Among other things, the complaint lays blame on "anthropomorphic mannerisms calibrated to convey human-like empathy." By December 2025, Google boffins were still grappling with the need to understand the impact of humanizing AI. In a paper [PDF] titled "How Tech Workers Contend with Hazards of Humanlikeness in Generative AI," Mark Díaz, Renee Shelby, Eric Corbett, and Andrew Smart from Google Research reported that tech workers participating in focus groups expressed similar reservations to those cited in the company's prior work. "Tech workers expressed significant concern that humanlikeness fosters a false sense of reliability and trust shaping their concerns as both users and developers, linking to fluid, natural language and tone, which can obscure errors," they wrote. So even as Google AI researchers highlight the need to understand how the "humanlikeness" of AI relates to perceived hazards and responsible AI development, Google Gemini Enterprise is moving straight to deployment. Live Avatars can be drawn from a pre-made library of characters or a custom avatar can be deployed if enterprise allow-listing is enabled. Noting how Live Avatar can make tool calls and fetch data in the background, as well as converse in 97 languages, making it suitable for scenarios like handling hotel guest check-ins, Chang and Zheng insist that Google created these Gemini-generated personas with trust and safety in mind. "We built Live Avatar with strict safeguards designed to respect identity, and keep AI-generated content transparent," they said. ®

那是 2021 年。到了 2025 年 8 月,将 AI 拟人化带来的风险走进了法庭:亚当·雷恩(Adam Raine)的父母起诉了 OpenAI,指控(PDF)ChatGPT 导致了他们儿子的自杀。诉状将矛头指向了“旨在传达人类般同理心的拟人化举止”等因素。

Google has augmented its live dialogue model Gemini 3.8 Live with a feature called "Live Avatar" that's capable of conjuring animated characters and realistic human simulacra to lip-sync Gemini's machine-generated speech. "Conversation is inherently multimodal: we listen, look, speak, and use facial expressions to communicate," explained Shuo-yiin Chang, research scientist at Google DeepMind, and CJ Zheng, Gemini software engineer, in a blog post. "Live Avatar brings these capabilities to enterprise agents. By processing visual and audio inputs simultaneously, it generates enriching conversations for a more comprehensive experience." Five years ago, DeepMind warned about various risks associated with the use of large language models, among them human interaction harms like anthropomorphising AI models – the attribution of human characteristics to software-based interactions. Anthropomorphism here applies broadly to visual human imitation as well as human-like speech and text. "Anthropomorphizing [language models] may inflate users' estimates of the conversational agent’s competencies," noted the more than 20 authors of DeepMind's paper [PDF], "Ethical and social risks of harm from Language Models." "For example, users may falsely infer that a conversational agent that appears human-like in language also displays other human-like characteristics, such as holding a coherent identity over time, or being capable of empathy, perspective-taking, and rational reasoning. As a result, they may place undue confidence, trust, or expectations in these agents." That was in 2021. In August 2025, the risks posed by humanizing AI reached the courtroom when the parents of Adam Raine sued OpenAI, alleging [PDF] that ChatGPT contributed to their son's suicide. Among other things, the complaint lays blame on "anthropomorphic mannerisms calibrated to convey human-like empathy." By December 2025, Google boffins were still grappling with the need to understand the impact of humanizing AI. In a paper [PDF] titled "How Tech Workers Contend with Hazards of Humanlikeness in Generative AI," Mark Díaz, Renee Shelby, Eric Corbett, and Andrew Smart from Google Research reported that tech workers participating in focus groups expressed similar reservations to those cited in the company's prior work. "Tech workers expressed significant concern that humanlikeness fosters a false sense of reliability and trust shaping their concerns as both users and developers, linking to fluid, natural language and tone, which can obscure errors," they wrote. So even as Google AI researchers highlight the need to understand how the "humanlikeness" of AI relates to perceived hazards and responsible AI development, Google Gemini Enterprise is moving straight to deployment. Live Avatars can be drawn from a pre-made library of characters or a custom avatar can be deployed if enterprise allow-listing is enabled. Noting how Live Avatar can make tool calls and fetch data in the background, as well as converse in 97 languages, making it suitable for scenarios like handling hotel guest check-ins, Chang and Zheng insist that Google created these Gemini-generated personas with trust and safety in mind. "We built Live Avatar with strict safeguards designed to respect identity, and keep AI-generated content transparent," they said. ®

直到 2025 年 12 月,谷歌的专家们仍在努力理解并应对将 AI 拟人化所带来的影响。

Google has augmented its live dialogue model Gemini 3.8 Live with a feature called "Live Avatar" that's capable of conjuring animated characters and realistic human simulacra to lip-sync Gemini's machine-generated speech. "Conversation is inherently multimodal: we listen, look, speak, and use facial expressions to communicate," explained Shuo-yiin Chang, research scientist at Google DeepMind, and CJ Zheng, Gemini software engineer, in a blog post. "Live Avatar brings these capabilities to enterprise agents. By processing visual and audio inputs simultaneously, it generates enriching conversations for a more comprehensive experience." Five years ago, DeepMind warned about various risks associated with the use of large language models, among them human interaction harms like anthropomorphising AI models – the attribution of human characteristics to software-based interactions. Anthropomorphism here applies broadly to visual human imitation as well as human-like speech and text. "Anthropomorphizing [language models] may inflate users' estimates of the conversational agent’s competencies," noted the more than 20 authors of DeepMind's paper [PDF], "Ethical and social risks of harm from Language Models." "For example, users may falsely infer that a conversational agent that appears human-like in language also displays other human-like characteristics, such as holding a coherent identity over time, or being capable of empathy, perspective-taking, and rational reasoning. As a result, they may place undue confidence, trust, or expectations in these agents." That was in 2021. In August 2025, the risks posed by humanizing AI reached the courtroom when the parents of Adam Raine sued OpenAI, alleging [PDF] that ChatGPT contributed to their son's suicide. Among other things, the complaint lays blame on "anthropomorphic mannerisms calibrated to convey human-like empathy." By December 2025, Google boffins were still grappling with the need to understand the impact of humanizing AI. In a paper [PDF] titled "How Tech Workers Contend with Hazards of Humanlikeness in Generative AI," Mark Díaz, Renee Shelby, Eric Corbett, and Andrew Smart from Google Research reported that tech workers participating in focus groups expressed similar reservations to those cited in the company's prior work. "Tech workers expressed significant concern that humanlikeness fosters a false sense of reliability and trust shaping their concerns as both users and developers, linking to fluid, natural language and tone, which can obscure errors," they wrote. So even as Google AI researchers highlight the need to understand how the "humanlikeness" of AI relates to perceived hazards and responsible AI development, Google Gemini Enterprise is moving straight to deployment. Live Avatars can be drawn from a pre-made library of characters or a custom avatar can be deployed if enterprise allow-listing is enabled. Noting how Live Avatar can make tool calls and fetch data in the background, as well as converse in 97 languages, making it suitable for scenarios like handling hotel guest check-ins, Chang and Zheng insist that Google created these Gemini-generated personas with trust and safety in mind. "We built Live Avatar with strict safeguards designed to respect identity, and keep AI-generated content transparent," they said. ®

在一篇名为《科技工作者如何应对生成式人工智能中的“类人化”风险》的论文(PDF格式)中,来自谷歌研究院的马克·迪亚兹(Mark Díaz)、蕾妮·谢尔比(Renee Shelby)、埃里克·科贝特(Eric Corbett)和安德鲁·斯马特(Andrew Smart)指出,参与焦点小组讨论的科技工作者表达了与该公司先前研究结果相似的担忧。他们写道:“科技工作者普遍担心,人工智能的‘类人化’特征会给人一种虚假的可靠性和信任感;这种特性既会影响用户的使用体验,也会影响开发者的决策过程。由于人工智能使用自然语言进行交流,这些‘类人化’特征可能会掩盖其中存在的错误。”尽管谷歌的研究人员强调需要深入理解人工智能的‘类人化’特性与潜在风险之间的关系,并推动负责任的AI开发,但谷歌的Gemini Enterprise产品仍直接进入了实际应用阶段。该产品支持使用预先制作好的角色库来创建虚拟化身,或者在企业允许的情况下创建自定义虚拟化身。张(Chang)和郑(Zheng)指出,Live Avatar能够在后台执行工具调用、获取数据,并支持97种语言的交流,因此非常适合用于酒店客人入住等场景。他们强调:“我们开发Live Avatar时采用了严格的保护措施,以确保用户的隐私和数据安全,并确保人工智能生成的内容具有透明度。”

Google has augmented its live dialogue model Gemini 3.8 Live with a feature called "Live Avatar" that's capable of conjuring animated characters and realistic human simulacra to lip-sync Gemini's machine-generated speech. "Conversation is inherently multimodal: we listen, look, speak, and use facial expressions to communicate," explained Shuo-yiin Chang, research scientist at Google DeepMind, and CJ Zheng, Gemini software engineer, in a blog post. "Live Avatar brings these capabilities to enterprise agents. By processing visual and audio inputs simultaneously, it generates enriching conversations for a more comprehensive experience." Five years ago, DeepMind warned about various risks associated with the use of large language models, among them human interaction harms like anthropomorphising AI models – the attribution of human characteristics to software-based interactions. Anthropomorphism here applies broadly to visual human imitation as well as human-like speech and text. "Anthropomorphizing [language models] may inflate users' estimates of the conversational agent’s competencies," noted the more than 20 authors of DeepMind's paper [PDF], "Ethical and social risks of harm from Language Models." "For example, users may falsely infer that a conversational agent that appears human-like in language also displays other human-like characteristics, such as holding a coherent identity over time, or being capable of empathy, perspective-taking, and rational reasoning. As a result, they may place undue confidence, trust, or expectations in these agents." That was in 2021. In August 2025, the risks posed by humanizing AI reached the courtroom when the parents of Adam Raine sued OpenAI, alleging [PDF] that ChatGPT contributed to their son's suicide. Among other things, the complaint lays blame on "anthropomorphic mannerisms calibrated to convey human-like empathy." By December 2025, Google boffins were still grappling with the need to understand the impact of humanizing AI. In a paper [PDF] titled "How Tech Workers Contend with Hazards of Humanlikeness in Generative AI," Mark Díaz, Renee Shelby, Eric Corbett, and Andrew Smart from Google Research reported that tech workers participating in focus groups expressed similar reservations to those cited in the company's prior work. "Tech workers expressed significant concern that humanlikeness fosters a false sense of reliability and trust shaping their concerns as both users and developers, linking to fluid, natural language and tone, which can obscure errors," they wrote. So even as Google AI researchers highlight the need to understand how the "humanlikeness" of AI relates to perceived hazards and responsible AI development, Google Gemini Enterprise is moving straight to deployment. Live Avatars can be drawn from a pre-made library of characters or a custom avatar can be deployed if enterprise allow-listing is enabled. Noting how Live Avatar can make tool calls and fetch data in the background, as well as converse in 97 languages, making it suitable for scenarios like handling hotel guest check-ins, Chang and Zheng insist that Google created these Gemini-generated personas with trust and safety in mind. "We built Live Avatar with strict safeguards designed to respect identity, and keep AI-generated content transparent," they said. ®