ChatPaper.aiChatPaper

通具身智能体:以人为中心的智能体AI新范式

ComBodied Agents: a New Paradigm of Human-Centric Agentic AI

August 11, 2026
作者: Qianggang Ding, Xingyao Wang, Rui Feng, Zhibin Wang, Feixiang Wang, Kelong Mao, Hao Sun, Zhiyao Luo, Jiankai Tang, Lei Li, Jiadong Guo, Minheng Ni, Weicong Lin, Chenxi Yang, Hongxiang Gao, Zhenghua Chen, Yang Bai, Min Wu, Jun Cheng, Huazhu Fu, Dacheng Tao, Bang Liu
cs.AI

摘要

当老年人漏服一剂药物后,软件智能体可以发送另一条提醒,具身智能体可以递送药物。然而,两者都无法解释该人是忘记了、感到困惑、出现了副作用,还是故意拒绝服药,也无法判断何种支持是适当的。这揭示了Agentic AI中的一个结构性缺口:数字智能体主要转换软件状态,而具身智能体转换物理状态;两者都未将人的动态状态及其能动性作为建模、干预和评估的首要对象。我们提出具身融合智能体(Combodied Agents),这是一种以人为中心的新范式,它通过软件工具、传感器、可穿戴设备、机器人和人类服务作为行动渠道而非最终目标,来感知、建模、预测并长期支持个体状态轨迹。我们将个人助手、健康智能体、AI伴侣和自适应人机系统中的碎片化能力统一为一个闭环:基于事件的多模态感知重建有意义的个人事件;纵向、可修正的记忆提供时间上下文;个人世界模型(Personal World Models)在备选决策和干预下估计未来个人状态与结果;可采纳的干预策略在同意、不确定性、安全性、可逆性和用户控制的前提下选择适度支持。来自人和环境的反馈持续更新该闭环。该框架无需构建完备的人类数字孪生(Human Digital Twin),而是采用目的受限、不确定性感知、用户可修正的表征。我们按人类状态目标、关系情境和智能体角色来组织设计空间,并提出以场景为中心的评估、能动性保持指标、基准测试要求、边缘原生个人模型和治理方向。具身融合智能体将Agentic AI从外部任务完成转向持续的人类福祉。
English
After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither explains whether the person forgot, is confused, has side effects, or deliberately refused, nor what support is appropriate. This reveals a structural gap in Agentic AI: Digital Agents primarily transform software states, while Embodied Agents transform physical states; neither makes a person's evolving state and agency the primary object of modeling, intervention, and evaluation. We introduce Combodied Agents, a human-centered paradigm that perceives, models, predicts, and supports individual human-state trajectories over time, using software tools, sensors, wearables, robots, and human services as action channels rather than end goals. We unify fragmented capabilities across personal assistants, health agents, AI companions, and adaptive human--AI systems into a closed loop: event-based multimodal perception reconstructs meaningful personal events; longitudinal, correctable memory provides temporal context; Personal World Models estimate future personal states and outcomes under alternative decisions and interventions; and an admissible intervention policy selects proportionate support under consent, uncertainty, safety, reversibility, and user control. Feedback from the person and environment updates the loop. Rather than requiring an exhaustive Human Digital Twin, the framework uses purpose-bounded, uncertainty-aware, user-correctable representations. We organize the design space by human-state targets, relational contexts, and agent roles, and propose scenario-centered evaluation, agency-preservation metrics, benchmark requirements, edge-native personal models, and governance directions. Combodied Agents shift Agentic AI from external task completion toward sustained human benefit.