ChatPaper.aiChatPaper

ARC:开放式真实世界交互中的公平相对优势比较

ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction

August 13, 2026
作者: Yongqi Tong, Tan Li Hui Faith, Choy Zhen Wen Marcus, Zhou Jin, Kewei Fu, Jiang-Ming Yang, Jianshe Li, Xin Zhang
cs.AI

摘要

开放式真实世界交互允许多种有效行为:智能体可以直接回答、请求澄清、提供进度更新或在行动前进行确认。这种灵活性打破了基于分组的强化学习(group-based RL)背后的一个核心假设:在同一分组内进行比较的轨迹(rollouts)不再保证行为上具有可比性。因此,奖励模型对交互风格的偏好可能扭曲相对优势(relative advantages),并将优化引向奖励偏好行为而非符合上下文的行为。我们将此形式化为奖励公平性问题,并提出ARC(Advantage Regularization via Conditioning,条件化优势正则化),一种训练方案,通过策略条件化的轨迹分组(strategy-conditioned rollout grouping),结合混合奖励(hybrid rewards)与熵正则化(entropy regularization),恢复更公平的相对比较。我们在所提出的\inter框架中研究ARC,该框架是一种新颖的响应式、可引导且执行感知的用户-智能体交互范式,将用户可见的通信与潜在推理和工具使用解耦。\inter还提供了构建\inter-86K所需的标注与蒸馏流水线,\inter-86K是我们用于监督训练和强化学习训练的策略标注训练语料库。实验上,ARC显著增强了核心τ/τ²工具使用基准的性能,而\inter将首令牌生成时间(time-to-first-token)从4.91秒降至1.27秒(相较于思维式基线)。这些结果共同表明,开放式交互学习中的核心瓶颈不仅在于智能体如何获得奖励,更在于其行为是否首先得到公平的比较。ARC的实现和\inter-86K训练数据将予以发布。
English
Open-ended real-world interaction admits multiple valid behaviors: an agent may answer directly, ask for clarification, provide progress updates, or confirm before acting. This flexibility breaks a core assumption behind group-based RL: rollouts compared within a group are no longer guaranteed to be behaviorally comparable. As a result, reward-model preferences over interaction style can distort relative advantages and steer optimization toward reward-preferred behaviors rather than context-appropriate ones. We formalize this as a reward fairness problem and propose ARC (Advantage Regularization via Conditioning), a training recipe that restores fairer relative comparison through strategy-conditioned rollout grouping, together with hybrid rewards and entropy regularization. We study ARC in our proposed \inter, a novel paradigm for responsive, steerable, and execution-aware user-agent interaction that decouples user-visible communication from latent reasoning and tool use. \inter\ also provides the annotation and distillation pipeline for constructing \inter-86K, our strategy-annotated training corpus for supervised and RL training. Empirically, ARC substantially strengthens the core τ/τ^2 tool-use benchmarks, while \inter\ reduces time-to-first-token from 4.91s to 1.27s relative to a think-style baseline. Together, these results suggest that a central bottleneck in open-ended interactive learning is not only how agents are rewarded, but whether their behaviors are compared fairly in the first place. The ARC implementation and \inter-86K training data will be released.