ARC:開放式真實世界互動中的公平相對優勢比較
ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction
August 13, 2026
作者: Yongqi Tong, Tan Li Hui Faith, Choy Zhen Wen Marcus, Zhou Jin, Kewei Fu, Jiang-Ming Yang, Jianshe Li, Xin Zhang
cs.AI
摘要
開放式真實世界互動允許多種有效行為:智能體可以直接回答、請求澄清、提供進度更新,或在行動前確認。這種靈活性打破了基於群組的強化學習背後的一項核心假設:群組內相互比較的 rollout 不再保證在行為上具有可比性。因此,獎勵模型對互動風格的偏好可能扭曲相對優勢,並將優化導向獎勵偏好的行為,而非符合情境的行為。我們將此形式化為獎勵公平性問題,並提出 ARC(Advantage Regularization via Conditioning,基於條件化的優勢正則化),這是一種透過策略條件化的 rollout 分組,並結合混合獎勵與熵正則化,以恢復更公平相對比較的訓練方案。我們在所提出的 \inter 中研究 ARC,\inter 是一個用於回應式、可引導且具執行感知能力的使用者-智能體互動的新範式,可將使用者可見的通訊與潛在推理及工具使用解耦。\inter 也提供了建構 \inter-86K 的標註與蒸餾流程;\inter-86K 是我們用於監督式與強化學習訓練的策略標註訓練語料庫。實驗上,ARC 大幅增強了核心的 τ/τ^2 工具使用基準,而 \inter 相對於思考式基線,將 time-to-first-token 從 4.91 秒降至 1.27 秒。這些結果共同表明,開放式互動學習中的一個核心瓶頸不僅在於智能體如何被獎勵,更在於其行為是否首先被公平地比較。ARC 的實作程式碼與 \inter-86K 訓練資料將予以釋出。
English
Open-ended real-world interaction admits multiple valid behaviors: an agent may answer directly, ask for clarification, provide progress updates, or confirm before acting. This flexibility breaks a core assumption behind group-based RL: rollouts compared within a group are no longer guaranteed to be behaviorally comparable. As a result, reward-model preferences over interaction style can distort relative advantages and steer optimization toward reward-preferred behaviors rather than context-appropriate ones. We formalize this as a reward fairness problem and propose ARC (Advantage Regularization via Conditioning), a training recipe that restores fairer relative comparison through strategy-conditioned rollout grouping, together with hybrid rewards and entropy regularization. We study ARC in our proposed \inter, a novel paradigm for responsive, steerable, and execution-aware user-agent interaction that decouples user-visible communication from latent reasoning and tool use. \inter\ also provides the annotation and distillation pipeline for constructing \inter-86K, our strategy-annotated training corpus for supervised and RL training. Empirically, ARC substantially strengthens the core τ/τ^2 tool-use benchmarks, while \inter\ reduces time-to-first-token from 4.91s to 1.27s relative to a think-style baseline. Together, these results suggest that a central bottleneck in open-ended interactive learning is not only how agents are rewarded, but whether their behaviors are compared fairly in the first place. The ARC implementation and \inter-86K training data will be released.