ChatPaper.aiChatPaper

ARC: オープンエンドな実世界インタラクションにおける公平な相対的優位性比較

ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction

August 13, 2026
著者: Yongqi Tong, Tan Li Hui Faith, Choy Zhen Wen Marcus, Zhou Jin, Kewei Fu, Jiang-Ming Yang, Jianshe Li, Xin Zhang
cs.AI

要旨

オープンエンドな実世界対話では、複数の妥当な行動が許容される。エージェントは、直接回答したり、明確化を求めたり、進捗を報告したり、行動前に確認したりできる。この柔軟性は、グループベースのRLの背後にある中核的な仮定を破る。すなわち、グループ内で比較されるロールアウトが行動的に比較可能であることはもはや保証されない。その結果、報酬モデルの対話スタイルに対する選好が相対的アドバンテージを歪め、最適化を文脈に適した行動ではなく報酬モデルが好む行動へと誘導し得る。我々はこれを報酬公平性問題として定式化し、ARC(Advantage Regularization via Conditioning:条件付けによるアドバンテージ正則化)を提案する。これは、方略条件付きロールアウトのグループ化とハイブリッド報酬およびエントロピー正則化を組み合わせることにより、より公平な相対比較を回復する訓練レシピである。我々は、提案する\inter{}においてARCを検証する。\inter{}は、ユーザーに見えるコミュニケーションを潜在的な推論やツール使用から分離する、応答性が高く、誘導可能で、実行認識型のユーザーとエージェントの対話の新しいパラダイムである。\inter{}はまた、教師あり学習およびRL訓練のための方略アノテーション付き訓練コーパスである\inter-86Kを構築するためのアノテーションおよび蒸留パイプラインも提供する。実験的には、ARCは中核となるτ/τ²ツール使用ベンチマークの性能を大幅に向上させる。一方、\inter{}は、思考型ベースラインと比較してtime-to-first-tokenを4.91秒から1.27秒に短縮する。これらの結果は総合的に、オープンエンドな対話型学習における中心的なボトルネックは、エージェントがどのように報酬を得るかだけではなく、そもそもその行動が公平に比較されるかどうかにあることを示唆している。ARCの実装と\inter-86Kの訓練データは公開される予定である。
English
Open-ended real-world interaction admits multiple valid behaviors: an agent may answer directly, ask for clarification, provide progress updates, or confirm before acting. This flexibility breaks a core assumption behind group-based RL: rollouts compared within a group are no longer guaranteed to be behaviorally comparable. As a result, reward-model preferences over interaction style can distort relative advantages and steer optimization toward reward-preferred behaviors rather than context-appropriate ones. We formalize this as a reward fairness problem and propose ARC (Advantage Regularization via Conditioning), a training recipe that restores fairer relative comparison through strategy-conditioned rollout grouping, together with hybrid rewards and entropy regularization. We study ARC in our proposed \inter, a novel paradigm for responsive, steerable, and execution-aware user-agent interaction that decouples user-visible communication from latent reasoning and tool use. \inter\ also provides the annotation and distillation pipeline for constructing \inter-86K, our strategy-annotated training corpus for supervised and RL training. Empirically, ARC substantially strengthens the core τ/τ^2 tool-use benchmarks, while \inter\ reduces time-to-first-token from 4.91s to 1.27s relative to a think-style baseline. Together, these results suggest that a central bottleneck in open-ended interactive learning is not only how agents are rewarded, but whether their behaviors are compared fairly in the first place. The ARC implementation and \inter-86K training data will be released.