ARC: 개방형 실세계 상호작용에서의 공정한 상대적 우위 비교
ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction
August 13, 2026
저자: Yongqi Tong, Tan Li Hui Faith, Choy Zhen Wen Marcus, Zhou Jin, Kewei Fu, Jiang-Ming Yang, Jianshe Li, Xin Zhang
cs.AI
초록
개방형 실제 상호작용에서는 여러 가지 유효한 행동이 허용된다. 에이전트는 직접 응답하거나, 명확화를 요청하거나, 진행 상황을 제공하거나, 행동하기 전에 확인을 받을 수 있다. 이러한 유연성은 그룹 기반 강화학습(RL)의 핵심 가정을 깨뜨린다. 그룹 내에서 비교되는 롤아웃이 더 이상 행동적으로 비교 가능하다는 보장이 없다. 그 결과, 상호작용 스타일에 대한 보상 모델의 선호는 상대적 이점을 왜곡하고, 최적화를 상황에 적합한 행동보다 보상이 선호하는 행동으로 이끌 수 있다. 우리는 이를 보상 공정성 문제로 공식화하고, 전략 조건화 롤아웃 그룹화, 하이브리드 보상, 엔트로피 정규화를 통해 더 공정한 상대 비교를 복원하는 훈련 기법인 ARC(Advantage Regularization via Conditioning, 조건화를 통한 이점 정규화)를 제안한다. 우리는 제안하는 INTER, 즉 사용자에게 보이는 의사소통과 잠재 추론 및 도구 사용을 분리하는 반응형, 유도 가능형, 실행 인식형 사용자-에이전트 상호작용을 위한 새로운 패러다임에서 ARC를 연구한다. 또한 INTER는 지도 학습 및 RL 훈련을 위한 전략 주석 훈련 말뭉치인 INTER-86K를 구축하기 위한 주석 및 증류 파이프라인을 제공한다. 실험적으로 ARC는 핵심 τ/τ^2 도구 사용 벤치마크를 크게 향상시키며, INTER는 사고형 기준선 대비 첫 토큰 생성 시간을 4.91초에서 1.27초로 줄인다. 종합하면, 이러한 결과는 개방형 상호작용 학습의 핵심 병목이 에이전트가 어떻게 보상을 받는지뿐만 아니라, 그들의 행동이 먼저 공정하게 비교되는지 여부라는 점을 시사한다. ARC 구현과 INTER-86K 훈련 데이터는 공개될 예정이다.
English
Open-ended real-world interaction admits multiple valid behaviors: an agent may answer directly, ask for clarification, provide progress updates, or confirm before acting. This flexibility breaks a core assumption behind group-based RL: rollouts compared within a group are no longer guaranteed to be behaviorally comparable. As a result, reward-model preferences over interaction style can distort relative advantages and steer optimization toward reward-preferred behaviors rather than context-appropriate ones. We formalize this as a reward fairness problem and propose ARC (Advantage Regularization via Conditioning), a training recipe that restores fairer relative comparison through strategy-conditioned rollout grouping, together with hybrid rewards and entropy regularization. We study ARC in our proposed \inter, a novel paradigm for responsive, steerable, and execution-aware user-agent interaction that decouples user-visible communication from latent reasoning and tool use. \inter\ also provides the annotation and distillation pipeline for constructing \inter-86K, our strategy-annotated training corpus for supervised and RL training. Empirically, ARC substantially strengthens the core τ/τ^2 tool-use benchmarks, while \inter\ reduces time-to-first-token from 4.91s to 1.27s relative to a think-style baseline. Together, these results suggest that a central bottleneck in open-ended interactive learning is not only how agents are rewarded, but whether their behaviors are compared fairly in the first place. The ARC implementation and \inter-86K training data will be released.