InternReviewer & InternAdvocate: 동료 검토 및 반박 과정에서 에이전트 강화 학습을 위한 객관적 보상과 평가
InternReviewer & InternAdvocate: Objective Reward and Evaluation for Agentic Reinforcement Learning in Peer Review and Rebuttal
July 21, 2026
저자: Xuerui Su, Liya Guo, Qizhi Pei, Qipeng Guo, Zhongbo Tian, Lijun Wu, Kai Chen, Zun Wang
cs.AI
초록
동료 심사와 반박과 같은 전문 학술 콘텐츠를 생성하는 것은 도메인 추론과 사실적 근거 간의 정교한 시너지를 요구한다. 본 연구는 전문 학술 에이전트인 InternReviewer와 InternAdvocate의 개발 및 평가를 위한 포괄적인 프레임워크를 제시한다. 먼저 대규모의 고품질 학술 데이터셋을 구축하고 고효율 arXiv 검색 도구를 통합하여 능동적 증거 수집을 가능하게 한다. 이러한 에이전트를 최적화하기 위해, 통합 목표 지표와 보상 체계에 기반한 에이전트 중심 강화 학습(RL) 패러다임을 구현한다. 이 체계는 참고문헌 기반 의미 정합성, 구조 준수, 그리고 실시간 상호작용 로그와 인용을 교차 검증하여 환각을 제거하는 엄격한 검증 메커니즘을 포함한 다차원 기준을 채택함으로써 주관적인 모델 기반 평가의 편향을 회피한다. 실험 결과는 이러한 폐루프 프레임워크 내에서 훈련된 에이전트가 추론의 깊이와 인용 정확성에 있어 유의미한 개선을 보였음을 입증한다.
English
Generating professional scholarly content, such as peer reviews and rebuttals, requires an intricate synergy between domain reasoning and factual grounding. This work presents a comprehensive framework for the development and evaluation of specialized scholarly agents, InternReviewer and InternAdvocate. We first establish a large-scale, high-quality scholarly dataset and integrate a high-efficiency arXiv retrieval tool to enable active evidence gathering. To optimize these agents, we implement an agentic Reinforcement Learning (RL) paradigm driven by a unified objective metric and reward system. This system avoids the biases of subjective model-based judging by employing multi-dimensional criteria, including reference-anchored semantic alignment, structural compliance, and a strict verification mechanism that cross-checks citations against real-time interaction logs to eliminate hallucinations. Experimental results demonstrate that agents trained within this closed-loop framework exhibit significant improvements in reasoning depth and citation accuracy.