InternReviewer与InternAdvocate:面向同行评审与反驳中智能体强化学习的客观奖励与评估
InternReviewer & InternAdvocate: Objective Reward and Evaluation for Agentic Reinforcement Learning in Peer Review and Rebuttal
July 21, 2026
作者: Xuerui Su, Liya Guo, Qizhi Pei, Qipeng Guo, Zhongbo Tian, Lijun Wu, Kai Chen, Zun Wang
cs.AI
摘要
生成专业学术内容,如同行评审与反驳意见,需要领域推理与事实依据之间错综复杂的协同配合。本工作提出了一个综合框架,用于开发和评估专业学术智能体——InternReviewer与InternAdvocate。我们首先构建了一个大规模、高质量的学术数据集,并集成了一种高效率的arXiv检索工具,以实现主动的证据收集。为优化这些智能体,我们实施了一种由统一目标指标和奖励系统驱动的智能体强化学习范式。该系统通过采用多维标准,包括基于参考文献锚定的语义对齐、结构合规性,以及一种将引文与实时交互日志交叉核对以消除幻觉的严格验证机制,从而避免了基于主观模型评判的偏差。实验结果表明,在此闭环框架内训练的智能体在推理深度和引文准确性方面均展现出显著提升。
English
Generating professional scholarly content, such as peer reviews and rebuttals, requires an intricate synergy between domain reasoning and factual grounding. This work presents a comprehensive framework for the development and evaluation of specialized scholarly agents, InternReviewer and InternAdvocate. We first establish a large-scale, high-quality scholarly dataset and integrate a high-efficiency arXiv retrieval tool to enable active evidence gathering. To optimize these agents, we implement an agentic Reinforcement Learning (RL) paradigm driven by a unified objective metric and reward system. This system avoids the biases of subjective model-based judging by employing multi-dimensional criteria, including reference-anchored semantic alignment, structural compliance, and a strict verification mechanism that cross-checks citations against real-time interaction logs to eliminate hallucinations. Experimental results demonstrate that agents trained within this closed-loop framework exhibit significant improvements in reasoning depth and citation accuracy.