ChatPaper.aiChatPaper

修辞如何欺骗AI审稿人以获奖励?——剖析基于AI的同行评审中的修辞敏感性

How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review

August 10, 2026
作者: Ming Li, Chenguang Wang, Xirui Li, Xinyue Zeng, Dianqi Li, Peng Shi, Dawei Zhou, Tianyi Zhou
cs.AI

摘要

随着大语言模型日益参与科学评估,我们研究了一种潜在的奖励作弊形式:在报告的科学内容保持不变的情况下,修辞选择如何影响AI评审判断,以及这些效应在不同评估条件下如何变化。我们构建了一个受控语料库,包含来源于120篇匿名ICLR 2026投稿的4,200篇完整论文手稿。两个LLM改写器沿六个修辞维度向相反方向改写,五个LLM评审员在标准和严格协议下评估由此产生的手稿。我们还测试了联合改写、递归改写和评审员引导的改写。我们的结果表明,修辞敏感性是结构化而非均匀的。证据框架和新颖性立场在总体评估中产生最大的正负对比,范围框架构成较弱的第二梯队;其余维度的影响较小或不太稳定。这一层级结构在不同的人工评定质量水平间保持一致,但分数变动强烈依赖于AI评审员的原始分数:较低分数趋于上升,较高分数趋于下降,方向性对比在中段范围内最为清晰。更精细的工作流程并不可靠地带来更大的收益。联合改写强烈依赖于改写器,评审员引导并不始终优于无引导的第二轮改写,而重复改写产生递减且依赖配置的收益。在各种条件下,改写器主要决定对立变体之间的分离程度,而评审员决定其分数效应的幅度和方向。严格评审使总体评估均值降低1.36分,但不持续改变修辞敏感性。这些发现揭示了修辞呈现何时影响AI科学评审,并促使构建对科学写作中内容保持性变化具有鲁棒性的评估系统。
English
As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions. We construct a controlled corpus of 4,200 full-paper manuscripts derived from 120 anonymized ICLR 2026 submissions. Two LLM rewriters transform six rhetorical dimensions in opposing directions, and five LLM reviewers evaluate the resulting manuscripts under standard and strict protocols. We also test joint, recursive, and reviewer-guided rewriting. Our results show that rhetorical sensitivity is structured rather than uniform. Evidence framing and novelty stance produce the largest positive-negative contrasts in overall assessment, with scope framing forming a weaker second tier; the remaining dimensions have smaller or less stable effects. This hierarchy persists across human-assessed quality levels, but score movement depends strongly on the AI reviewer's original score: lower scores tend to rise, higher scores tend to fall, and directional contrasts are clearest in the middle ranges. More elaborate workflows do not reliably yield larger gains. Joint rewriting is strongly rewriter-dependent, reviewer guidance does not consistently outperform an unguided second pass, and repeated rewriting yields diminishing, configuration-dependent returns. Across conditions, the rewriter primarily determines the separation between opposing variants, whereas the reviewer determines the magnitude and sign of their score effects. Strict review lowers mean OA by 1.36 points without consistently changing rhetorical sensitivity. These findings identify when rhetorical presentation influences AI scientific review and motivate evaluation systems robust to content-preserving variation in scientific writing.