レトリックはいかにしてAIレビュアーを報酬ハックし得るのか?——AIベースのピアレビューにおける修辞的感受性の解明
How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review
August 10, 2026
著者: Ming Li, Chenguang Wang, Xirui Li, Xinyue Zeng, Dianqi Li, Peng Shi, Dawei Zhou, Tianyi Zhou
cs.AI
要旨
大規模言語モデル(LLM)が科学的評価にますます参加する中で、我々は報酬ハッキングの潜在的な一形態、すなわち、報告された科学的内容が保持されている場合に修辞的選択がAIレビュー判断をどのように形成するか、またその効果が評価条件によってどのように異なるかを調査する。120件の匿名化されたICLR 2026投稿論文から作成した4,200件の全文論文原稿からなる統制コーパスを構築した。2つのLLMリライターが6つの修辞的次元を正反対の方向に変換し、5つのLLMレビュアーが、標準および厳格なプロトコルの下で、得られた原稿を評価する。さらに、共同書き換え、再帰的書き換え、レビュアー誘導型書き換えも検証する。本研究の結果は、修辞的感度が一律ではなく構造化されていることを示している。エビデンスのフレーミングと新規性スタンスは、総合評価において最大の正負のコントラストを生み出し、スコープのフレーミングはより弱い第2層を形成する。残りの次元は、より小さい、または安定性の低い効果を持つ。この階層構造は、人間が評価した品質レベル全体で持続するが、スコアの変動はAIレビュアーの元のスコアに強く依存する。低いスコアは上昇傾向にあり、高いスコアは下降傾向にあり、方向性のあるコントラストは中間域で最も明確になる。より精巧なワークフローは、確実に大きな利得をもたらすわけではない。共同書き換えはリライター依存性が強く、レビュアーによる誘導は誘導なしの2回目のパスを一貫して上回るわけではなく、書き換えの繰り返しは設定に依存した収穫逓減をもたらす。すべての条件を通じて、リライターは対立するバリアント間の分離を主に決定し、レビュアーはそのスコア効果の大きさと符号を決定する。厳格なレビューは、修辞的感度を一貫して変えることなく、平均総合評価(OA)を1.36ポイント低下させる。これらの知見は、修辞的提示がAIによる科学的レビューに影響を与える状況を特定し、科学的記述における内容を保持した変動に対して頑健な評価システムを動機づける。
English
As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions. We construct a controlled corpus of 4,200 full-paper manuscripts derived from 120 anonymized ICLR 2026 submissions. Two LLM rewriters transform six rhetorical dimensions in opposing directions, and five LLM reviewers evaluate the resulting manuscripts under standard and strict protocols. We also test joint, recursive, and reviewer-guided rewriting. Our results show that rhetorical sensitivity is structured rather than uniform. Evidence framing and novelty stance produce the largest positive-negative contrasts in overall assessment, with scope framing forming a weaker second tier; the remaining dimensions have smaller or less stable effects. This hierarchy persists across human-assessed quality levels, but score movement depends strongly on the AI reviewer's original score: lower scores tend to rise, higher scores tend to fall, and directional contrasts are clearest in the middle ranges. More elaborate workflows do not reliably yield larger gains. Joint rewriting is strongly rewriter-dependent, reviewer guidance does not consistently outperform an unguided second pass, and repeated rewriting yields diminishing, configuration-dependent returns. Across conditions, the rewriter primarily determines the separation between opposing variants, whereas the reviewer determines the magnitude and sign of their score effects. Strict review lowers mean OA by 1.36 points without consistently changing rhetorical sensitivity. These findings identify when rhetorical presentation influences AI scientific review and motivate evaluation systems robust to content-preserving variation in scientific writing.