ChatPaper.aiChatPaper

修辭如何能獎勵劫持AI審稿人?剖析基於AI的同行評審中的修辭敏感性

How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review

August 10, 2026
作者: Ming Li, Chenguang Wang, Xirui Li, Xinyue Zeng, Dianqi Li, Peng Shi, Dawei Zhou, Tianyi Zhou
cs.AI

摘要

隨著大型語言模型日益參與科學評估,我們研究了一種潛在的獎勵駭取形式:在所報告的科學內容保持不變的情況下,修辭選擇如何塑造AI審稿判斷,以及這些效應如何隨評估條件而變化。我們建構了一個受控語料庫,包含4,200篇完整論文手稿,源自120篇匿名化的ICLR 2026投稿。兩個LLM改寫器在相反方向上轉換六個修辭維度,五個LLM審稿人在標準與嚴格協議下評估生成的手稿。我們也測試了聯合、遞迴及審稿人引導的改寫。結果顯示,修辭敏感性具有結構性而非均質性。證據框架與新穎性立場在總體評估中產生最大的正負對比,範圍框架構成較弱的第二層級;其餘維度的效應較小或較不穩定。此層級結構在人類評估的品質等級之間持續存在,但分數變化在很大程度上取決於AI審稿人的原始分數:較低的分數傾向於上升,較高的分數傾向於下降,而方向性對比在中間範圍最為清晰。更複雜的工作流程並不能可靠地帶來更大的收益。聯合改寫強烈依賴於改寫器的選擇,審稿人引導並不持續優於無引導的第二輪改寫,而重複改寫產生遞減且依賴配置的收益。在各種條件下,改寫器主要決定對立變體之間的分離程度,而審稿人則決定其分數效應的大小與符號。嚴格審查將平均總體評估(OA)降低1.36分,但不會持續改變修辭敏感性。這些發現識別了修辭呈現何時影響AI科學審查,並促使建構對科學寫作中保留內容的變異具有魯棒性的評估系統。
English
As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions. We construct a controlled corpus of 4,200 full-paper manuscripts derived from 120 anonymized ICLR 2026 submissions. Two LLM rewriters transform six rhetorical dimensions in opposing directions, and five LLM reviewers evaluate the resulting manuscripts under standard and strict protocols. We also test joint, recursive, and reviewer-guided rewriting. Our results show that rhetorical sensitivity is structured rather than uniform. Evidence framing and novelty stance produce the largest positive-negative contrasts in overall assessment, with scope framing forming a weaker second tier; the remaining dimensions have smaller or less stable effects. This hierarchy persists across human-assessed quality levels, but score movement depends strongly on the AI reviewer's original score: lower scores tend to rise, higher scores tend to fall, and directional contrasts are clearest in the middle ranges. More elaborate workflows do not reliably yield larger gains. Joint rewriting is strongly rewriter-dependent, reviewer guidance does not consistently outperform an unguided second pass, and repeated rewriting yields diminishing, configuration-dependent returns. Across conditions, the rewriter primarily determines the separation between opposing variants, whereas the reviewer determines the magnitude and sign of their score effects. Strict review lowers mean OA by 1.36 points without consistently changing rhetorical sensitivity. These findings identify when rhetorical presentation influences AI scientific review and motivate evaluation systems robust to content-preserving variation in scientific writing.