수사학은 어떻게 AI 리뷰어를 보상 해킹할 수 있는가? AI 기반 동료 평가에서 수사적 민감성 분석
How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review
August 10, 2026
저자: Ming Li, Chenguang Wang, Xirui Li, Xinyue Zeng, Dianqi Li, Peng Shi, Dawei Zhou, Tianyi Zhou
cs.AI
초록
대규모 언어 모델(LLM)이 과학적 평가에 참여하는 비중이 점점 커짐에 따라, 본 연구는 보상 해킹(reward hacking)의 잠재적 형태를 조사한다. 구체적으로, 보고된 과학적 내용이 보존될 때 수사적 선택이 AI 리뷰 판단을 어떻게 형성하는지, 그리고 이러한 효과가 평가 조건에 따라 어떻게 달라지는지를 분석한다. 120개의 익명화된 ICLR 2026 제출 논문에서 파생된 4,200개의 전체 논문 원고로 구성된 통제 코퍼스를 구축하였다. 두 개의 LLM 재작성 모델이 여섯 가지 수사적 차원을 반대 방향으로 변형하고, 다섯 개의 LLM 평가자가 표준 및 엄격 프로토콜 하에서 결과 원고를 평가한다. 또한 결합, 재귀, 평가자 안내 재작성 방식도 테스트하였다. 실험 결과, 수사적 민감성은 균일하지 않고 구조화되어 있음이 확인되었다. 증거 프레이밍과 참신성 입장이 종합 평가(OA)에서 가장 큰 긍정-부정 대비를 보였으며, 범위 프레이밍은 이보다 약한 두 번째 계층을 형성하였다. 나머지 차원들은 더 작거나 덜 안정적인 효과를 보였다. 이러한 위계는 인간 평가 품질 수준에 걸쳐 유지되었지만, 점수 변동은 AI 평가자의 원래 점수에 크게 의존하였다. 낮은 점수는 상승하는 경향이 있고, 높은 점수는 하락하는 경향이 있으며, 방향성 대비는 중간 점수대에서 가장 뚜렷하게 나타났다. 더 정교한 워크플로우가 확실히 더 큰 이득을 가져오지는 않았다. 결합 재작성은 사용된 재작성 모델에 크게 의존하였고, 평가자 안내 방식은 안내 없는 두 번째 재작성보다 일관되게 우수하지 않았으며, 반복적 재작성은 구성에 따라 달라지는 체감 효과를 보였다. 모든 조건에 걸쳐, 재작성 모델은 반대 변형 간의 분리를 주로 결정한 반면, 평가자는 점수 효과의 크기와 방향을 결정하였다. 엄격 리뷰는 평균 OA를 1.36점 낮추었지만, 수사적 민감성에는 일관된 변화를 주지 않았다. 본 연구는 수사적 표현이 AI 과학 리뷰에 영향을 미치는 조건을 식별함으로써, 과학적 글쓰기의 내용 보존적 변형에 강건한 평가 시스템의 필요성을 제기한다.
English
As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions. We construct a controlled corpus of 4,200 full-paper manuscripts derived from 120 anonymized ICLR 2026 submissions. Two LLM rewriters transform six rhetorical dimensions in opposing directions, and five LLM reviewers evaluate the resulting manuscripts under standard and strict protocols. We also test joint, recursive, and reviewer-guided rewriting. Our results show that rhetorical sensitivity is structured rather than uniform. Evidence framing and novelty stance produce the largest positive-negative contrasts in overall assessment, with scope framing forming a weaker second tier; the remaining dimensions have smaller or less stable effects. This hierarchy persists across human-assessed quality levels, but score movement depends strongly on the AI reviewer's original score: lower scores tend to rise, higher scores tend to fall, and directional contrasts are clearest in the middle ranges. More elaborate workflows do not reliably yield larger gains. Joint rewriting is strongly rewriter-dependent, reviewer guidance does not consistently outperform an unguided second pass, and repeated rewriting yields diminishing, configuration-dependent returns. Across conditions, the rewriter primarily determines the separation between opposing variants, whereas the reviewer determines the magnitude and sign of their score effects. Strict review lowers mean OA by 1.36 points without consistently changing rhetorical sensitivity. These findings identify when rhetorical presentation influences AI scientific review and motivate evaluation systems robust to content-preserving variation in scientific writing.