arXiv: 2607.14713

マルチエージェント討論は研究論文に対するAIフィードバックを改善するか?

Does Multi-Agent Debate Improve AI Feedback on Research Papers?

July 16, 2026
著者: Tomas Havranek, Zuzana Irsova
econ.GNecon.GNcs.CLcs.MA

要旨

おそらくそうではない。少なくとも経済学におけるメタアナリシスに関しては。事前登録され、身元が隠蔽された同一論文内実験において、44件のメタアナリシスの著者らは、自らの論文を改善するための有用性に基づき、3種類のAIレポートを順位付けした。すなわち、フロンティアモデルによる単回処理、および我々が構築し勝利を期待した2種類のマルチエージェント討論ツールである。全レポートは共通の長さとテンプレートに従った。著者らは単回処理を好み、mad-researchに対して0.66順位ポイント(95%信頼区間0.32~1.00)、paper-workshopに対して0.57順位ポイント(同0.16~0.95)の差をつけた。ただし、paper-workshopはおよそ30倍のトークンを消費した。自身のジャーナル査読レポートを想起した著者は、それをほぼ常に1位に置き、最下位にすることはなかった。一方、別の実験では、3名のAI判定者はほぼ常に実際のジャーナル査読レポートを最下位に置いた。3件のAIレポートの中では、Gemini(いずれのレポートも執筆していないモデルファミリーの判定者)が、著者に代わってpaper-workshopを1位にランク付けし、単回処理の選好を逆転させた。この逆転は、AI判定者を著者の代わりに用いることへの警告となる。本稿では完成論文に対する認識された有用性を測定するものであり、AIが論文を査読すべきか否かは別の問題である。
English
Probably not, at least for meta-analyses in economics. In a pre-registered, identity-masked, within-paper experiment, the authors of 44 meta-analyses ranked three AI reports on their own paper by usefulness for improving it: a single pass by a frontier model against two multi-agent debate tools we built and expected to win. All reports were held to a common length and template. The authors preferred the single pass, by 0.66 rank points over mad-research (95% CI 0.32 to 1.00) and 0.57 over paper-workshop (0.16 to 0.95), though paper-workshop spent roughly thirty times the tokens. Authors who recalled their journal referee report usually placed it first and never last; in a separate exercise, three AI judges almost always placed the real journal referee report last. Among the three AI reports, Gemini (the judge whose model family wrote none of the reports) would have ranked paper-workshop first in the authors' place, reversing the single-pass preference. The reversal warns against substituting an AI judge for the author. We measure perceived usefulness for finished papers; whether AI should referee papers is a separate question.
PDFJuly 19, 2026