arXiv: 2607.14713

多智能體辯論是否改善了對研究論文的AI反饋?

Does Multi-Agent Debate Improve AI Feedback on Research Papers?

July 16, 2026
作者: Tomas Havranek, Zuzana Irsova
econ.GNecon.GNcs.CLcs.MA

摘要

很可能不會,至少在經濟學的元分析中不會。在一項預先註冊、身分隱藏、且為論文內實驗的研究中,44篇元分析的作者針對自身論文,根據改進實用性對三份AI報告進行排序:包括單次前沿模型生成的報告,以及我們建構並預期會勝出的兩種多代理辯論工具。所有報告均維持相同長度與模板。作者偏愛單次前沿模型,其排名比mad-research高出0.66個排名點(95%信賴區間0.32至1.00),比paper-workshop高出0.57個排名點(0.16至0.95),儘管paper-workshop消耗的token量約為前者的三十倍。回憶起自己期刊審稿報告的作者,通常將其列為第一且從不列為最後;在另一項獨立測試中,三位AI評審幾乎總是將真實的期刊審稿報告列為最後。在三份AI報告中,Gemini(其模型家族未撰寫任何報告的評審)若站在作者立場,會將paper-workshop排在第一,逆轉了對單次模型的偏好。此逆轉結果警示我們不應以AI評審取代作者。我們衡量的是針對已完成論文的感知實用性;至於AI是否應擔任論文審稿人,則是另一問題。
English
Probably not, at least for meta-analyses in economics. In a pre-registered, identity-masked, within-paper experiment, the authors of 44 meta-analyses ranked three AI reports on their own paper by usefulness for improving it: a single pass by a frontier model against two multi-agent debate tools we built and expected to win. All reports were held to a common length and template. The authors preferred the single pass, by 0.66 rank points over mad-research (95% CI 0.32 to 1.00) and 0.57 over paper-workshop (0.16 to 0.95), though paper-workshop spent roughly thirty times the tokens. Authors who recalled their journal referee report usually placed it first and never last; in a separate exercise, three AI judges almost always placed the real journal referee report last. Among the three AI reports, Gemini (the judge whose model family wrote none of the reports) would have ranked paper-workshop first in the authors' place, reversing the single-pass preference. The reversal warns against substituting an AI judge for the author. We measure perceived usefulness for finished papers; whether AI should referee papers is a separate question.
PDFJuly 19, 2026