arXiv: 2607.14713
多智能体辩论能否提升对研究论文的AI反馈?
Does Multi-Agent Debate Improve AI Feedback on Research Papers?
July 16, 2026
作者: Tomas Havranek, Zuzana Irsova
econ.GNecon.GNcs.CLcs.MA
摘要
在经济学领域,至少针对荟萃分析而言,答案很可能是否定的。在一项预先注册、身份隐匿的论文内部实验中,44项荟萃分析的作者根据对改进自身论文的有用性,对三份AI报告进行了排序:一份由前沿模型单次生成,另外两份则来自我们搭建且预期胜出的多智能体辩论工具。所有报告均遵循统一的长度和模板。作者们更偏好单次生成的结果,相较mad-research高出0.66个排名点(95%置信区间0.32至1.00),相较paper-workshop高出0.57个排名点(0.16至0.95),尽管paper-workshop消耗的令牌数约为前者的三十倍。那些记得自己期刊审稿报告的作者通常将其排在首位,且从未置于末位;而在另一项独立测试中,三位AI评审几乎总是将真实的期刊审稿报告排在最后。在三份AI报告中,Gemini(其模型家族未参与任何报告撰写的评审)若处于作者位置会将paper-workshop排在第一,从而逆转了单次生成的偏好。这一逆转警示我们,不应以AI评审替代作者本人。我们测量的是对已完成论文的感知有用性;至于AI是否应担任论文审稿人,则是一个独立的问题。
English
Probably not, at least for meta-analyses in economics. In a pre-registered, identity-masked, within-paper experiment, the authors of 44 meta-analyses ranked three AI reports on their own paper by usefulness for improving it: a single pass by a frontier model against two multi-agent debate tools we built and expected to win. All reports were held to a common length and template. The authors preferred the single pass, by 0.66 rank points over mad-research (95% CI 0.32 to 1.00) and 0.57 over paper-workshop (0.16 to 0.95), though paper-workshop spent roughly thirty times the tokens. Authors who recalled their journal referee report usually placed it first and never last; in a separate exercise, three AI judges almost always placed the real journal referee report last. Among the three AI reports, Gemini (the judge whose model family wrote none of the reports) would have ranked paper-workshop first in the authors' place, reversing the single-pass preference. The reversal warns against substituting an AI judge for the author. We measure perceived usefulness for finished papers; whether AI should referee papers is a separate question.