Naar het automatiseren van wetenschappelijke peer review met Google's Paper Assistant Tool
Towards Automating Scientific Review with Google's Paper Assistant Tool
June 26, 2026
Auteurs: Rajesh Jayaram, Drew Tyler, David Woodruff, Corinna Cortes, Yossi Matias, Vahab Mirrokni, Vincent Cohen-Addad
cs.AI
Samenvatting
Kunstmatige intelligentie drijft een revolutie in wetenschappelijke ontdekkingen aan, waarbij alles van hypothesegeneratie tot het bewijzen van wiskundige stellingen wordt versneld. Deze snelle versnelling creëert echter een systemische uitdaging: traditionele menselijke peer review kan niet opschalen om gelijke tred te houden met de instroom van AI-ondersteunde wetenschap. Om deze spanning op te lossen, moeten we uiteindelijk ook AI inzetten om het verificatie- en reviewproces zelf te versnellen. Om de discussie rond deze overgang te kaderen, stellen we een taxonomie voor die bestaat uit vier progressieve niveaus van AI-menselijke samenwerking bij wetenschappelijke evaluatie, en bespreken we verschillende afwegingen die bij elk niveau komen kijken.
Als stap in de richting van deze toekomst introduceren we de Paper Assistant Tool (PAT), een agentisch AI-raamwerk ontworpen voor diepgaande wetenschappelijke review en verificatie. PAT neemt volledige wetenschappelijke manuscripten op en produceert een uitgebreide evaluatie, waarbij theoretische resultaten worden gecontroleerd, experimenten worden gevalideerd, verbeteringen worden voorgesteld en mogelijke fouten worden geïdentificeerd. Door gebruik te maken van inferentieschalingstechnieken kan PAT diepere problemen identificeren dan een enkele modelaanroep alleen, wat een verbetering van 34% oplevert ten opzichte van zero-shot recall voor wiskundige fouten in de SPOT-benchmark. Proefimplementaties van PAT als een pre-submissietool voor auteurs op twee grote Computer Science-conferenties -- STOC en ICML -- tonen het vermogen aan om kritieke fouten te identificeren en substantiële verbeteringen aan onderzoekspapers voor te stellen. Door fouten vroegtijdig op te sporen, verlicht PAT de cognitieve last die op recensenten rust, terwijl hun controle over de uitkomsten van het reviewproces behouden blijft.
English
Artificial intelligence is driving a revolution in scientific discovery, accelerating everything from hypothesis generation to mathematical theorem proving. However, this rapid acceleration is creating a systemic challenge: traditional human peer review cannot scale to match the influx of AI-assisted science. Ultimately, to resolve this tension, we must also deploy AI to accelerate the verification and review process itself. To frame the discussion around this transition, we propose a taxonomy consisting of four progressive levels of AI-human collaboration in scientific evaluation, and discuss various trade-offs involved with each.
As a step toward this future, we introduce the Paper Assistant Tool (PAT), an agentic AI framework built for deep scientific review and verification. PAT ingests full scientific manuscripts and produces a comprehensive evaluation, checking theoretical results, validating experiments, suggesting improvements, and identifying potential flaws. By utilizing inference scaling techniques, PAT is able to identify deeper issues than a single model call alone, achieving a 34% improvement over zero-shot recall on mathematical errors in the SPOT benchmark. Pilot deployments of PAT as a pre-submission tool for authors at two major Computer Science conferences -- STOC and ICML -- demonstrate its ability to identify critical errors and suggest substantive improvements to research papers. By catching errors early, PAT eases the cognitive burden placed on referees, while preserving their control over the outcomes of the review process.