ChatPaper.aiChatPaper

問題就是問題:邁向可擴展的數學發現

The Problem Is the Problem: Towards Scalable Mathematical Discovery

August 17, 2026
作者: Zeyu Zheng, Shengtong Zhang, Jeremy Avigad, Prasad Tetali, Sean Welleck
cs.AI

摘要

AI系統正日益具備對數學研究的貢獻能力。在研究實踐中,前沿模型的推理能力是有限的資源,而專家數學審查則受到更嚴格的限制。因此,妥善配置這些稀缺資源對於提升AI輔助數學發現的效率至關重要。在當前大多數AI數學研究工作流程中,人力投入集中在開始與結束兩個階段,即選擇合適的研究問題,以及後續審查所產出的成果。這兩個階段正成為研究級數學的瓶頸。我們通過提出一種新的人機協作發現典範來解決這些問題。人類輸入不再是事先選定的單一問題,而是專家具有興趣與專業知識的研究方向。系統隨後在廣泛的文獻語料庫中搜索該方向下的候選問題。受搜索與推薦系統啟發,我們構建了尋找、嘗試與推薦(Find, Attempt, and Recommend, FAR)系統,這是一個從文獻到審查的級聯流程,可自動化搜索合適的問題,並將人類注意力集中於已通過多階段篩選的成果上。在一項組合學試點研究中,該流程從5,245篇組合學論文出發,提取出6,453個候選猜想或開放問題,並將其篩選至4,717個表面良置且仍開放的猜想。隨後進行的推理與自動分流階段浮現出598個潛在解答,並選出77個項目供作者團隊審查。其中,我們發現了許多有趣的成果,包括關於Davies–Jenssen–Perkins–Roberts、Erdős–Straus、Ikenmeyer–Pak–Panova及Lund–Saraf–Wolf猜想與問題的結果。這些成果證明了這種新型人機協作模式在數學發現中的有效性。
English
AI systems are increasingly capable of contributing to mathematical research. In research practice, frontier-model reasoning is a limited resource, and expert mathematical review is even more sharply constrained. Allocating these scarce resources well is therefore central to making AI-assisted mathematical discovery efficient. In most current AI-for-math workflows, human effort is concentrated at the beginning and end, in selecting suitable research problems and later reviewing the resulting artifacts. These two stages are becoming bottlenecks for research-level mathematics. We address them by proposing a new human-AI discovery paradigm. The human input is no longer a single problem selected in advance, but a research direction in which the experts have interest and expertise. The system then searches a broad literature corpus for candidate problems in that direction. Inspired by search and recommender systems, we build Find, Attempt, and Recommend (FAR), a literature-to-review cascade that automates the search for suitable problems and focuses human attention on artifacts that have passed several stages of filtering. In a combinatorics pilot, the pipeline starts from 5,245 combinatorics papers, recovers 6,453 candidate conjectures or open problems, and filters them to 4,717 apparently well-posed and still-open conjectures. Subsequent reasoning and automated triage stages surface 598 potential resolutions and select 77 items for author-team review. Among them, we identify many interesting discoveries, including results on conjectures and questions of Davies--Jenssen--Perkins--Roberts, Erdős--Straus, Ikenmeyer--Pak--Panova, and Lund--Saraf--Wolf. These results demonstrate the effectiveness of this new mode of human-AI collaboration for mathematical discovery.