問題が問題である:スケーラブルな数学的発見に向けて
The Problem Is the Problem: Towards Scalable Mathematical Discovery
August 17, 2026
著者: Zeyu Zheng, Shengtong Zhang, Jeremy Avigad, Prasad Tetali, Sean Welleck
cs.AI
要旨
AIシステムは、数学研究に貢献する能力をますます高めている。研究実践において、フロンティアモデルによる推論は限られたリソースであり、専門家による数学レビューはさらに深刻な制約を受けている。したがって、これらの希少なリソースを適切に配分することが、AI支援による数学的発見を効率的にするうえで中心的な課題となる。現在のAIを活用した数学研究ワークフローのほとんどでは、人間の労力は最初と最後、すなわち適切な研究問題の選定と、その後に生成される成果物のレビューに集中している。この2つの段階は、研究レベルの数学にとってボトルネックになりつつある。我々はこの課題に対処するため、新しい人間とAIの協働による発見パラダイムを提案する。人間の入力はもはや事前に選定された単一の問題ではなく、専門家が関心と専門性を持つ研究の方向性である。システムはその方向性に沿った候補問題を、広範な文献コーパスから検索する。検索システムとレコメンダーシステムに着想を得て、我々はFind、Attempt、Recommend(FAR)を構築する。これは文献からレビューへと至るカスケードであり、適切な問題の探索を自動化し、複数のフィルタリング段階を通過した成果物に人間の注意を集中させる。組合せ論のパイロット研究では、このパイプラインは5,245本の組合せ論論文から開始し、6,453件の候補となる予想または未解決問題を抽出し、それらをフィルタリングして、見かけ上適切に定式化され現在も未解決の予想4,717件に絞り込む。その後の推論と自動トリアージの段階により、598件の潜在的な解決が表面化し、著者チームによるレビューのために77項目が選定される。その中から、我々は多くの興味深い発見を特定する。これには、Davies–Jenssen–Perkins–Roberts、Erdős–Straus、Ikenmeyer–Pak–Panova、Lund–Saraf–Wolfによる予想や問題に関する結果が含まれる。これらの結果は、数学的発見のための人間とAIの協働におけるこの新しい形態の有効性を示している。
English
AI systems are increasingly capable of contributing to mathematical research. In research practice, frontier-model reasoning is a limited resource, and expert mathematical review is even more sharply constrained. Allocating these scarce resources well is therefore central to making AI-assisted mathematical discovery efficient. In most current AI-for-math workflows, human effort is concentrated at the beginning and end, in selecting suitable research problems and later reviewing the resulting artifacts. These two stages are becoming bottlenecks for research-level mathematics. We address them by proposing a new human-AI discovery paradigm. The human input is no longer a single problem selected in advance, but a research direction in which the experts have interest and expertise. The system then searches a broad literature corpus for candidate problems in that direction. Inspired by search and recommender systems, we build Find, Attempt, and Recommend (FAR), a literature-to-review cascade that automates the search for suitable problems and focuses human attention on artifacts that have passed several stages of filtering. In a combinatorics pilot, the pipeline starts from 5,245 combinatorics papers, recovers 6,453 candidate conjectures or open problems, and filters them to 4,717 apparently well-posed and still-open conjectures. Subsequent reasoning and automated triage stages surface 598 potential resolutions and select 77 items for author-team review. Among them, we identify many interesting discoveries, including results on conjectures and questions of Davies--Jenssen--Perkins--Roberts, Erdős--Straus, Ikenmeyer--Pak--Panova, and Lund--Saraf--Wolf. These results demonstrate the effectiveness of this new mode of human-AI collaboration for mathematical discovery.