문제는 문제다: 확장 가능한 수학적 발견을 위하여
The Problem Is the Problem: Towards Scalable Mathematical Discovery
August 17, 2026
저자: Zeyu Zheng, Shengtong Zhang, Jeremy Avigad, Prasad Tetali, Sean Welleck
cs.AI
초록
AI 시스템은 점점 더 수학 연구에 기여할 수 있게 되었다. 연구 현장에서 최첨단 모델 추론은 제한된 자원이며, 전문가의 수학적 검토는 더욱 심각하게 제약되어 있다. 따라서 AI 지원 수학적 발견의 효율성을 높이기 위해서는 이러한 희소 자원을 잘 배분하는 것이 핵심이다. 대부분의 현재 AI-수학 워크플로우에서 인간의 노력은 시작과 끝, 즉 적절한 연구 문제를 선정하고 이후에 생성된 산출물을 검토하는 데 집중된다. 이 두 단계는 연구 수준의 수학에서 병목 현상이 되고 있다. 우리는 새로운 인간-AI 발견 패러다임을 제안함으로써 이를 해결한다. 인간의 입력은 더 이상 사전에 선정된 단일 문제가 아니라, 전문가들이 관심과 전문성을 가진 연구 방향이다. 이후 시스템은 그 방향의 후보 문제를 찾기 위해 광범위한 문헌 말뭉치를 탐색한다. 검색 및 추천 시스템에서 영감을 얻어, 우리는 Find, Attempt, Recommend(FAR)를 구축한다. FAR은 적절한 문제 탐색을 자동화하고 여러 단계의 필터링을 통과한 산출물에 인간의 주의를 집중시키는 문헌-검토 캐스케이드이다. 조합론 파일럿 연구에서, 이 파이프라인은 5,245편의 조합론 논문에서 시작하여 6,453개의 후보 추측 또는 미해결 문제를 추출하고, 이를 4,717개의 겉보기에 잘 정식화되었고 아직 미해결인 추측으로 필터링한다. 이후의 추론 및 자동 분류 단계는 598개의 잠재적 해결을 표면화하고 77개 항목을 저자 팀 검토용으로 선별한다. 그중에서 우리는 Davies–Jenssen–Perkins–Roberts, Erdős–Straus, Ikenmeyer–Pak–Panova, Lund–Saraf–Wolf의 추측 및 질문에 대한 결과를 포함하여 많은 흥미로운 발견을 확인했다. 이러한 결과는 수학적 발견을 위한 인간-AI 협업의 새로운 방식이 효과적임을 입증한다.
English
AI systems are increasingly capable of contributing to mathematical research. In research practice, frontier-model reasoning is a limited resource, and expert mathematical review is even more sharply constrained. Allocating these scarce resources well is therefore central to making AI-assisted mathematical discovery efficient. In most current AI-for-math workflows, human effort is concentrated at the beginning and end, in selecting suitable research problems and later reviewing the resulting artifacts. These two stages are becoming bottlenecks for research-level mathematics. We address them by proposing a new human-AI discovery paradigm. The human input is no longer a single problem selected in advance, but a research direction in which the experts have interest and expertise. The system then searches a broad literature corpus for candidate problems in that direction. Inspired by search and recommender systems, we build Find, Attempt, and Recommend (FAR), a literature-to-review cascade that automates the search for suitable problems and focuses human attention on artifacts that have passed several stages of filtering. In a combinatorics pilot, the pipeline starts from 5,245 combinatorics papers, recovers 6,453 candidate conjectures or open problems, and filters them to 4,717 apparently well-posed and still-open conjectures. Subsequent reasoning and automated triage stages surface 598 potential resolutions and select 77 items for author-team review. Among them, we identify many interesting discoveries, including results on conjectures and questions of Davies--Jenssen--Perkins--Roberts, Erdős--Straus, Ikenmeyer--Pak--Panova, and Lund--Saraf--Wolf. These results demonstrate the effectiveness of this new mode of human-AI collaboration for mathematical discovery.