ResearchStudio-Idea: 머신러닝 컨퍼런스 결과물로부터 도출된 증거 기반의 연구 아이디어 창출 스킬 모음
ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes
July 5, 2026
저자: Qihao Zhao, Yangyu Huang, Yalun Dai, Lingao Xiao, Jianjun Gao, Xin Zhang, Wenshan Wu, Scarlett Li, Yang He, Yan Lu, Yap Kim Hui
cs.AI
초록
대규모 언어 모델은 연구 아이디어 발상에 대한 접근성을 높였지만, 효과적인 아이디어 개발은 단순히 후보 방향을 생성하는 것 이상을 요구한다. 연구자는 문제를 현재 문헌에 근거시키고, 의미 있는 병목 지점을 식별하며, 기존 해결책과의 차별성을 확보하고, 구현에 착수하기 전에 위험을 평가해야 한다. 본 연구에서는 연구 아이디어 발상의 첫 단계를 위한 재사용 가능한 스킬 제품군으로 ResearchStudio-Idea를 제시한다. 이 제품군에는 독립형 다중 소스 문헌 검색 스킬인 Paper-Search, 참신성 주장에 대한 독립형 선행 기술 충돌 검사기인 Scoop-Check, 그리고 증거 근거 부여, 패턴 기반 생성, 충돌 검색, 감사, 아이디어 카드 렌더링을 하나의 워크플로우로 구성하는 종단 간 스킬인 IdeaSpark가 포함된다. IdeaSpark는 2021년부터 2025년 사이 ICLR, ICML, NeurIPS에서 수집된 1,947개의 머신러닝 학회 논문(구두 논문, 별도로 추적된 고인용 하위 집합, 그리고 불합격된 제출물 포함)으로 구성된 말뭉치를 기반으로 구축되었다. 이러한 결과물에 대한 분석을 통해 31개의 반복적인 아이디어 발상 하위 패턴이 발견되었고, 이는 15개의 재사용 가능한 아이디어 발상 패턴으로 통합되었다. 각 패턴은 연구 맥락, 병목 유형, 차별화 전략, 지원 선례, 그리고 일반적인 실패 모드를 포함하는 구조화된 카드로 구현된다. IdeaSpark는 주어진 연구 문제와 증거 묶음을 바탕으로 증거의 준비 상태를 평가하고, 주변 연구 맥락을 재구성하며, 해결되지 않은 병목 지점을 식별하고, 관련 패턴을 선택하며, 하나의 후보 방향을 구체화하고, 잠재적으로 충돌하는 선행 연구를 검색하며, 결과 기반 감사를 수행한다. 이 워크플로우는 재사용 가능한 아이디어 발상 패턴을 추적 가능한 연구 제안서로 변환한다. 블라인드 자동 평가 결과, IdeaSpark는 스킬이 없거나 일반 스킬을 사용한 기준선보다 일관되게 더 강력한 연구 제안서를 생성하면서도 경쟁력 있는 참신성을 유지함을 보여준다.
English
Large language models have made research ideation increasingly accessible, yet effective idea development requires more than generating candidate directions. Researchers must ground a problem in current literature, identify meaningful bottlenecks, differentiate from existing solutions, and evaluate risks before committing to implementation. We present ResearchStudio-Idea as a reusable skill suite for this first mile of research ideation. The suite includes Paper-Search, a standalone multi-source literature search skill; Scoop-Check, a standalone prior-art collision checker for novelty claims; and IdeaSpark, the end-to-end skill that composes evidence grounding, pattern-guided generation, collision retrieval, audit, and idea-card rendering into one workflow. IdeaSpark is constructed from a corpus of 1,947 machine learning conference papers collected from ICLR, ICML, and NeurIPS between 2021 and 2025, including Oral papers, a separately tracked high-citation subset, and rejected submissions. Analysis of these outcomes reveals 31 recurring ideation sub-patterns, consolidated into 15 reusable ideation patterns. Each pattern is operationalized as a structured card containing research contexts, bottleneck types, differentiation strategies, supporting precedents, and common failure modes. Given a research problem and an evidence bundle, IdeaSpark evaluates evidence readiness, reconstructs the surrounding research context, identifies unresolved bottlenecks, selects relevant patterns, instantiates one candidate direction, retrieves potentially conflicting prior work, and performs outcome-informed auditing. This workflow transforms reusable ideation patterns into traceable research proposals. Blind automated-judge evaluations show that IdeaSpark consistently produces stronger research proposals than no-skill and generic-skill baselines while maintaining competitive novelty.