ResearchStudio-Idea:一個基於機器學習會議成果的證據驅動研究構想技能套件
ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes
July 5, 2026
作者: Qihao Zhao, Yangyu Huang, Yalun Dai, Lingao Xiao, Jianjun Gao, Xin Zhang, Wenshan Wu, Scarlett Li, Yang He, Yan Lu, Yap Kim Hui
cs.AI
摘要
大型語言模型讓研究構想的發想變得越來越容易,但有效的構想發展不僅僅是生成候選方向。研究人員必須將問題紮根於現有文獻,找出關鍵瓶頸,區別於現有解決方案,並在投入實作前評估風險。我們提出 ResearchStudio-Idea,作為一個可重複使用的技能套件,專門針對研究構想的起點里程。該套件包含 Paper-Search(一個獨立的多來源文獻搜尋技能)、Scoop-Check(一個獨立的既有成果碰撞檢查器,用於評估新穎性主張),以及 IdeaSpark(一個端到端的技能,結合了證據根基、模式引導生成、碰撞檢索、審計與構想卡片輸出於單一工作流程)。IdeaSpark 是基於 2021 至 2025 年間從 ICLR、ICML 與 NeurIPS 收集的 1,947 篇機器學習會議論文所建構,包含口頭報告論文、一個單獨追蹤的高引用子集,以及被拒稿的投稿。透過分析這些結果,我們歸納出 31 個反覆出現的構想子模式,並整合為 15 個可重複使用的構想模式。每個模式以結構化卡片的形式呈現,內容包括研究背景、瓶頸類型、差異化策略、支持的先例與常見失敗模式。給定一個研究問題與一組證據素材,IdeaSpark 會評估證據的完備性、重建周圍的研究背景、找出未解決的瓶頸、選擇相關模式、具體化一個候選方向、檢索可能衝突的既有研究,並執行以結果為導向的審計。這個工作流程將可重複使用的構想模式轉化為可追蹤的研究提案。透過盲審自動評判的評估結果顯示,IdeaSpark 持續產出比無技能及通用技能基準更強的研究提案,同時維持具競爭力的新穎性。
English
Large language models have made research ideation increasingly accessible, yet effective idea development requires more than generating candidate directions. Researchers must ground a problem in current literature, identify meaningful bottlenecks, differentiate from existing solutions, and evaluate risks before committing to implementation. We present ResearchStudio-Idea as a reusable skill suite for this first mile of research ideation. The suite includes Paper-Search, a standalone multi-source literature search skill; Scoop-Check, a standalone prior-art collision checker for novelty claims; and IdeaSpark, the end-to-end skill that composes evidence grounding, pattern-guided generation, collision retrieval, audit, and idea-card rendering into one workflow. IdeaSpark is constructed from a corpus of 1,947 machine learning conference papers collected from ICLR, ICML, and NeurIPS between 2021 and 2025, including Oral papers, a separately tracked high-citation subset, and rejected submissions. Analysis of these outcomes reveals 31 recurring ideation sub-patterns, consolidated into 15 reusable ideation patterns. Each pattern is operationalized as a structured card containing research contexts, bottleneck types, differentiation strategies, supporting precedents, and common failure modes. Given a research problem and an evidence bundle, IdeaSpark evaluates evidence readiness, reconstructs the surrounding research context, identifies unresolved bottlenecks, selects relevant patterns, instantiates one candidate direction, retrieves potentially conflicting prior work, and performs outcome-informed auditing. This workflow transforms reusable ideation patterns into traceable research proposals. Blind automated-judge evaluations show that IdeaSpark consistently produces stronger research proposals than no-skill and generic-skill baselines while maintaining competitive novelty.