ChatPaper.aiChatPaper

ICDAR 2026 原子層沉積/蝕刻(ALD/E)科學圖表資訊抽取競賽

ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures

July 29, 2026
作者: Fahad Ahmed, Sören Auer, Jennifer D'Souza
cs.AI

摘要

利用多模態AI進行科學圖表理解與推理,需要將視覺感知與領域特定推理相結合,以提取有意義的知識,而這些知識往往不會呈現在研究論文的文字中。Sci-ImageMiner基準資料集,搭配社群驅動的競賽,透過策劃一個涵蓋四項互補端到端任務的全面性、專家註解資料集,提升了先前科學競賽的標準。該競賽吸引了68位活躍參與者,並在2026年1月9日至2026年4月8日期間收到1,263份公開/私人提交。我們的結果顯示,最先進的多模態模型在分類與摘要任務上表現良好,但在資料提取與科學推理方面,特別是在視覺問答中,則面臨困難。這些發現揭示了關鍵限制,並凸顯了改善領域感知多模態AI系統的挑戰與機會。整體而言,Sci-ImageMiner基準與競賽為推進科學圖表理解與推理研究建立了一個嚴謹的平台,同時展現了最先進方法在具挑戰性且複雜的研究領域中的潛力。
English
Scientific figure comprehension and reasoning using multimodal AI requires integrating visual perception with domain-specific reasoning to extract meaningful knowledge, often not presented in the text of a research publication. The Sci-ImageMiner benchmark dataset, accompanied by a community-driven competition, raises the bar over prior scientific competitions by curating a comprehensive, expert-annotated dataset across four end-to-end complementary tasks. The competition attracted 68 active participants and 1,263 public/private submissions from 9th January 2026 to 8th April 2026. Our results show that state-of-the-art multimodal models perform well on classification and summarization tasks but struggle with data extraction and scientific reasoning, particularly in visual question-answering. These findings reveal key limitations and highlight challenges and opportunities for improving domain-aware multimodal AI systems. Overall, the Sci-ImageMiner benchmark and competition establish a rigorous platform for advancing research in scientific figure comprehension and reasoning and demonstrate the potential of state-of-the-art approaches for a challenging and complex research area.