ICDAR 2026 原子层沉积/刻蚀(ALD/E)科学图表信息提取竞赛
ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures
July 29, 2026
作者: Fahad Ahmed, Sören Auer, Jennifer D'Souza
cs.AI
摘要
使用多模态AI进行科学图形理解与推理,需要将视觉感知与领域特定推理相结合,以提取有意义的知识,而这些知识往往不会呈现在研究论文的文本中。Sci-ImageMiner基准数据集配合社区驱动的竞赛,通过构建一个涵盖四项端到端互补任务的综合性专家标注数据集,将先前科学竞赛的标准提升到了新的高度。该竞赛吸引了68名活跃参与者,在2026年1月9日至2026年4月8日期间共收到1,263份公开/私有提交。我们的结果表明,最先进的多模态模型在分类和摘要任务上表现良好,但在数据提取和科学推理方面存在困难,尤其是在视觉问答任务中。这些发现揭示了关键局限性,并凸显了改进领域感知多模态AI系统的挑战与机遇。总体而言,Sci-ImageMiner基准与竞赛为推进科学图形理解与推理研究建立了一个严谨的平台,同时展示了最先进方法在这一具有挑战性和复杂性的研究领域中的潜力。
English
Scientific figure comprehension and reasoning using multimodal AI requires integrating visual perception with domain-specific reasoning to extract meaningful knowledge, often not presented in the text of a research publication. The Sci-ImageMiner benchmark dataset, accompanied by a community-driven competition, raises the bar over prior scientific competitions by curating a comprehensive, expert-annotated dataset across four end-to-end complementary tasks. The competition attracted 68 active participants and 1,263 public/private submissions from 9th January 2026 to 8th April 2026. Our results show that state-of-the-art multimodal models perform well on classification and summarization tasks but struggle with data extraction and scientific reasoning, particularly in visual question-answering. These findings reveal key limitations and highlight challenges and opportunities for improving domain-aware multimodal AI systems. Overall, the Sci-ImageMiner benchmark and competition establish a rigorous platform for advancing research in scientific figure comprehension and reasoning and demonstrate the potential of state-of-the-art approaches for a challenging and complex research area.