ChatPaper.aiChatPaper

通向通用科学AI的路径:科学图像的多模态理解

A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images

August 14, 2026
作者: Jennifer D'Souza, Fahad Ahmed, Cecilia Andrea Bustamante Andrade, Lina Frolova, Poorani Gnanasambandan, Dilshad Hussain, Muhammad Uzair Khan, Nkembeng Kevin Nkengfoa, Paul Praveen J., Fabio Priante, Sjoerd Franciscus van der Werf, Thomas Frederik Jan van Roeden
cs.AI

摘要

科学图表编码了关键的实验证据,然而数字图书馆和多模态人工智能系统仍难以检索和解读这些图表。ALD/E-ImageMiner基准和ICDAR 2026原子层沉积/刻蚀科学图表信息提取竞赛提供了来自205篇出版物的1,951幅图表,经专家标注涵盖分类、数据表抽取、摘要生成和视觉问答等任务。在这些配套论文中,我们提出了关于该基准如何指导未来科学图像挑战的前瞻性视角。我们考察了其任务如何从视觉和定量读取到领域驱动的推理和证据论证来探测能力,以及基于布鲁姆分类法的问题设计如何支持更深层次的科学理解。我们提出将"基于图像的科学概念理解"作为长期基准目标,未来方向包括更广泛的领域和图表类型、上下文和跨文档综合、假设评估、来源追溯、不确定性、反事实依据以及开放式多模态研究。这一视角将ICDAR 2026竞赛与更广泛的机器可操作科学视觉知识和可验证多模态科学人工智能议程联系起来。
English
Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret. The ALD/E-ImageMiner benchmark and ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching Scientific Figures provide 1,951 figures from 205 publications, expert-annotated for classification, data table extraction, summarization, and visual question answering. In these companion proceedings, we present a forward-looking perspective on how the benchmark can guide future scientific-image challenges. We examine how its tasks probe capabilities from visual and quantitative reading to domain-grounded reasoning and evidential justification, and how Bloom-informed question design can support deeper scientific understanding. We propose "scientific conceptual understanding from images" as a long-term benchmark objective, with future directions including broader domains and figure types, contextual and cross-document synthesis, hypothesis evaluation, provenance, uncertainty, counterfactual grounding, and open-ended multimodal research. This perspective connects the ICDAR 2026 challenge to a broader agenda for machine-actionable scientific visual knowledge and verifiable multimodal scientific AI.