ChatPaper.aiChatPaper

通用科學人工智慧的發展路徑:科學影像的多模態理解

A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images

August 14, 2026
作者: Jennifer D'Souza, Fahad Ahmed, Cecilia Andrea Bustamante Andrade, Lina Frolova, Poorani Gnanasambandan, Dilshad Hussain, Muhammad Uzair Khan, Nkembeng Kevin Nkengfoa, Paul Praveen J., Fabio Priante, Sjoerd Franciscus van der Werf, Thomas Frederik Jan van Roeden
cs.AI

摘要

科學圖形與表格編碼了關鍵的實驗證據,但對於數位圖書館和多模態人工智慧系統而言,它們仍然難以檢索和解釋。ALD/E-ImageMiner 基準測試與 ICDAR 2026 原子層沉積/蝕刻科學圖形資訊抽取競賽提供了來自 205 篇出版物的 1,951 個圖形,並由專家進行註釋,涵蓋分類、資料表格抽取、摘要和視覺問答。在這些配套論文中,我們提出了前瞻性的視角,探討該基準如何引導未來的科學圖形挑戰。我們檢視其任務如何探測從視覺和量化閱讀到領域基礎推理和證據證明的能力,以及受 Bloom 分類法啟發的問題設計如何支持更深入的科學理解。我們提出「從圖像中進行科學概念理解」作為長期基準目標,未來方向包括更廣泛的領域和圖形類型、上下文與跨文件綜合、假設評估、來源追溯、不確定性、反事實基礎,以及開放式多模態研究。此視角將 ICDAR 2026 挑戰與更廣泛的、可機器操作的科學視覺知識和可驗證的多模態科學人工智慧議程聯繫起來。
English
Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret. The ALD/E-ImageMiner benchmark and ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching Scientific Figures provide 1,951 figures from 205 publications, expert-annotated for classification, data table extraction, summarization, and visual question answering. In these companion proceedings, we present a forward-looking perspective on how the benchmark can guide future scientific-image challenges. We examine how its tasks probe capabilities from visual and quantitative reading to domain-grounded reasoning and evidential justification, and how Bloom-informed question design can support deeper scientific understanding. We propose "scientific conceptual understanding from images" as a long-term benchmark objective, with future directions including broader domains and figure types, contextual and cross-document synthesis, hypothesis evaluation, provenance, uncertainty, counterfactual grounding, and open-ended multimodal research. This perspective connects the ICDAR 2026 challenge to a broader agenda for machine-actionable scientific visual knowledge and verifiable multimodal scientific AI.