汎用科学AIへの道筋:科学画像のマルチモーダル理解
A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images
August 14, 2026
著者: Jennifer D'Souza, Fahad Ahmed, Cecilia Andrea Bustamante Andrade, Lina Frolova, Poorani Gnanasambandan, Dilshad Hussain, Muhammad Uzair Khan, Nkembeng Kevin Nkengfoa, Paul Praveen J., Fabio Priante, Sjoerd Franciscus van der Werf, Thomas Frederik Jan van Roeden
cs.AI
要旨
科学図表は本質的な実験的証拠をコード化する一方で、デジタルライブラリやマルチモーダルAIシステムにとって、その検索と解釈は依然として困難である。ALD/E-ImageMinerベンチマークとICDAR 2026 原子層堆積/エッチング科学図表からの情報抽出コンペティションは、205件の出版論文から得られた1,951点の図表に対し、分類、データ表抽出、要約、視覚的質問応答のための専門家アノテーションを提供する。本併設プロシーディングスにおいて、我々はこのベンチマークが将来の科学画像チャレンジをどのように導き得るかについて、将来を見据えた展望を提示する。本ベンチマークのタスクが、視覚的・定量的読み取りから領域に根ざした推論や証拠に基づく正当化に至る能力をどのように探求するか、またブルーム分類学に基づく問題設計がより深い科学的理解をいかに支援できるかを考察する。我々は「画像からの科学的概念理解」を長期的なベンチマーク目標として提案し、将来の方向性として、より広範な領域と図表タイプ、文脈的・文書横断的統合、仮説評価、来歴、不確実性、反事実的根拠付け、オープンエンドなマルチモーダル研究を含める。本展望は、ICDAR 2026チャレンジを、機械処理可能な科学的視覚知識と検証可能なマルチモーダル科学AIのためのより広範なアジェンダに接続するものである。
English
Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret. The ALD/E-ImageMiner benchmark and ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching Scientific Figures provide 1,951 figures from 205 publications, expert-annotated for classification, data table extraction, summarization, and visual question answering. In these companion proceedings, we present a forward-looking perspective on how the benchmark can guide future scientific-image challenges. We examine how its tasks probe capabilities from visual and quantitative reading to domain-grounded reasoning and evidential justification, and how Bloom-informed question design can support deeper scientific understanding. We propose "scientific conceptual understanding from images" as a long-term benchmark objective, with future directions including broader domains and figure types, contextual and cross-document synthesis, hypothesis evaluation, provenance, uncertainty, counterfactual grounding, and open-ended multimodal research. This perspective connects the ICDAR 2026 challenge to a broader agenda for machine-actionable scientific visual knowledge and verifiable multimodal scientific AI.