ChatPaper.aiChatPaper

범용 과학 AI로의 경로: 과학 이미지의 멀티모달 이해

A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images

August 14, 2026
저자: Jennifer D'Souza, Fahad Ahmed, Cecilia Andrea Bustamante Andrade, Lina Frolova, Poorani Gnanasambandan, Dilshad Hussain, Muhammad Uzair Khan, Nkembeng Kevin Nkengfoa, Paul Praveen J., Fabio Priante, Sjoerd Franciscus van der Werf, Thomas Frederik Jan van Roeden
cs.AI

초록

과학적 도표와 표는 필수적인 실험 증거를 담고 있지만, 디지털 도서관과 다중모달 AI 시스템이 이를 검색하고 해석하기는 여전히 어렵다. ALD/E-ImageMiner 벤치마크와 원자층증착/식각 과학 도표의 정보 추출을 위한 ICDAR 2026 대회는 205개 출판물에서 수집한 1,951개의 도표를 제공하며, 분류, 데이터 표 추출, 요약, 시각적 질의응답을 위해 전문가 주석을 수행하였다. 본 동반 논문에서 우리는 이 벤치마크가 향후 과학적 이미지 과제를 어떻게 안내할 수 있는지에 대한 전향적 관점을 제시한다. 우리는 이 과제들이 시각적·정량적 판독에서부터 도메인 기반 추론 및 증거 기반 정당화에 이르는 역량을 어떻게 검증하는지, 그리고 블룸 분류학에 기반한 질문 설계가 더 깊은 과학적 이해를 어떻게 지원할 수 있는지를 살펴본다. 우리는 "이미지로부터의 과학적 개념적 이해"를 장기적 벤치마크 목표로 제안하며, 향후 방향으로 더 광범위한 도메인과 도표 유형, 맥락적 및 문서 간 종합, 가설 평가, 출처 추적, 불확실성, 반사실적 근거, 개방형 다중모달 연구를 포함한다. 본 관점은 ICDAR 2026 과제를 기계가 처리 가능한 과학적 시각 지식과 검증 가능한 다중모달 과학적 AI를 위한 더 광범위한 의제로 연결한다.
English
Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret. The ALD/E-ImageMiner benchmark and ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching Scientific Figures provide 1,951 figures from 205 publications, expert-annotated for classification, data table extraction, summarization, and visual question answering. In these companion proceedings, we present a forward-looking perspective on how the benchmark can guide future scientific-image challenges. We examine how its tasks probe capabilities from visual and quantitative reading to domain-grounded reasoning and evidential justification, and how Bloom-informed question design can support deeper scientific understanding. We propose "scientific conceptual understanding from images" as a long-term benchmark objective, with future directions including broader domains and figure types, contextual and cross-document synthesis, hypothesis evaluation, provenance, uncertainty, counterfactual grounding, and open-ended multimodal research. This perspective connects the ICDAR 2026 challenge to a broader agenda for machine-actionable scientific visual knowledge and verifiable multimodal scientific AI.