ChatPaper.aiChatPaper

ICDAR 2026 원자층 증착/식각(ALD/E) 과학적 그림 정보 추출 대회

ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures

July 29, 2026
저자: Fahad Ahmed, Sören Auer, Jennifer D'Souza
cs.AI

초록

과학적 도표 이해와 추론을 위한 멀티모달 AI는 시각적 인식과 도메인 특화 추론을 통합하여, 연구 논문의 본문에는 제시되지 않는 의미 있는 지식을 추출해야 한다. Sci-ImageMiner 벤치마크 데이터셋은 커뮤니티 주도 대회와 함께, 네 가지 상호 보완적인 종단 간 작업에 걸친 포괄적이고 전문가 주석 기반 데이터셋을 구축함으로써 기존 과학 경쟁 대회보다 높은 기준을 제시한다. 본 대회는 2026년 1월 9일부터 2026년 4월 8일까지 68명의 활발한 참가자와 1,263건의 공개/비공개 제출물을 유치했다. 연구 결과, 최첨단 멀티모달 모델은 분류 및 요약 작업에서는 우수한 성능을 보였으나, 특히 시각 질의응답에서 데이터 추출과 과학적 추론에 어려움을 겪는 것으로 나타났다. 이러한 발견은 핵심적 한계를 드러내며, 도메인 인지 멀티모달 AI 시스템 개선을 위한 과제와 기회를 강조한다. 전반적으로 Sci-ImageMiner 벤치마크와 대회는 과학적 도표 이해 및 추론 연구를 발전시키기 위한 엄격한 플랫폼을 확립하며, 도전적이고 복잡한 연구 영역에서 최첨단 접근법의 잠재력을 입증한다.
English
Scientific figure comprehension and reasoning using multimodal AI requires integrating visual perception with domain-specific reasoning to extract meaningful knowledge, often not presented in the text of a research publication. The Sci-ImageMiner benchmark dataset, accompanied by a community-driven competition, raises the bar over prior scientific competitions by curating a comprehensive, expert-annotated dataset across four end-to-end complementary tasks. The competition attracted 68 active participants and 1,263 public/private submissions from 9th January 2026 to 8th April 2026. Our results show that state-of-the-art multimodal models perform well on classification and summarization tasks but struggle with data extraction and scientific reasoning, particularly in visual question-answering. These findings reveal key limitations and highlight challenges and opportunities for improving domain-aware multimodal AI systems. Overall, the Sci-ImageMiner benchmark and competition establish a rigorous platform for advancing research in scientific figure comprehension and reasoning and demonstrate the potential of state-of-the-art approaches for a challenging and complex research area.