ChatPaper.aiChatPaper

ICDAR 2026 原子層堆積/エッチング(ALD/E)科学図からの情報抽出コンペティション

ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures

July 29, 2026
著者: Fahad Ahmed, Sören Auer, Jennifer D'Souza
cs.AI

要旨

マルチモーダルAIを用いた科学図の理解と推論には、視覚的知覚とドメイン固有の推論の統合が必要であり、研究論文のテキストにはしばしば明示されない有意義な知識を抽出する。コミュニティ主導のコンペティションを伴うSci-ImageMinerベンチマークデータセットは、専門家によるアノテーションが施された包括的なデータセットを4つのエンドツーエンドの相補的タスクにわたってキュレーションすることにより、従来の科学系コンペティションの水準を引き上げるものである。本コンペティションには、2026年1月9日から2026年4月8日までの期間に68名のアクティブな参加者が集まり、1,263件の公開・非公開の提出が行われた。我々の結果は、最先端のマルチモーダルモデルが分類タスクと要約タスクでは良好な性能を発揮する一方、データ抽出と科学推論、特に視覚的質問応答においては困難に直面することを示している。これらの知見は、主要な限界を明らかにし、ドメイン認識型マルチモーダルAIシステムの改善に向けた課題と機会を浮き彫りにする。全体として、Sci-ImageMinerベンチマークとコンペティションは、科学図の理解と推論に関する研究を前進させるための厳格なプラットフォームを確立し、挑戦的で複雑な研究領域における最先端アプローチの可能性を示すものである。
English
Scientific figure comprehension and reasoning using multimodal AI requires integrating visual perception with domain-specific reasoning to extract meaningful knowledge, often not presented in the text of a research publication. The Sci-ImageMiner benchmark dataset, accompanied by a community-driven competition, raises the bar over prior scientific competitions by curating a comprehensive, expert-annotated dataset across four end-to-end complementary tasks. The competition attracted 68 active participants and 1,263 public/private submissions from 9th January 2026 to 8th April 2026. Our results show that state-of-the-art multimodal models perform well on classification and summarization tasks but struggle with data extraction and scientific reasoning, particularly in visual question-answering. These findings reveal key limitations and highlight challenges and opportunities for improving domain-aware multimodal AI systems. Overall, the Sci-ImageMiner benchmark and competition establish a rigorous platform for advancing research in scientific figure comprehension and reasoning and demonstrate the potential of state-of-the-art approaches for a challenging and complex research area.