CoVA-SFT: 視覚的抽象化の連鎖のための大規模データセット
CoVA-SFT: A Large-Scale Dataset for Chain of Visual Abstractions
August 29, 2026
著者: Tsung-Han Wu, Heekyung Lee, Anya Ji, Haoming Chen, Trevor Darrell, Joseph E. Gonzalez, David M. Chan
cs.AI
要旨
Chain-of-thought(CoT)推論は、大規模言語モデル(LLM)が問題を中間ステップに分解することを可能にし、LLMを劇的に改善してきた。CoTは言語タスクに対して広く有効である一方、テキストのみのCoTはモデルに対して、視覚的な問題を不自然な散文へと系列化することを強いる。視覚入力を処理するためのアーキテクチャ上の解決策は存在するものの、研究コミュニティには、純粋にテキストのみの推論問題を解く際に、モデルが内部の視覚的ワークスペースを構築・維持する方法を教えるための、大規模かつ多段階で自己修正されたデータセットが不足している。この限界に対処するため、我々はCoVA-SFTを導入する。これは、5つの異なるレイアウトファミリーと17の複雑なタスクにわたる222K以上のマルチモーダル推論ステップを含む、51.9Kサンプルからなる高度に構造化されたコーパスである。さらに、再現可能な評価のために同じタスクを網羅する1,700のホールドアウトテストサンプルからなる付随ベンチマークであるCoVA-Benchも導入する。CoVA-SFTは、明示的な根拠の定式化、エージェント的レンダリング、検証ループを提供することで、マルチモーダル言語モデルにテキストと視覚的抽象化を交互に配置することを教える。我々は、CoVA-SFTで微調整されたモデルがCoVA-Bench上のすべてのインターリーブ型CoTベースラインを平均で2倍以上上回ることを実証し、このデータセットを検証する。ただし、それらは依然として強力なテキストのみのCoTベースラインには及ばず、今後の研究における未解決の課題を浮き彫りにしている。
English
Chain-of-thought (CoT) reasoning has dramatically improved large language models (LLMs) by allowing them to decompose problems into intermediate steps. While CoT is widely effective for linguistic tasks, text-only CoT forces models to serialize visual problems into awkward prose. Although architectural solutions exist to process visual inputs, the community lacks a massive, multi-step, self-corrected dataset to teach models how to build and maintain internal visual workspaces when solving purely textual reasoning problems. To address this limitation, we introduce CoVA-SFT, a highly structured corpus of 51.9K samples containing over 222K multimodal reasoning steps across 5 distinct layout families and 17 complex tasks, and CoVA-Bench, a companion benchmark of 1,700 held-out test samples spanning the same tasks for reproducible evaluation. By providing explicit rationale formulations, agentic renderings, and verification loops, CoVA-SFT teaches multimodal language models to interleave text and visual abstractions. We validate the dataset by demonstrating that models fine-tuned on CoVA-SFT outperform all interleaved CoT baselines by more than 2x on average on CoVA-Bench, though they still fall short of strong text-only CoT baselines, highlighting open challenges for future work.