Aphanta: マルチモーダル推論のためのタスク整合画像編集中間生成物の診断
Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning
August 27, 2026
著者: Hengyuan Xu, Wei Cheng, Yumeng Ji, Xuanyang Zhang, Xianfang Zeng, Gang Yu, Xingjun Ma
cs.AI
要旨
明示的な視覚的中間表現は、マルチモーダル大規模言語モデル(MLLM)が空間的証拠や更新された視覚状態を外在化することを可能にするが、その有用性は、画像エディタが必要とされる変換を忠実に実現できるかどうかに依存する。本稿では、MLLM→画像エディタ→MLLMというパイプラインのための自動タスク発見と閉ループ診断フレームワークであるAphantaを紹介する。Aphantaは、直接推論、エディタ生成中間表現を用いた推論、理想化された参照中間表現を用いた推論という3つの条件を評価することで、現在のエディタの実用的有用性から潜在的な視覚的改善余地を分離する。20の候補タスクと複数のエディタ・MLLMの組み合わせにわたって、有用性はタスクに強く依存することがわかった。改善効果は、視覚的手がかりの注入、グラウンディング、反事実的状態の実現に集中する一方、記号に敏感な構築や構造的外挿を必要とする中間表現は、かなり信頼性が低い。選択されたポジティブタスク部分集合では、統合したQwenパイプラインにより平均タスクスコアが0.343から0.445へ改善した(+10.2ポイント、相対+29.7%)。一方、全研究では、フィルタリングされたタスクや失敗したタスクも保持し、その適用限界を明らかにしている。これらの結果は、画像編集を普遍的な推論メカニズムではなく、専門的な視覚的ワークスペースとして位置づけ、Aphantaを、タスクと表現の整合性、エディタの実現能力、下流パイプラインの有用性を測定するための再利用可能なプロトコルとして確立するものである。
English
Explicit visual intermediates can help multimodal large language models (MLLMs) externalize spatial evidence and updated visual states, but their utility depends on whether an image editor can faithfully realize the required transformation. We introduce Aphanta, an automated task-discovery and closed-loop diagnostic framework for the MLLM -> image editor -> MLLM pipeline. Aphanta evaluates three conditions---direct reasoning, reasoning with an editor-generated intermediate, and reasoning with an idealized reference intermediate---to separate potential visual headroom from the practical utility of current editors. Across 20 candidate tasks and multiple editor--MLLM combinations, we find that utility is strongly task-conditioned. Gains concentrate in visual cue injection, grounding, and counterfactual state realization, whereas intermediates requiring symbol-sensitive construction or structural extrapolation are substantially less reliable. On the selected positive-task subset, our consolidated Qwen pipeline improves the mean task score from 0.343 to 0.445 (+10.2 points; +29.7% relative), while the full study also retains filtered and unsuccessful tasks to expose the boundary. These results position image editing as a specialized visual workspace rather than a universal reasoning mechanism, and establish Aphanta as a reusable protocol for measuring task--representation alignment, editor realization, and downstream pipeline utility.