ChatPaper.aiChatPaper

Aphanta:診斷任務對齊的圖像編輯中間表徵以進行多模態推理

Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning

August 27, 2026
作者: Hengyuan Xu, Wei Cheng, Yumeng Ji, Xuanyang Zhang, Xianfang Zeng, Gang Yu, Xingjun Ma
cs.AI

摘要

顯式視覺中間表徵可以幫助多模態大型語言模型(MLLMs)外化空間證據與更新的視覺狀態,但其效用取決於影像編輯器能否忠實實現所需的轉換。我們提出Aphanta,一個針對「MLLM → 影像編輯器 → MLLM」管線的自動化任務發現與閉環診斷框架。Aphanta評估三種條件:直接推理、使用編輯器生成的中間表徵進行推理,以及使用理想化參考中間表徵進行推理,以區分潛在的視覺提升空間與當前編輯器的實際效用。在20個候選任務及多種「編輯器-MLLM」組合中,我們發現效用具有強烈的任務條件依存性。增益集中於視覺線索注入、視覺定位與反事實狀態實現,而需要符號敏感建構或結構外推的中間表徵則明顯較不可靠。在選定的正向任務子集上,我們整合的Qwen管線將平均任務分數從0.343提升至0.445(+10.2分;相對提升29.7%);同時,完整研究亦保留了被篩選與未成功的任務,以揭示其界限。這些結果將影像編輯定位為一個專門的視覺工作空間,而非通用的推理機制,並確立Aphanta作為一個可重用的協議,用於衡量任務-表徵對齊、編輯器實現能力以及下游管線效用。
English
Explicit visual intermediates can help multimodal large language models (MLLMs) externalize spatial evidence and updated visual states, but their utility depends on whether an image editor can faithfully realize the required transformation. We introduce Aphanta, an automated task-discovery and closed-loop diagnostic framework for the MLLM -> image editor -> MLLM pipeline. Aphanta evaluates three conditions---direct reasoning, reasoning with an editor-generated intermediate, and reasoning with an idealized reference intermediate---to separate potential visual headroom from the practical utility of current editors. Across 20 candidate tasks and multiple editor--MLLM combinations, we find that utility is strongly task-conditioned. Gains concentrate in visual cue injection, grounding, and counterfactual state realization, whereas intermediates requiring symbol-sensitive construction or structural extrapolation are substantially less reliable. On the selected positive-task subset, our consolidated Qwen pipeline improves the mean task score from 0.343 to 0.445 (+10.2 points; +29.7% relative), while the full study also retains filtered and unsuccessful tasks to expose the boundary. These results position image editing as a specialized visual workspace rather than a universal reasoning mechanism, and establish Aphanta as a reusable protocol for measuring task--representation alignment, editor realization, and downstream pipeline utility.