ChatPaper.aiChatPaper

Aphanta:面向多模态推理的任务对齐图像编辑中间产物诊断

Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning

August 27, 2026
作者: Hengyuan Xu, Wei Cheng, Yumeng Ji, Xuanyang Zhang, Xianfang Zeng, Gang Yu, Xingjun Ma
cs.AI

摘要

显式视觉中间表示可以帮助多模态大语言模型(MLLMs)外化空间证据和更新后的视觉状态,但其效用取决于图像编辑器能否忠实地实现所需的变换。我们提出了Aphanta,一个面向MLLM → 图像编辑器 → MLLM流程的自动化任务发现与闭环诊断框架。Aphanta评估三种条件——直接推理、使用编辑器生成的中间表示进行推理、以及使用理想化参考中间表示进行推理——以将潜在的视觉提升空间与当前编辑器的实际效用分离开来。在20个候选任务和多种编辑器–大语言模型组合中,我们发现效用强烈依赖于任务条件。增益集中在视觉线索注入、视觉定位和反事实状态实现上,而需要符号敏感构造或结构外推的中间表示则可靠性显著较低。在选定的正向任务子集上,我们整合的Qwen流程将平均任务得分从0.343提升至0.445(+10.2个百分点;相对提升29.7%),同时完整研究也保留了被过滤和未成功的任务以揭示边界。这些结果将图像编辑定位为一种专门的视觉工作空间,而非通用的推理机制,并将Aphanta确立为一种可复用的协议,用于衡量任务–表征对齐、编辑器实现能力以及下游流程效用。
English
Explicit visual intermediates can help multimodal large language models (MLLMs) externalize spatial evidence and updated visual states, but their utility depends on whether an image editor can faithfully realize the required transformation. We introduce Aphanta, an automated task-discovery and closed-loop diagnostic framework for the MLLM -> image editor -> MLLM pipeline. Aphanta evaluates three conditions---direct reasoning, reasoning with an editor-generated intermediate, and reasoning with an idealized reference intermediate---to separate potential visual headroom from the practical utility of current editors. Across 20 candidate tasks and multiple editor--MLLM combinations, we find that utility is strongly task-conditioned. Gains concentrate in visual cue injection, grounding, and counterfactual state realization, whereas intermediates requiring symbol-sensitive construction or structural extrapolation are substantially less reliable. On the selected positive-task subset, our consolidated Qwen pipeline improves the mean task score from 0.343 to 0.445 (+10.2 points; +29.7% relative), while the full study also retains filtered and unsuccessful tasks to expose the boundary. These results position image editing as a specialized visual workspace rather than a universal reasoning mechanism, and establish Aphanta as a reusable protocol for measuring task--representation alignment, editor realization, and downstream pipeline utility.