Aphanta: 멀티모달 추론을 위한 Task-정렬 이미지-편집 중간 결과물 진단
Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning
August 27, 2026
저자: Hengyuan Xu, Wei Cheng, Yumeng Ji, Xuanyang Zhang, Xianfang Zeng, Gang Yu, Xingjun Ma
cs.AI
초록
명시적 시각 중간물은 멀티모달 대규모 언어 모델(MLLM)이 공간적 증거와 갱신된 시각 상태를 외부화하도록 도울 수 있지만, 그 유용성은 이미지 편집기가 요구된 변환을 충실히 구현할 수 있는지에 달려 있다. 본 연구는 MLLM → 이미지 편집기 → MLLM 파이프라인을 위한 자동화된 과업 발견 및 폐루프 진단 프레임워크인 Aphanta를 소개한다. Aphanta는 세 가지 조건—직접 추론, 편집기 생성 중간물을 사용한 추론, 이상화된 참조 중간물을 사용한 추론—을 평가하여 잠재적 시각적 개선 여지와 현재 편집기의 실질적 유용성을 분리한다. 20개 후보 과업과 다양한 편집기–MLLM 조합에 걸쳐, 유용성은 과업 조건에 크게 의존함을 발견했다. 이득은 시각적 단서 주입, 접지, 반사실적 상태 구현에 집중된 반면, 기호에 민감한 구성이나 구조적 외삽을 요구하는 중간물은 훨씬 덜 신뢰할 수 있었다. 선정된 긍정적 과업 부분집합에서, 통합된 Qwen 파이프라인은 평균 과업 점수를 0.343에서 0.445로 향상시켰으며(+10.2점, +29.7% 상대적), 전체 연구는 경계를 드러내기 위해 필터링된 과업과 실패한 과업도 함께 유지한다. 이러한 결과는 이미지 편집을 보편적 추론 메커니즘이라기보다 전문화된 시각적 작업 공간으로 위치시키며, Aphanta를 과업-표상 정렬, 편집기 구현, 하류 파이프라인 유용성을 측정하는 재사용 가능한 프로토콜로 확립한다.
English
Explicit visual intermediates can help multimodal large language models (MLLMs) externalize spatial evidence and updated visual states, but their utility depends on whether an image editor can faithfully realize the required transformation. We introduce Aphanta, an automated task-discovery and closed-loop diagnostic framework for the MLLM -> image editor -> MLLM pipeline. Aphanta evaluates three conditions---direct reasoning, reasoning with an editor-generated intermediate, and reasoning with an idealized reference intermediate---to separate potential visual headroom from the practical utility of current editors. Across 20 candidate tasks and multiple editor--MLLM combinations, we find that utility is strongly task-conditioned. Gains concentrate in visual cue injection, grounding, and counterfactual state realization, whereas intermediates requiring symbol-sensitive construction or structural extrapolation are substantially less reliable. On the selected positive-task subset, our consolidated Qwen pipeline improves the mean task score from 0.343 to 0.445 (+10.2 points; +29.7% relative), while the full study also retains filtered and unsuccessful tasks to expose the boundary. These results position image editing as a specialized visual workspace rather than a universal reasoning mechanism, and establish Aphanta as a reusable protocol for measuring task--representation alignment, editor realization, and downstream pipeline utility.