ChatPaper.aiChatPaper

동시적 이미지 이해 및 생성: 자기 교정 결합 마르코프 점프 과정

Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes

July 14, 2026
저자: Minh-Quan Le, Armand Comas, Alexandros Lattas, Stylianos Moschoglou, Pedro Vélez, Amit Raj, Aaron Germuth, Thabo Beeler, Dimitris Samaras, Di Qiu
cs.AI

초록

인간의 인지는 이해와 생성을 분리하지 않는다. 칠판 앞의 교사는 말하고 동시에 그림을 그리며, 각 양식(modality)은 서로를 재구성한다. 본 논문에서는 이러한 결합 루프(coupled loop)를 인공 시스템에 도입한다. 마스크 확산 모델(Masked Diffusion Models, MDMs)은 이 작업에 이상적으로 적합하지만, 기존 샘플러는 텍스트와 이미지를 교차(interleaved)하여 디코딩하거나 이전 단계 기록만 공유하고 동일 단계 내에서 다른 양식의 최신 결정은 공유하지 않는 병렬 분기에서 독립적으로 업데이트한다. 여기에 MDM의 재마스킹(remasking) 불가능성이 더해져, 교차 양식 간 모순(cross-modal contradiction)은 탐지되지도 않고 수리되지도 않는다. 우리는 자기 수정 결합 마르코프 점프 프로세스(Self-Correcting Coupled Markov Jump Processes, SC-CMJP)를 도입한다. 이 프레임워크에서는 한 양식의 전이율(transition rate)이 교차 양식 주의(cross-modal attention)에 의해 가중된 다른 양식의 신뢰 점수(confidence score)의 범함수(functional)이다. 나아가, 재마스킹 점프(remasking jump)는 교차 양식 증거(cross-modal evidence)가 반대 방향으로 전환되는 순간 약속(commitment)을 철회한다. SC-CMJP와 함께, 우리는 CO$_2$Jump (Self-text{CO}rrecting text{CO}upled text{Jump})를 도입한다. 이는 공동 다중 모달 생성(joint multimodal generation)을 위한 새로운 훈련 없는 단일 패스 샘플러(single-pass sampler)이다. 훈련 및 평가 목적을 위해, 우리는 세 가지 대규모 공동 다중 모달 생성 코퍼스(corpus)인 JEdit-1M, JMaze-200K, JNono-200K를 생성하였으며, 이에 대응하는 분포 내(in-distribution) 및 분포 외(out-of-distribution) 벤치마크와 함께 공개할 예정이다. CO$_2$Jump는 이미지 이해 및 편집뿐만 아니라 시각적 추론(미로 및 노노그램 풀이)에서 최고의 공동 성능을 달성한다. 샘플러의 성능은 잡음 제거 단계(denoising step) 수에 따라 단조롭게 증가하며, 이는 교차 양식 결합의 이점이 궤적(trajectory) 전체에 걸쳐 누적된다는 증거이다. 프로젝트 페이지: https://coupled-jump.github.io
English
Human cognition does not separate understanding and generation. A teacher at a whiteboard speaks and draws together, each modality reshapes the other. In this paper, we bring this coupled loop to artificial systems. Masked Diffusion Models (MDMs) are ideally suited to this task, yet existing samplers either decode text and image interleavedly or independently update them in parallel branches that share only previous-step history, but not the other modality's latest decisions within the same step; combined with MDMs' inability to remask, cross-modal contradictions are neither detected nor repaired. We introduce Self-Correcting Coupled Markov Jump Processes (SC-CMJP), a framework in which one modality's transition rates are functionals of the other modality's confidence score, as weighted by cross-modal attention. Furthermore, a remasking jump retracts commitments the moment cross-modal evidence turns against them. In conjunction with SC-CMJP, we introduce CO_2Jump (Self-text{CO}rrecting text{CO}upled text{Jump}), a novel training-free single-pass sampler for joint multimodal geneneration. For training and evaluation purposes, we have created and will release three large-scale joint multimodal generation corpora: JEdit-1M, JMaze-200K, JNono-200K, with matching in- and out-of-distribution benchmarks. CO_2Jump achieves best joint performance for image understanding and editing as well as visual reasoning (maze and nonogram solving). The performance of the sampler scales monotonically with the number of denoising steps, evidence that the benefits of cross-modal coupling compound across the trajectory. Project page: https://coupled-jump.github.io