并发图像理解与生成:自校正耦合马尔可夫跳变过程
Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes
July 14, 2026
作者: Minh-Quan Le, Armand Comas, Alexandros Lattas, Stylianos Moschoglou, Pedro Vélez, Amit Raj, Aaron Germuth, Thabo Beeler, Dimitris Samaras, Di Qiu
cs.AI
摘要
人类认知并非将理解与生成割裂开来。当教师在白板上边讲解边绘图时,两种模态相互重塑。本文将此耦合循环引入人工系统。掩码扩散模型(MDM)天然适合此任务,但现有采样器要么交错解码文本与图像,要么在并行分支中独立更新——这些分支仅共享前一步历史信息,却不包含同一步内另一模态的最新决策;加之MDM不具备重掩码能力,跨模态矛盾既无法被检测也无法被修复。我们提出自校正耦合马尔可夫跳跃过程(SC-CMJP)框架:在该框架中,一种模态的转移速率是另一模态置信度分数的泛函,该置信度由跨模态注意力加权。此外,当跨模态证据出现矛盾时,重掩码跳跃会立即撤回先前承诺。基于SC-CMJP,我们提出CO₂Jump(自校正耦合跳跃),一种无需训练的单遍采样器,专用于联合多模态生成。为训练与评估,我们创建并将发布三个大规模联合多模态生成语料库:JEdit-1M、JMaze-200K、JNono-200K,附带匹配的分布内与分布外基准测试。CO₂Jump在图像理解与编辑、视觉推理(迷宫与非图求解)中均取得最优联合性能。该采样器的性能随去噪步数单调提升,证明跨模态耦合的优势沿轨迹持续累积。项目页面:https://coupled-jump.github.io
English
Human cognition does not separate understanding and generation. A teacher at a whiteboard speaks and draws together, each modality reshapes the other. In this paper, we bring this coupled loop to artificial systems. Masked Diffusion Models (MDMs) are ideally suited to this task, yet existing samplers either decode text and image interleavedly or independently update them in parallel branches that share only previous-step history, but not the other modality's latest decisions within the same step; combined with MDMs' inability to remask, cross-modal contradictions are neither detected nor repaired. We introduce Self-Correcting Coupled Markov Jump Processes (SC-CMJP), a framework in which one modality's transition rates are functionals of the other modality's confidence score, as weighted by cross-modal attention. Furthermore, a remasking jump retracts commitments the moment cross-modal evidence turns against them. In conjunction with SC-CMJP, we introduce CO_2Jump (Self-text{CO}rrecting text{CO}upled text{Jump}), a novel training-free single-pass sampler for joint multimodal geneneration. For training and evaluation purposes, we have created and will release three large-scale joint multimodal generation corpora: JEdit-1M, JMaze-200K, JNono-200K, with matching in- and out-of-distribution benchmarks. CO_2Jump achieves best joint performance for image understanding and editing as well as visual reasoning (maze and nonogram solving). The performance of the sampler scales monotonically with the number of denoising steps, evidence that the benefits of cross-modal coupling compound across the trajectory. Project page: https://coupled-jump.github.io