ChatPaper.aiChatPaper

並行影像理解與生成:自我修正的耦合馬可夫跳躍過程

Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes

July 14, 2026
作者: Minh-Quan Le, Armand Comas, Alexandros Lattas, Stylianos Moschoglou, Pedro Vélez, Amit Raj, Aaron Germuth, Thabo Beeler, Dimitris Samaras, Di Qiu
cs.AI

摘要

人類認知並不區分理解與生成。站在白板前的教師邊說邊畫,每一種模態都在重塑另一種模態。在本論文中,我們將這種耦合循環引入人工系統。遮罩擴散模型非常適合這項任務,然而現有的取樣器要交錯解碼文字與圖像,要透過平行分支獨立更新,且這些分支僅共享前一步的歷史資訊,卻不共享同一時間步內另一模態的最新決策;再搭配上遮罩擴散模型無法重新遮罩的特性,跨模態的矛盾既無法被偵測也無法被修復。我們提出「自校正耦合馬可夫跳躍過程」,這套框架中,一個模態的轉移速率是另一個模態信心分數的函數,並以跨模態注意力加權。此外,一個重新遮罩跳躍會在建構跨模態證據反轉時撤回已做出的承諾。結合此框架,我們推出「CO_2Jump」,這是一種無需訓練、單遍取樣的耦合模態生成取樣器。為進行訓練與評估,我們建立了三個大規模的聯合多模態生成語料庫:JEdit-1M、JMaze-200K 以及 JNono-200K,並附上對應的分佈內與分佈外基準測試。CO_2Jump 在圖像理解與編輯,以及視覺推理(迷宮與黑白非邏輯解題)上均達到最佳的聯合表現。該取樣器的效能隨去噪步數單調遞增,顯示跨模態耦合的效益在整條軌跡中持續疊加。專案頁面:https://coupled-jump.github.io
English
Human cognition does not separate understanding and generation. A teacher at a whiteboard speaks and draws together, each modality reshapes the other. In this paper, we bring this coupled loop to artificial systems. Masked Diffusion Models (MDMs) are ideally suited to this task, yet existing samplers either decode text and image interleavedly or independently update them in parallel branches that share only previous-step history, but not the other modality's latest decisions within the same step; combined with MDMs' inability to remask, cross-modal contradictions are neither detected nor repaired. We introduce Self-Correcting Coupled Markov Jump Processes (SC-CMJP), a framework in which one modality's transition rates are functionals of the other modality's confidence score, as weighted by cross-modal attention. Furthermore, a remasking jump retracts commitments the moment cross-modal evidence turns against them. In conjunction with SC-CMJP, we introduce CO_2Jump (Self-text{CO}rrecting text{CO}upled text{Jump}), a novel training-free single-pass sampler for joint multimodal geneneration. For training and evaluation purposes, we have created and will release three large-scale joint multimodal generation corpora: JEdit-1M, JMaze-200K, JNono-200K, with matching in- and out-of-distribution benchmarks. CO_2Jump achieves best joint performance for image understanding and editing as well as visual reasoning (maze and nonogram solving). The performance of the sampler scales monotonically with the number of denoising steps, evidence that the benefits of cross-modal coupling compound across the trajectory. Project page: https://coupled-jump.github.io