同時画像理解と生成:自己補正結合マルコフ跳躍過程
Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes
July 14, 2026
著者: Minh-Quan Le, Armand Comas, Alexandros Lattas, Stylianos Moschoglou, Pedro Vélez, Amit Raj, Aaron Germuth, Thabo Beeler, Dimitris Samaras, Di Qiu
cs.AI
要旨
人間の認知は、理解と生成を分離しない。ホワイトボードの前で教師は話しながら描き、それぞれのモダリティが互いを再形成する。本論文では、この結合ループを人工システムに導入する。マスク拡散モデル(Masked Diffusion Models; MDMs)はこの課題に理想的に適合するが、既存のサンプラーはテキストと画像を交互に復号するか、あるいは前のステップの履歴のみを共有し同一ステップ内で他方のモダリティの最新の決定を共有しない並列分岐で独立に更新する。これに加え、MDMが再マスクできないことにより、クロスモーダルな矛盾は検出も修正もされない。本稿では、自己修正結合マルコフ跳躍過程(Self-Correcting Coupled Markov Jump Processes; SC-CMJP)を導入する。これは、一方のモダリティの遷移速度が、クロスモーダル注意機構で重み付けされた他方のモダリティの信頼度スコアの汎関数となる枠組みである。さらに、再マスク跳躍により、クロスモーダルな証拠が逆転した瞬間にコミットメントを撤回する。SC-CMJPと併せて、学習不要のシングルパスサンプラーであるCO₂Jump(Self-Correcting COupled Jump)を導入する。これは、マルチモーダル共同生成のための新たなサンプラーである。訓練および評価のために、大規模なマルチモーダル共同生成コーパス3種(JEdit-1M、JMaze-200K、JNono-200K)を作成し、公開する。これらには、ドメイン内およびドメイン外のベンチマークが付随する。CO₂Jumpは、画像理解・編集、ならびに視覚的推論(迷路およびノノグラムの解法)において最良の共同性能を達成する。サンプラーの性能はノイズ除去ステップ数に単調に比例し、クロスモーダル結合の利点が軌跡全体にわたって累積されることを示す。プロジェクトページ:https://coupled-jump.github.io
English
Human cognition does not separate understanding and generation. A teacher at a whiteboard speaks and draws together, each modality reshapes the other. In this paper, we bring this coupled loop to artificial systems. Masked Diffusion Models (MDMs) are ideally suited to this task, yet existing samplers either decode text and image interleavedly or independently update them in parallel branches that share only previous-step history, but not the other modality's latest decisions within the same step; combined with MDMs' inability to remask, cross-modal contradictions are neither detected nor repaired. We introduce Self-Correcting Coupled Markov Jump Processes (SC-CMJP), a framework in which one modality's transition rates are functionals of the other modality's confidence score, as weighted by cross-modal attention. Furthermore, a remasking jump retracts commitments the moment cross-modal evidence turns against them. In conjunction with SC-CMJP, we introduce CO_2Jump (Self-text{CO}rrecting text{CO}upled text{Jump}), a novel training-free single-pass sampler for joint multimodal geneneration. For training and evaluation purposes, we have created and will release three large-scale joint multimodal generation corpora: JEdit-1M, JMaze-200K, JNono-200K, with matching in- and out-of-distribution benchmarks. CO_2Jump achieves best joint performance for image understanding and editing as well as visual reasoning (maze and nonogram solving). The performance of the sampler scales monotonically with the number of denoising steps, evidence that the benefits of cross-modal coupling compound across the trajectory. Project page: https://coupled-jump.github.io