ChatPaper.aiChatPaper

Gelijktijdig Beeldbegrip en -generatie: Zelfcorrigerende Gekoppelde Markov-sprongprocessen

Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes

July 14, 2026
Auteurs: Minh-Quan Le, Armand Comas, Alexandros Lattas, Stylianos Moschoglou, Pedro Vélez, Amit Raj, Aaron Germuth, Thabo Beeler, Dimitris Samaras, Di Qiu
cs.AI

Samenvatting

De menselijke cognitie scheidt begrip en generatie niet. Een leraar voor een whiteboard spreekt en tekent tegelijkertijd, waarbij elke modaliteit de andere hervormt. In dit artikel brengen we deze gekoppelde lus naar kunstmatige systemen. Gemaskeerde Diffusiemodellen (MDMs) zijn ideaal geschikt voor deze taak, maar bestaande samplers decoderen tekst en beeld ofwel afwisselend, ofwel updaten ze onafhankelijk in parallelle takken die alleen de geschiedenis van de vorige stap delen, maar niet de nieuwste beslissingen van de andere modaliteit binnen dezelfde stap; in combinatie met het onvermogen van MDMs om te hermaskeren worden cross-modale tegenstrijdigheden noch gedetecteerd noch hersteld. We introduceren Zelfcorrigerende Gekoppelde Markov-sprongprocessen (SC-CMJP), een raamwerk waarin de overgangssnelheden van de ene modaliteit functionalen zijn van de vertrouwensscore van de andere modaliteit, gewogen door cross-modale aandacht. Bovendien trekt een hermaskeringssprong toezeggingen in op het moment dat cross-modale evidentie ertegen keert. In combinatie met SC-CMJP introduceren we CO\(_2\)Jump (zelfcorrigerende gekoppelde sprong), een nieuwe trainingsvrije single-pass sampler voor gezamenlijke multimodale generatie. Voor trainings- en evaluatiedoeleinden hebben we drie grootschalige corpora voor gezamenlijke multimodale generatie gecreëerd en zullen we deze vrijgeven: JEdit-1M, JMaze-200K, JNono-200K, met bijbehorende in-distribution en out-of-distribution benchmarks. CO\(_2\)Jump behaalt de beste gezamenlijke prestaties voor beeldbegrip en -bewerking, evenals visueel redeneren (doolhof- en nonogram-oplossen). De prestaties van de sampler schalen monotoon met het aantal ontruisingsstappen, wat bewijs levert dat de voordelen van cross-modale koppeling zich in de loop van het traject opstapelen. Projectpagina: https://coupled-jump.github.io
English
Human cognition does not separate understanding and generation. A teacher at a whiteboard speaks and draws together, each modality reshapes the other. In this paper, we bring this coupled loop to artificial systems. Masked Diffusion Models (MDMs) are ideally suited to this task, yet existing samplers either decode text and image interleavedly or independently update them in parallel branches that share only previous-step history, but not the other modality's latest decisions within the same step; combined with MDMs' inability to remask, cross-modal contradictions are neither detected nor repaired. We introduce Self-Correcting Coupled Markov Jump Processes (SC-CMJP), a framework in which one modality's transition rates are functionals of the other modality's confidence score, as weighted by cross-modal attention. Furthermore, a remasking jump retracts commitments the moment cross-modal evidence turns against them. In conjunction with SC-CMJP, we introduce CO_2Jump (Self-text{CO}rrecting text{CO}upled text{Jump}), a novel training-free single-pass sampler for joint multimodal geneneration. For training and evaluation purposes, we have created and will release three large-scale joint multimodal generation corpora: JEdit-1M, JMaze-200K, JNono-200K, with matching in- and out-of-distribution benchmarks. CO_2Jump achieves best joint performance for image understanding and editing as well as visual reasoning (maze and nonogram solving). The performance of the sampler scales monotonically with the number of denoising steps, evidence that the benefits of cross-modal coupling compound across the trajectory. Project page: https://coupled-jump.github.io