Multimodale Continue Redenering via Asymmetrisch Wederzijds Variationeel Leren
Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning
July 1, 2026
Auteurs: Shijie Li, Yilin Gao, Siyuan Yang, Tieyuan Chen, Chaofan Gan, Zhihao He, Zicheng Zhao, Yuyu Guo, Weiyao Lin, Hang Yu
cs.AI
Samenvatting
Multimodale Grote Taalmodellen (MLLMs) worden vaak beperkt door een taalruimte-knelpunt, waardoor complexe visuele redeneringen worden gedwongen in discrete tokens die perceptuele nuances kunnen verliezen. Een veelbelovend alternatief is continue latente redenering, waarbij het doel is om impliciete redeneerpaden te ontdekken die de multimodale query en het uiteindelijke antwoord overbruggen. Dit introduceert echter een ernstige trainings-inferentiemismatch: een posterior tijdens de training, die conditioneel is op het juiste antwoord, kan antwoordafhankelijke shortcuts exploiteren. Standaard variationele training dwingt vervolgens de prior tijdens inferentie om een posterior na te bootsen die toegang heeft tot informatie die tijdens het testen niet beschikbaar is, wat leidt tot slechte prestaties. Om dit aan te pakken, stellen we Asymmetric Mutual Variational Learning (AMVL) voor, een raamwerk dat deze mismatch oplost via een bidirectionele kalibratiedoelstelling. Een forward KL-divergentie traint de doel-agnostische prior om overeen te komen met de posterior, terwijl een nieuwe reverse KL-divergentie tegelijkertijd de posterior regulariseert, voorkomend dat deze instort in inferentie-incompatibele regio's en deze "antwoordlekkage" vermindert. We bieden een theoretische analyse die deze lekkage formaliseert als prior-contaminatie en bewijzen dat onze dual-KL-doelstelling deze vermindert. We implementeren AMVL in een latent-geïntegreerde MLLM en tonen aan dat het consistent beter presteert dan sterke discrete en latente redeneringsbaselines, met een verbetering van de gemiddelde score op de complexe BLINK-benchmark van +10,83 en winsten tot +32,00 op individuele redeneertaken, waarbij analyses een verbeterde stabiliteit van de latentruimte bevestigen.
English
Multimodal Large Language Models (MLLMs) are often constrained by a language-space bottleneck, forcing complex visual reasoning into discrete tokens which can lose perceptual nuance. A promising alternative is continuous latent reasoning, where the goal is to discover implicit reasoning pathways that bridge the multimodal query and the final answer. However, this introduces a severe train-inference mismatch: a training-time posterior, conditioned on the ground-truth answer, can exploit answer-dependent shortcuts. Standard variational training then forces the inference-time prior to mimic a posterior that has access to information unavailable at test time, leading to poor performance. To address this, we propose Asymmetric Mutual Variational Learning (AMVL), a framework that resolves this mismatch via a bidirectional calibration objective. A forward KL divergence trains the target-agnostic prior to match the posterior, while a novel reverse KL divergence simultaneously regularizes the posterior, preventing it from collapsing into inference-incompatible regions and mitigating this ``answer leakage''. We provide theoretical analysis formalizing this leakage as prior contamination and prove that our dual-KL objective reduces it. We instantiate AMVL in a latent-integrated MLLM and show that it consistently outperforms strong discrete and latent-reasoning baselines, improving the average score on the complex BLINK benchmark by +10.83 and achieving gains of up to +32.00 on individual reasoning tasks, with analyses confirming improved latent-space stability.