ChatPaper.aiChatPaper

Focus op wat ertoe doet: Saillantie-gestuurde nauwkeurige routing voor Diffusie MoE

Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoE

June 25, 2026
Auteurs: Haoyou Deng, Keyu Yan, Chaojie Mao, Xiang Wang, Yu Liu, Changxin Gao, Nong Sang
cs.AI

Samenvatting

Mixture-of-Experts (MoE)-architecturen zijn naar voren gekomen als een krachtig paradigma voor het schalen van diffusiemodellen in visuele generatie. Recente vooruitgang heeft zich gericht op het adaptief toewijzen van rekenbronnen over diverse tokens om efficiëntie en prestaties te verbeteren. Echter, wij identificeren een routeringstoewijzingsprobleem in bestaande diffusie-MoE-raamwerken: de router slaagt er niet in om nauwkeurig meer rekenbronnen toe te wijzen aan opvallende tokens. Onze analyse wijt dit falen aan de afhankelijkheid van de router van door ruis aangetaste latente kenmerken gedurende het ontruisingsproces. Dergelijke stochastische ruis verbergt de kritische structurele en textuurinformatie, waardoor de router wordt verhinderd om opvallende tokens effectief te onderscheiden. Om dit aan te pakken, stellen we SharpMoE voor, een post-training raamwerk met een nauwkeurig routeringsmechanisme dat saillantie benut, waarbij schone latente kenmerken worden gebruikt als een ruisvrij stuur signaal voor routering. Door de door ruis vervormde invoer te omzeilen, voorziet SharpMoE de router van duidelijke saillantie-aanwijzingen, waardoor identificatie van opvallende tokens mogelijk is, zelfs in fasen met hoge ruis. Verder introduceren we een trajectrouteringsverlies om de toewijzing van rekenkracht gedurende het meerstaps ontruisingstraject te beperken, wat zorgt voor een nauwkeurige toewijzing van bronnen langs de generatie-uitrol. Uitgebreide experimenten tonen aan dat SharpMoE dient als een veelzijdige, plug-and-play-oplossing die de vooraf getrainde, geconvergeerde MoE-modellen verder verbetert en state-of-the-art prestaties in visuele generatie bereikt.
English
Mixture-of-Experts (MoE) architectures have emerged as a powerful paradigm for scaling diffusion models in visual generation. Recent advancements have focused on adaptively allocating computational resources across diverse tokens to improve efficiency and performance. However, we identify a routing assignment problem in existing diffusion MoE frameworks: the router fails to accurately allocate more computational resources to salient tokens. Our analysis attributes this failure to the router's reliance on noise-corrupted latent features throughout the denoising process. Such stochastic noise obscures the critical structural and textural information, thereby preventing the router from effectively distinguishing salient tokens. To address this, we propose SharpMoE, a post-training framework with a saliency-harnessing accurate routing mechanism, which utilizes clean latent features as a noise-free guidance signal for routing. By bypassing the noise-distorted inputs, SharpMoE provides the router with clear saliency guidance, enabling the identification of salient tokens even in high-noise stages. Furthermore, we introduce a trajectory routing loss to constrain the compute allocation throughout the multi-step denoising trajectory, ensuring precise resource allocation along the generation rollout. Extensive experiments demonstrate that SharpMoE serves as a versatile, plug-and-play solution that further enhances the pretrained, converged MoE models, achieving state-of-the-art performance in visual generation.