寄存器對像素空間擴散Transformer至關重要
Registers Matter for Pixel-Space Diffusion Transformers
July 6, 2026
作者: Nikita Starodubcev, Ilia Sudakov, Ilya Drobyshevskiy, Artem Babenko, Dmitry Baranchuk
cs.AI
摘要
視覺Transformer(ViTs)已知會出現高範數的補丁標記異常值,導致特徵圖品質下降,而寄存器標記能有效緩解此問題。隨著擴散模型越來越多採用Transformer架構並轉向像素空間訓練,其形式與ViT日益接近,這引發了一個問題:寄存器標記對擴散Transformer(DiTs)是否同樣有用?在本研究中,我們證明DiT與ViT在一個關鍵方面有所不同:DiT不會出現補丁標記異常值,但仍可從寄存器中受益。有趣的是,寄存器在像素空間DiT中的效果優於在潛空間DiT中的效果。透過分析中間表徵,我們發現寄存器標記在高雜訊水準下會產生更乾淨的特徵圖,這可能有助於它們在像素空間生成中的有效性。我們進一步觀察到,近期像素空間DiT架構隱含地納入了類似寄存器的機制,這或許能部分解釋其強大的實證表現。基於這些觀察,我們提出寄存器引導(Register Guidance)技術,該技術能放大負責改善視覺結構與一致性的寄存器標記之貢獻。
English
Vision Transformers (ViTs) are known to exhibit high-norm patch-token outliers that degrade feature map quality, a problem effectively mitigated by register tokens. As diffusion models increasingly adopt transformer architectures and move toward pixel-space training, they become closer in form to ViTs, raising the question of whether register tokens are also useful for Diffusion Transformers (DiTs). In this work, we show that DiTs differ from ViTs in a key respect: they do not exhibit patch-token outliers but still benefit from registers. Interestingly, registers are more effective in pixel-space DiTs than in latent-space DiTs. By analyzing intermediate representations, we find that register tokens produce cleaner feature maps at high noise levels, which may contribute to their effectiveness in pixel-space generation. We further observe that recent pixel-space DiT architectures implicitly incorporate register-like mechanisms, which may partially account for their strong empirical performance. Motivated by these observations, we propose Register Guidance, a technique that amplifies the contribution of register tokens responsible for improving visual structure and coherence.