ピクセル空間拡散トランスフォーマーにおけるレジスタの重要性
Registers Matter for Pixel-Space Diffusion Transformers
July 6, 2026
著者: Nikita Starodubcev, Ilia Sudakov, Ilya Drobyshevskiy, Artem Babenko, Dmitry Baranchuk
cs.AI
要旨
Vision Transformers(ViT)は、高ノルムのパッチトークン外れ値を示し、特徴マップの品質を低下させることが知られており、この問題はレジスタトークンによって効果的に緩和される。拡散モデルがトランスフォーマーアーキテクチャを採用し、画素空間での学習へと移行するにつれて、その形式はViTに近づき、レジスタトークンがDiffusion Transformers(DiT)にも有効かどうかという疑問が生じる。本研究では、DiTがViTとは重要な点で異なることを示す。すなわち、DiTはパッチトークン外れ値を示さないが、それでもレジスタの恩恵を受ける。興味深いことに、レジスタは潜在空間DiTよりも画素空間DiTにおいてより効果的である。中間表現を分析することで、レジスタトークンが高ノイズレベルでよりクリーンな特徴マップを生成することを発見し、これが画素空間生成におけるその有効性に寄与している可能性がある。さらに、最近の画素空間DiTアーキテクチャは暗黙的にレジスタのようなメカニズムを取り入れており、これがその高い実証的性能を部分的に説明している可能性があることを観察する。これらの観察に動機づけられ、視覚的構造と一貫性を向上させる役割を担うレジスタトークンの寄与を増幅する手法であるRegister Guidanceを提案する。
English
Vision Transformers (ViTs) are known to exhibit high-norm patch-token outliers that degrade feature map quality, a problem effectively mitigated by register tokens. As diffusion models increasingly adopt transformer architectures and move toward pixel-space training, they become closer in form to ViTs, raising the question of whether register tokens are also useful for Diffusion Transformers (DiTs). In this work, we show that DiTs differ from ViTs in a key respect: they do not exhibit patch-token outliers but still benefit from registers. Interestingly, registers are more effective in pixel-space DiTs than in latent-space DiTs. By analyzing intermediate representations, we find that register tokens produce cleaner feature maps at high noise levels, which may contribute to their effectiveness in pixel-space generation. We further observe that recent pixel-space DiT architectures implicitly incorporate register-like mechanisms, which may partially account for their strong empirical performance. Motivated by these observations, we propose Register Guidance, a technique that amplifies the contribution of register tokens responsible for improving visual structure and coherence.