ChatPaper.aiChatPaper

레지스터는 픽셀 공간 확산 트랜스포머에서 중요하다

Registers Matter for Pixel-Space Diffusion Transformers

July 6, 2026
저자: Nikita Starodubcev, Ilia Sudakov, Ilya Drobyshevskiy, Artem Babenko, Dmitry Baranchuk
cs.AI

초록

비전 트랜스포머(ViT)는 높은 노름(norm)을 가진 패치 토큰 이상치(outlier)를 나타내어 특징 맵 품질을 저하시키는 것으로 알려져 있으며, 이 문제는 레지스터 토큰(register token)에 의해 효과적으로 완화됩니다. 확산 모델이 점차 트랜스포머 아키텍처를 채택하고 픽셀 공간 학습으로 나아감에 따라, 이들은 형태적으로 ViT에 더 가까워지며, 레지스터 토큰이 확산 트랜스포머(DiT)에도 유용한지에 대한 의문이 제기됩니다. 본 연구에서 우리는 DiT가 ViT와 중요한 측면에서 다르다는 것을 보여줍니다. 즉, DiT는 패치 토큰 이상치를 나타내지 않지만 레지스터의 이점을 여전히 누립니다. 흥미롭게도, 레지스터는 잠재 공간 DiT보다 픽셀 공간 DiT에서 더 효과적입니다. 중간 표현을 분석함으로써, 레지스터 토큰이 높은 노이즈 수준에서 더 깨끗한 특징 맵을 생성하며, 이는 픽셀 공간 생성에서의 효과성에 기여할 수 있음을 발견했습니다. 우리는 또한 최근의 픽셀 공간 DiT 아키텍처가 암시적으로 레지스터와 유사한 메커니즘을 포함하고 있으며, 이것이 강력한 경험적 성능을 부분적으로 설명할 수 있음을 관찰합니다. 이러한 관찰에 기반하여, 우리는 시각적 구조와 일관성을 개선하는 데 기여하는 레지스터 토큰의 기여를 증폭시키는 기술인 레지스터 가이던스(Register Guidance)를 제안합니다.
English
Vision Transformers (ViTs) are known to exhibit high-norm patch-token outliers that degrade feature map quality, a problem effectively mitigated by register tokens. As diffusion models increasingly adopt transformer architectures and move toward pixel-space training, they become closer in form to ViTs, raising the question of whether register tokens are also useful for Diffusion Transformers (DiTs). In this work, we show that DiTs differ from ViTs in a key respect: they do not exhibit patch-token outliers but still benefit from registers. Interestingly, registers are more effective in pixel-space DiTs than in latent-space DiTs. By analyzing intermediate representations, we find that register tokens produce cleaner feature maps at high noise levels, which may contribute to their effectiveness in pixel-space generation. We further observe that recent pixel-space DiT architectures implicitly incorporate register-like mechanisms, which may partially account for their strong empirical performance. Motivated by these observations, we propose Register Guidance, a technique that amplifies the contribution of register tokens responsible for improving visual structure and coherence.