ChatPaper.aiChatPaper

텍스트 템플릿 토큰은 확산 트랜스포머에서 암시적 의미 레지스터입니다.

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers

July 21, 2026
저자: Maohua Li, Qirui Li, Yanke Zhou, Yiduo Li, Zhaosheng Chi, Chao Xu, Cuifeng Shen, Yixuan Xu, Hanlin Tang, Kan Liu, Tao Lan, Lin Qu, Shao-Qun Zhang
cs.AI

초록

텍스트-이미지 확산 트랜스포머(DiT)는 텍스트와 이미지 토큰을 공동으로 처리하지만, 노이즈 제거 과정에서의 내부 연산은 여전히 제대로 이해되지 않고 있다. 본 연구는 토큰 범위, 헤드, 계층 전반에 걸친 주의집중 분해와 표적 개입을 결합한 현대 대규모 DiT를 위한 인과적 해석 가능성 프레임워크를 제안한다. 이를 활용해 프롬프트-내용 토큰과 구조적 템플릿 토큰을 분리한 결과, 구조적 토큰은 인코더 출력에서 프롬프트 특정 정보를 거의 포함하지 않는 것으로 나타났다. 그러나 놀랍게도 이들은 DiT 내에서 지배적인 이미지-텍스트 주의집중 싱크(attention sink)로 작용하며 객체 정체성을 인과적으로 유지하고, 암묵적 의미 레지스터 역할을 수행했다. 본 연구는 이러한 정체성이 간접적으로 획득됨을 보여주는데, 프롬프트 의미론은 먼저 이미지 잠재 변수에 주입된 후 프롬프트 토큰에서 직접 전달되지 않고 템플릿 토큰으로 다시 읽혀진다. 위 발견에서 영감을 받아, 본 연구는 DiT를 위한 훈련 없는 가지치기 규칙을 설계했다. 프롬프트 토큰에 가장 강하게 주의를 기울이는 헤드는 불필요하며, 이를 제거하면 GenEval에서 단 1.4포인트 하락만으로 주의집중 FLOP의 20%를 제거한다. 또한 DiT에서 생성 연산이 헤드와 깊이에 걸쳐 어떻게 구성되는지, 의미 라우팅과 시각적 합성을 분리하고 정체성 형성에서 전파 및 정제로 진행되는 과정을 밝혀낸다. 본 연구는 입력에서 의미를 인코딩하는 토큰이 생성 중에 이를 유지하는 토큰과 반드시 일치할 필요가 없음을 보여줄 뿐만 아니라, DiT의 내부 메커니즘에 대한 인과적 관점을 제공한다.
English
Text-to-image diffusion transformers (DiTs) jointly process text and image tokens, yet their internal computation during denoising remains poorly understood. We introduce a causal interpretability framework for modern large-scale DiTs that combines attention decomposition with targeted interventions across token spans, heads, and layers. Using it to separate prompt-content tokens from structural template tokens, we find that the structural tokens carry little prompt-specific information at the encoder output. Yet surprisingly, they emerge as dominant image-to-text attention sinks and causally maintain object identity inside the DiT, acting as implicit semantic registers. We show that they acquire this identity indirectly, with prompt semantics first injected into the image latents and then read back into the template tokens rather than transferred directly from the prompt tokens. Inspired by the above findings, we design a training-free pruning rule for DiTs. Heads that attend most strongly to prompt tokens are dispensable, and pruning them removes 20% of attention FLOPs with only a 1.4-point drop on GenEval. We further reveal how generative computation in DiTs is organized across heads and depth, separating semantic routing from visual synthesis and progressing from identity formation to propagation and refinement. Our work not only reveals that the tokens encoding semantics at input need not be those that maintain it during generation, but also provides a causal view of internal mechanisms in DiTs.