ChatPaper.aiChatPaper

潜在的時計:拡散言語モデルにおける潜在時間モデリング

Subliminal Clocks: Latent Time Modelling in Diffusion Language Models

July 20, 2026
著者: Maximo Eduardo Rulli, Thomas Vaitses Fontanari, Simone Petruzzi, Federico Alvetreti, Giorgio Strano, Donato Crisostomi, Giorgos Nikolaou, Tommaso Mencattini, Andrea Santilli, Emanuele Rodolà, Simone Scardapane, Alessio Devoto
cs.AI

要旨

拡散言語モデル(DLM)は、近年、自己回帰モデルに代わる有望な手法として登場した。標準的な拡散ベースの手法とは異なり、DLMはタイムステップに明示的に条件付けられていないため、自然な疑問が生じる:これらのモデルは内部でノイズ除去の進行状況を表現しているのか、またそのような情報は下流でどのように利用されるのか?本研究では、DLMが実際に残差ストリーム内に拡散タイムステップに関連する潜在表現を符号化していることを示す。この信号は層を横断したプローブを用いて確実に抽出可能であり、ノイズ除去の進行状況が内部活性化から復号可能であることがわかる。さらに、推定されたタイムステップに関連する低次元部分空間に沿ってモデルを操縦することで、ノイズ除去の進行に関するモデルの概念を系統的に変調でき、その結果、モデルの信頼度とエントロピーに予測可能な変化が生じることを実証する。最後に、特定された表現の幾何学を解析し、活性化空間において構造化され解釈可能な特性を示すこと、またそのような信号がこれらのモデルによってどのように処理されるかを明らかにする。
English
Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive models. Unlike standard diffusion-based approaches, DLMs are not explicitly conditioned on a timestep, raising a natural question: do these models internally represent denoising progress, and how is such information used downstream? In this work, we show that DLMs do in fact encode a latent representation related to the diffusion timestep within their residual streams. We find that this signal can be reliably extracted using probes across layers, indicating that denoising progress is decodable from internal activations. We further demonstrate that steering the model along a low-dimensional subspace associated with the inferred timestep allows us to systematically modulate its notion of denoising progress, leading to predictable changes in model confidence and entropy. Finally, we analyse the geometry of the identified representation, showing that it exhibits structured and interpretable properties in activation space, and shedding light on how such a signal is processed by these models.