潛意識時鐘:擴散語言模型中的潛在時間建模
Subliminal Clocks: Latent Time Modelling in Diffusion Language Models
July 20, 2026
作者: Maximo Eduardo Rulli, Thomas Vaitses Fontanari, Simone Petruzzi, Federico Alvetreti, Giorgio Strano, Donato Crisostomi, Giorgos Nikolaou, Tommaso Mencattini, Andrea Santilli, Emanuele Rodolà, Simone Scardapane, Alessio Devoto
cs.AI
摘要
擴散語言模型(Diffusion Language Models, DLMs)近期已成為自迴歸模型之外一個有前景的替代方案。與標準的基於擴散的方法不同,DLMs 並非明確地以時間步為條件,這引發了一個自然問題:這些模型是否在內部表徵去噪進度,而此類資訊又如何在下游被利用?在本工作中,我們證明 DLMs 確實在其殘差流中編碼了與擴散時間步相關的潛在表徵。我們發現,此訊號可透過跨層的探測器可靠地提取,顯示去噪進度可從內部激活中解碼。我們進一步證明,沿著與推斷時間步相關的低維子空間引導模型,能讓我們系統性地調節其對去噪進度的認知,從而導致模型置信度與熵的可預測變化。最後,我們分析所辨識表徵的幾何結構,顯示其在激活空間中具有結構化且可解釋的性質,並闡明這類訊號如何被這些模型處理。
English
Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive models. Unlike standard diffusion-based approaches, DLMs are not explicitly conditioned on a timestep, raising a natural question: do these models internally represent denoising progress, and how is such information used downstream? In this work, we show that DLMs do in fact encode a latent representation related to the diffusion timestep within their residual streams. We find that this signal can be reliably extracted using probes across layers, indicating that denoising progress is decodable from internal activations. We further demonstrate that steering the model along a low-dimensional subspace associated with the inferred timestep allows us to systematically modulate its notion of denoising progress, leading to predictable changes in model confidence and entropy. Finally, we analyse the geometry of the identified representation, showing that it exhibits structured and interpretable properties in activation space, and shedding light on how such a signal is processed by these models.