Subliminal Clocks: Latent Time Modelling in Diffusion Language Models

Maximo Eduardo Rulli, Thomas Vaitses Fontanari, Simone Petruzzi, Federico Alvetreti, Giorgio Strano, Donato Crisostomi, Giorgos Nikolaou, Tommaso Mencattini, Andrea Santilli, Emanuele Rodolà, Simone Scardapane, Alessio Devoto

arXiv:2607.01774 · 2026-07-23 공개 · arXiv · PDF

diffusion-language-models model-interpretability residual-streams activation-probing subspace-steering diffusion-timestep latent-time-modelling denoising-progress

Abstract

Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive models. Unlike standard diffusion-based approaches, DLMs are not explicitly conditioned on a timestep, raising a natural question: do these models internally represent denoising progress, and how is such information used downstream? In this work, we show that DLMs do in fact encode a latent representation related to the diffusion timestep within their residual streams. We find that this signal can be reliably extracted using probes across layers, indicating that denoising progress is decodable from internal activations. We further demonstrate that steering the model along a low-dimensional subspace associated with the inferred timestep allows us to systematically modulate its notion of denoising progress, leading to predictable changes in model confidence and entropy. Finally, we analyse the geometry of the identified representation, showing that it exhibits structured and interpretable properties in activation space, and shedding light on how such a signal is processed by these models.

한국어 요약

한 줄 요약

DLMs가 잔여 스트림 내에서 확산 단계 정보를 내재적으로 인코딩하고 이를 제어할 수 있음을 밝힘.

핵심 기여도

핵심 아이디어

DLMs는 확산 시간에 명시적으로 조건을 주지 않지만, 내부적으로는 denoising progress를 추적하고 이를 계산에 활용한다는 가설을 제기한다. 이 연구는 잔여 스트림 내의 활성화를 분석하여 이 정보가 레이어별로 일관되게 인코딩됨을 확인하고, 이를 추출하여 모델의 확신도와 엔트로피를 제어할 수 있음을 보인다. 특히, 추출된 신호는 mean vector 방향을 따라 유도될 수 있으며, 이는 단순히 설명적인 것이 아니라 모델의 계산에 기능적으로 관련됨을 나타낸다.

기술적 접근법

주요 결과

의의 및 한계

이 연구는 DLMs가 내부적으로 denoising progress를 추적하고 이를 계산에 활용함을 실증적으로 보여주어, 모델의 내부 동작 이해에 기여한다. 특히, **masked token 비율**을 기반으로 한 확산 진행도 추정이 모델 내부에서 구조화된 방식으로 표현됨을 밝힘. 그러나 신호가 어떻게 계산되는지, 즉 **mask-ratio estimation**, **attention-MLP 상호작용**, 또는 **분산 시퀀스 통계** 중 어떤 경로를 통해 생성되는지는 명시되지 않으며, 이는 향후 연구 주제로 제시된다.

실용적 활용

이 연구는 DLMs의 내부 신호를 해석하고 제어하는 데 기초가 되며, 생성 과정의 신뢰도 조절, 텍스트 생성의 다양성 조정, 또는 모델의 안정성 향상에 활용될 수 있다. 특히, **low-dimensional steering** 기법은 모델의 내재적 표현을 조정하여 예측 가능한 결과를 유도하는 데 유용할 수 있다.