Learning to (Learn at Test Time): RNNs with Expressive Hidden States

Yu Sun, Xinhao Li, Karan Dalal, Jiarui Xu, Arjun Vikram, Ge Zhang, Yann Dubois, Xinlei Chen, Xiaolong Wang, Oluwasanmi Koyejo, Tatsunori Hashimoto, Carlos Guestrin

arXiv:2407.04620 · 2026-07-27 공개 · arXiv · PDF

transformer long-context mamba self-attention sequence-modeling test-time-training rnn ttt-linear

Abstract

Self-attention performs well in long context but has quadratic complexity. Existing RNN layers have linear complexity, but their performance in long context is limited by the expressive power of their hidden states. We present a practical framework for instantiating sequence modeling layers with linear complexity and expressive hidden states. The key idea is to make the hidden state a machine learning model itself, and the update rule a step of self-supervised learning. Since the hidden state is updated by training even on test sequences, our layers are called Test-Time Training (TTT) layers. We consider two instantiations: TTT-Linear and TTT-MLP, whose hidden state is a linear model and a two-layer MLP respectively. We evaluate our instantiations at the scale of 125M to 1.3B parameters, comparing with a strong Transformer and Mamba, a modern RNN. Similar to Transformer, TTT-Linear and TTT-MLP can keep reducing perplexity by conditioning on more tokens, while Mamba cannot after 16k context. TTT-MLP still faces challenges in memory I/O, but shows larger potential in long context, pointing to a promising direction for future research.

한국어 요약

한 줄 요약

TTT 레이어는 테스트 시점에 학습하는 RNN 구조로, 긴 문맥에서 성능을 유지하면서도 선형 복잡도를 달성한다.

핵심 기여도

핵심 아이디어

기존 RNN은 고정 크기의 히든 상태로 문맥을 압축하기 때문에 긴 시퀀스에서 정보 손실이 발생한다. 이에 반해, TTT 레이어는 히든 상태 자체를 머신러닝 모델로 구현하고, 테스트 시점에서도 자기 감독 학습을 통해 업데이트하도록 설계했다. 이는 히든 상태가 시퀀스의 구조와 관계를 학습적으로 압축할 수 있게 해준다. TTT-Linear는 선형 모델, TTT-MLP는 2층 MLP를 히든 상태로 사용하며, 이는 각각 다른 표현력과 성능을 보인다. 핵심 아이디어는 "학습을 (테스트 시점에) 학습하는 것"으로, 이는 기존 네트워크 아키텍처 개념을 재정의하는 새로운 접근법이다.

기술적 접근법

주요 결과

의의 및 한계

TTT 레이어는 RNN의 선형 복잡도 장점을 유지하면서, 긴 문맥에서의 표현력 한계를 해결하는 새로운 방향을 제시한다. 특히 TTT-MLP는 MLP 기반 히든 상태로 더 복잡한 관계를 학습할 수 있어, Mamba와 같은 기존 RNN보다 긴 문맥에서 우수한 성능을 보인다. 그러나 TTT-MLP는 메모리 I/O 문제를 여전히 겪고 있어, 효율적인 구현이 필요하다. 또한, TTT-MLP는 짧은 문맥에서는 TTT-Linear보다 성능이 떨어지는 경향이 있어, 모델 복잡도와 문맥 길이 간의 균형이 중요하다.

실용적 활용

TTT 레이어는 대규모 언어 모델에서 긴 문맥 처리가 필요한 상황, 예를 들어 문서 요약, 대화 시스템, 코드 생성 등에 적용 가능하다. 특히, Mamba와 같은 기존 RNN이 한계를 보이는 16k 이상의 긴 문맥에서 유용하며, 선형 복잡도를 유지하면서도 Transformer의 성능에 근접할 수 있어, 실용적이고 효율적인 대안이 될 수 있다.