The pitfalls of next-token prediction

Gregor Bachmann, Vaishnavh Nagarajan

arXiv:2403.06963 · 2026-07-27 공개 · arXiv · PDF

transformer mamba next-token-prediction autoregressive-inference model-failure teacher-forcing teacherless-training planning-task

Abstract

Can a mere next-token predictor faithfully model human intelligence? We crystallize this emerging concern and correct popular misconceptions surrounding it, and advocate a simple multi-token objective. As a starting point, we argue that the two often-conflated phases of next-token prediction -- autoregressive inference and teacher-forced training -- must be treated distinctly. The popular criticism that errors can compound during autoregressive inference, crucially assumes that teacher-forcing has learned an accurate next-token predictor. This assumption sidesteps a more deep-rooted problem we expose: in certain classes of tasks, teacher-forcing can simply fail to learn an accurate next-token predictor in the first place. We describe a general mechanism of how teacher-forcing can fail, and design a minimal planning task where both the Transformer and the Mamba architecture empirically fail in that manner -- remarkably, despite the task being straightforward to learn. Finally, we provide preliminary evidence that this failure can be resolved using _teacherless_ training, a simple modification using dummy tokens that predicts multiple tokens in advance. We hope this finding can ground future debates and inspire explorations beyond the next-token prediction paradigm. We make our code available under https://github.com/gregorbachmann/Next-Token-Failures

한국어 요약

한 줄 요약

Next-token prediction이 인간 지능을 충분히 모델링하지 못할 수 있음을 밝히고, teacher-forcing의 한계를 실증적으로 분석한다.

핵심 기여도

핵심 아이디어

Next-token prediction은 인간의 문제 해결 방식과 본질적으로 다를 수 있다. 인간은 실행 전에 계획을 상상하고 수정하지만, next-token predictor는 단순히 다음 토큰만 예측한다. 이 논문은 teacher-forcing이 next-token predictor를 정확히 학습하지 못하는 경우가 있음을 지적한다. 이는 "Clever Hans cheat"라고 명명된 메커니즘으로 설명되는데, teacher-forcing이 ground truth의 일부를 제공받아 미래 토큰을 학습하는 과정에서 단서를 악용하게 된다. 이로 인해 초기 토큰 학습이 어려워지고, 전체적인 예측 정확도가 저하된다. 이는 특히 "lookahead task"에서 두드러지며, Transformer와 Mamba 아키텍처 모두 실패함을 보여준다.

기술적 접근법

주요 결과

의의 및 한계

이 연구는 next-token prediction이 인간 지능을 모델링하는 데 한계가 있을 수 있음을 실증적으로 보여준다. teacher-forcing이 학습 과정에서 자체적으로 오류를 유발할 수 있다는 점은 기존 연구에서 간과되었던 핵심 문제다. 그러나 실험은 매우 제한된 task에만 적용되었으며, 실제 복잡한 언어 생성 상황에서의 일반화 가능성은 명시되지 않았다. 또한, teacherless training의 효과는 아직 초기 단계로, 더 많은 연구가 필요하다.

실용적 활용

이 연구는 next-token prediction을 기반으로 한 언어 모델의 학습 전략을 재검토할 필요성을 제시한다. 특히, planning이 필요한 생성 작업(예: 스토리 작성, 코드 생성)에서 teacherless training과 같은 대안 접근법이 유용할 수 있다. 또한, 모델이 단순히 다음 토큰을 예측하는 데만 의존하지 않고, 전체적인 계획을 고려하는 방식으로 학습할 수 있도록 하는 새로운 학습 목표 개발에 영감을 줄 수 있다.