Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations

Jiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang, Rui Li, Xuan Cao, Leon Gao, Zhaojie Gong, Fangda Gu, Michael He, Yin-Hua Lu, Yu Shi

arXiv:2402.17152 · 2026-07-27 공개 · arXiv · PDF

recommendation-systems large-scale-models high-cardinality transformer-based hstu-architecture ndcg-metrics carbon-footprint sequential-transducers

Abstract

Large-scale recommendation systems are characterized by their reliance on high cardinality, heterogeneous features and the need to handle tens of billions of user actions on a daily basis. Despite being trained on huge volume of data with thousands of features, most Deep Learning Recommendation Models (DLRMs) in industry fail to scale with compute. Inspired by success achieved by Transformers in language and vision domains, we revisit fundamental design choices in recommendation systems. We reformulate recommendation problems as sequential transduction tasks within a generative modeling framework ("Generative Recommenders"), and propose a new architecture, HSTU, designed for high cardinality, non-stationary streaming recommendation data. HSTU outperforms baselines over synthetic and public datasets by up to 65.8% in NDCG, and is 5.3x to 15.2x faster than FlashAttention2-based Transformers on 8192 length sequences. HSTU-based Generative Recommenders, with 1.5 trillion parameters, improve metrics in online A/B tests by 12.4% and have been deployed on multiple surfaces of a large internet platform with billions of users. More importantly, the model quality of Generative Recommenders empirically scales as a power-law of training compute across three orders of magnitude, up to GPT-3/LLaMa-2 scale, which reduces carbon footprint needed for future model developments, and further paves the way for the first foundational models in recommendations.

한국어 요약

한 줄 요약

1.5trillion 파라미터를 가진 HSTU 기반 Generative Recommenders가 기존 DLRM 대비 12.4% 성능 향상과 5.3x~15.2x 속도 개선을 달성했다.

핵심 기여도

핵심 아이디어

기존 DLRM은 고카디널리티, 이질적 특성, 비정적 데이터 처리에 어려움을 겪는다. 본 연구는 추천 문제를 순차적 트랜스덕션 문제로 재정의하고, **Generative Recommenders**라는 새로운 프레임워크를 제안한다. 이는 사용자 행동을 새로운 모달리티로 간주하여, 순차적 생성 방식으로 모델을 학습하게 한다. **HSTU**는 고카디널리티, 비정적 데이터를 처리하기 위해 어텐션 메커니즘을 수정한 모듈로, 기존 Transformer 대비 빠른 속도를 보인다. 또한, **M-FALCON** 알고리즘은 미니배치를 활용해 추론 비용을 동일 수준 유지하면서도 모델 복잡도를 285배까지 증가시킬 수 있다.

기술적 접근법

주요 결과

의의 및 한계

**Generative Recommenders**는 추천 시스템의 첫 번째 Foundation Model 기반을 제시하며, 사용자 행동을 생성 모델링의 새로운 모달리티로 활용할 수 있음을 보여준다. 또한, 기존 DLRM의 스케일링 한계를 극복하고, GPT-3 수준의 모델까지 확장 가능함을 입증했다. 그러나, 1.5trillion 파라미터 모델의 학습 및 배포는 여전히 높은 컴퓨팅 자원을 요구하며, 실시간 추론 성능에 대한 구체적 제약은 명시되지 않았다. 또한, 모델의 일반화 능력은 특정 플랫폼 내에서만 검증되었으며, 다른 도메인으로의 확장 가능성은 추가 연구가 필요하다.

실용적 활용

HSTU 기반 GR은 수십억 사용자를 보유한 대형 인터넷 플랫폼에서 이미 배포되어 있으며, 추천, 검색, 광고 등 다양한 도메인에서 통합된 추천 시스템 구축에 활용될 수 있다. 특히, 사용자 행동을 기반으로 한 생성 모델링은 개인화된 콘텐츠 추천, 실시간 맞춤형 마케팅, 사용자 행동 분석 등에 실용적 가치를 제공한다.