Do LLMs Understand Sequential Structure? A Controlled Study of Inference and Generation

Jerry Wang, Zhengxiang Wang, Ting Yu Liu, Hsin-Ling Hsu, Yi-Cheng Lai, Tengfei Ma

arXiv:2610.04977 · 2026-10-10 공개 · arXiv · PDF

llm generative-models context-length conditional-dependencies behavioral-simulation sequential-structure rock-paper-scissors markov-rules

Abstract

Large language models (LLMs) are increasingly used as interactive agents and simulators, yet it remains unclear whether they can recover latent sequential structure beyond surface action frequencies. This distinction is critical for behavioral simulation, where actions are often shaped by prior context rather than marginal frequencies alone. We study this question using controlled two-player Rock--Paper--Scissors interactions and a one-player stochastic n-gram continuation task. Across these experiments, we test whether LLMs can identify latent strategies, follow simple Markov rules, and sustain higher-order conditional dependencies. Our framework separates distribution matching from conditional rule following. Results show that longer context does not improve identification, correct recognition does not ensure faithful simulation, and higher-order dependencies substantially degrade rule recovery. Apparent behavioral fidelity can therefore mask incorrect generative mechanisms.

한국어 요약

한 줄 요약

LLMs가 표면 행동 통계 이상의 순차 구조를 이해하는지 실험적으로 검증한 연구.

핵심 기여도

핵심 아이디어

LLMs는 순차적 상호작용 환경에서 사용되지만, 이들이 단순한 표면 행동 빈도 이상의 조건부 구조를 이해하고 생성할 수 있는지는 명확하지 않았다. 본 연구는 RPS 게임과 n-gram 생성을 통해 LLM이 **잠재 전략을 식별**, **마르코프 규칙을 따르기**, **고차 조건부 의존성을 유지**할 수 있는지 실험적으로 검증한다. 특히, 표면 행동 분포와 실제 생성 메커니즘 사이의 차이를 분리하여 평가하는 것이 핵심이다. 연구는 **distribution matching**과 **conditional rule following**을 구분하는 프레임워크를 제시하며, 이는 기존 연구에서 혼동되었던 능력을 명확히 분리한다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 LLM이 순차적 행동을 시뮬레이션할 때 **표면 행동과 실제 생성 메커니즘 사이의 차이를 명확히 분리**할 수 있는지 평가하는 데 기여한다. 특히, **고차 조건부 구조의 복원 능력**이 LLM의 순차 생성 능력의 한계를 드러내며, 이는 인공지능 기반 시뮬레이션의 신뢰도에 영향을 미친다. 한계로는 **정확한 실험 설정과 모델 성능 수치가 일부 누락**되었으며, **더 다양한 LLM 아키텍처와 데이터셋에 대한 일반화 가능성**이 제한적임을 지적할 수 있다.

실용적 활용

이 연구는 **사용자 행동 시뮬레이션**, **다중 에이전트 상호작용**, **복잡한 순차적 의사결정 시스템** 등에서 LLM의 신뢰도를 평가하는 데 활용될 수 있다. 특히, **고차 조건부 의존성을 요구하는 시스템**(예: 금융 시장 예측, 의료 진단 시스템)에서 LLM의 한계를 명확히 인식하고 보완 전략을 수립하는 데 기초 자료가 될 수 있다.