The Impact of Reasoning Step Length on Large Language Models

Mingyu Jin, Qinkai Yu, Dong Shu, Haiyan Zhao, Wenyue Hua, Yanda Meng, Yongfeng Zhang, Mengnan Du

arXiv:2401.04925 · 2026-07-27 공개 · arXiv · PDF

large-language-models chain-of-thought prompt-engineering reasoning-abilities task-dependency reasoning-step-length incorrect-rationales empirical-experiments

Abstract

Chain of Thought (CoT) is significant in improving the reasoning abilities of large language models (LLMs). However, the correlation between the effectiveness of CoT and the length of reasoning steps in prompts remains largely unknown. To shed light on this, we have conducted several empirical experiments to explore the relations. Specifically, we design experiments that expand and compress the rationale reasoning steps within CoT demonstrations while keeping all other factors constant. We have the following key findings. First, the results indicate that lengthening the reasoning steps in prompts, even without adding new information into the prompt, considerably enhances LLMs' reasoning abilities across multiple datasets. Alternatively, shortening the reasoning steps, even while preserving the key information, significantly diminishes the reasoning abilities of models. This finding highlights the importance of the number of steps in CoT prompts and provides practical guidance to make better use of LLMs' potential in complex problem-solving scenarios. Second, we also investigated the relationship between the performance of CoT and the rationales used in demonstrations. Surprisingly, the result shows that even incorrect rationales can yield favorable outcomes if they maintain the requisite length of inference. Third, we observed that the advantages of increasing reasoning steps are task-dependent: simpler tasks require fewer steps, whereas complex tasks gain significantly from longer inference sequences. The code is available at https://github.com/MingyuJ666/The-Impact-of-Reasoning-Step-Length-on-Large-Language-Models

한국어 요약

한 줄 요약

Chain of Thought (CoT)의 추론 단계 길이가 LLM의 추론 성능에 큰 영향을 미친다는 것을 실험적으로 밝혔다.

핵심 기여도

핵심 아이디어

CoT는 인간의 단계적 추론을 모방한 프롬프팅 기법으로, LLM의 추론 능력을 향상시키는 데 효과적이다. 그러나 CoT의 효과가 단계 수에 얼마나 의존하는지는 명확하지 않았다. 본 연구는 CoT 내 추론 단계의 길이를 조절하면서 다른 요소는 고정한 실험을 통해 이 관계를 탐구했다. 핵심 통찰은 추론 단계의 길이가 LLM의 성능에 큰 영향을 미친다는 점이다. 특히, 추론 단계가 길수록 모델이 더 복잡한 문제를 해결하는 데 유리하며, 단계 수가 줄어들면 성능이 급격히 하락한다. 또한, 추론 과정이 정확하지 않더라도 단계 수가 충분하면 성능이 개선된다는 점에서, 추론의 정확성보다는 단계 수가 더 중요한 역할을 한다는 결론이 도출되었다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 CoT의 단계 수가 LLM의 추론 성능에 직접적인 영향을 미친다는 것을 명확히 밝혀내어, CoT 기법의 최적화 전략에 실질적인 지침을 제공한다. 특히, 추론 단계의 정확성보다는 길이가 더 중요한 역할을 한다는 점은 기존 인식과 상반되며, 추론 프롬프팅 설계에 새로운 관점을 제시한다. 그러나 본 연구는 특정 모델(GPT-3.5, GPT-4)에만 국한되었으며, 다른 LLM(예: LLaMA, Claude)에 대한 일반화 가능성은 명시되지 않았다. 또한, 추론 단계 수가 무한히 늘어날수록 성능이 지속적으로 향상되는지에 대한 분석은 포함되지 않았다.

실용적 활용

본 연구는 복잡한 문제 해결을 요구하는 산업 분야(예: 금융, 법률, 과학 연구)에서 LLM의 활용을 최적화하는 데 기여할 수 있다. 특히, 추론 단계 수를 조절하여 모델의 성능을 조정할 수 있으므로, 사용자 맞춤형 추론 프롬프팅 설계에 활용할 수 있다. 또한, 추론 과정의 정확성보다는 길이가 더 중요한 점을 고려해, 추론 단계를 자동 생성하는 도구 개발에도 활용 가능하다.