Recursive Introspection: Teaching Language Model Agents How to Self-Improve

Yuxiao Qu, Tianjun Zhang, Naman Garg, Aviral Kumar

arXiv:2407.18219 · 2026-07-27 공개 · arXiv · PDF

reinforcement-learning math-reasoning llm-finetuning imitation-learning self-improvement mistral llama2 llama3

Abstract

A central piece in enabling intelligent agentic behavior in foundation models is to make them capable of introspecting upon their behavior, reasoning, and correcting their mistakes as more computation or interaction is available. Even the strongest proprietary large language models (LLMs) do not quite exhibit the ability of continually improving their responses sequentially, even in scenarios where they are explicitly told that they are making a mistake. In this paper, we develop RISE: Recursive IntroSpEction, an approach for fine-tuning LLMs to introduce this capability, despite prior work hypothesizing that this capability may not be possible to attain. Our approach prescribes an iterative fine-tuning procedure, which attempts to teach the model how to alter its response after having executed previously unsuccessful attempts to solve a hard test-time problem, with optionally additional environment feedback. RISE poses fine-tuning for a single-turn prompt as solving a multi-turn Markov decision process (MDP), where the initial state is the prompt. Inspired by principles in online imitation learning and reinforcement learning, we propose strategies for multi-turn data collection and training so as to imbue an LLM with the capability to recursively detect and correct its previous mistakes in subsequent iterations. Our experiments show that RISE enables Llama2, Llama3, and Mistral models to improve themselves with more turns on math reasoning tasks, outperforming several single-turn strategies given an equal amount of inference-time computation. We also find that RISE scales well, often attaining larger benefits with more capable models. Our analysis shows that RISE makes meaningful improvements to responses to arrive at the correct solution for challenging prompts, without disrupting one-turn abilities as a result of expressing more complex distributions.

한국어 요약

한 줄 요약

RISE는 LLM이 단계적으로 자기 개선할 수 있도록 훈련하는 반복적 미세조정 알고리즘으로, 수학 추론 성능을 향상시킨다.

핵심 기여도

핵심 아이디어

기존 LLM은 단일 턴 내에서 최선의 답변을 생성하지만, 복잡한 문제에서는 반복적 피드백을 통해 스스로 개선하는 능력이 필요하다. RISE는 이 자기 개선 능력을 반복적 미세조정을 통해 학습하도록 설계되었다. 핵심 아이디어는 단일 턴 프롬프트를 다중 턴의 MDP로 모델링하고, 온라인 흉내 학습과 강화 학습의 원리를 결합하여 모델이 자신의 오류를 반복적으로 탐지하고 수정하도록 유도하는 것이다. RISE는 학습자 모델이 자신의 분포 내에서 샘플링한 후보 답변 중 최고 품질을 선택하는 best-of-N 전략을 사용하여, 학습자가 스스로 개선하는 과정을 시뮬레이션한다.

기술적 접근법

주요 결과

의의 및 한계

RISE는 LLM이 단순히 단일 턴 내에서 최적 답변을 생성하는 것을 넘어, 반복적 턴을 통해 스스로 개선하는 능력을 습득하도록 유도하는 학습 프레임워크로, 수학 추론 등 복잡한 문제 해결에 기여한다. 특히, RISE는 더 강력한 모델일수록 더 큰 성능 향상을 보이는 점에서 확장성도 뛰어나다. 그러나 RISE는 특정 유형의 문제(예: 논리적 추론)에 효과적일 수 있으며, 다른 유형의 문제에서는 동일한 성능 향상을 보장하지 못할 수 있다. 또한, 학습 데이터가 모델 자체에서 생성되기 때문에 외부 피드백 없이 개선이 제한될 수 있다.

실용적 활용

RISE는 수학 문제 해결, 코드 생성, 복잡한 질의 응답 등 반복적 개선이 필요한 LLM 기반 에이전트 개발에 활용 가능하다. 특히, 사용자 피드백 없이도 모델이 스스로 오류를 수정하고 정확도를 높이는 시스템 구축에 유용하다. 산업적으로는 자동화된 문제 해결, 챗봇, 인공지능 튜터 등에 적용 가능하다.