Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models

Jean de Dieu Nyandwi, Leena Mathur, Yonatan Bisk, Robert Hawkins, Graham Neubig

arXiv:2608.13760 · 2026-08-17 공개 · arXiv · PDF

vision-language reasoning-models self-correction self-awareness knowledge-alignment confidence-calibration reasoning-training llm-vlm

Abstract

Which reasoning behaviors are associated with correct answers in reasoning models, and does reasoning-oriented training amplify those behaviors? This distinction is important because reasoning-oriented training can make traces look more deliberative without amplifying the behaviors most tied to model correctness. We quantify this mismatch with Behavioral Lift, a metric that measures how much correctness changes when a behavior is present versus absent in a model's reasoning trace. Across 15 models and 6 benchmarks spanning text-only and vision-language reasoning, we annotate 15,282 traces with a taxonomy whose core behaviors are defined for both LLM and VLM traces. We find evidence for an Amplification-Lift Gap, in which thinking models strongly amplify self-correction, hypothesis testing, and uncertainty acknowledgment, while the highest-lift behaviors are confidence calibration, knowledge alignment, and self-awareness. Confidence calibration is among the strongest positive signals of correctness in both modalities, yet is barely amplified; uncertainty acknowledgment is amplified by 3--7$\times$, yet is weakly or negatively associated with correctness. We find that reasoning-oriented training does not preferentially amplify the highest-Lift behaviors, motivating process-level objectives that reward calibrated and grounded reasoning rather than surface form alone.

한국어 요약

한 줄 요약

15개 모델을 대상으로 15,282개 추론 트레이스를 분석하여, 사고 훈련이 정확도와 강하게 연관된 행동을 증폭시키지 않는다는 사실을 밝힘.

핵심 기여도

핵심 아이디어

기존 사고 훈련은 추론 트레이스가 더 길고 복잡해지도록 유도하지만, 이는 반드시 정확한 추론과 연관되지 않는다.
이 연구는 **Behavioral Lift**라는 새로운 지표를 통해, 행동의 빈도가 높은 것과 정확도에 긍정적인 영향을 주는 것이 다른 개념임을 밝힘.
예를 들어, **Self-correction**은 추론이 실패한 후에 나타나는 경우가 많아, 빈도가 높다고 해서 성능이 좋은 것은 아님.
반면, **Confidence Calibration**은 추론의 강도에 따라 자신감 수준을 조절하는 행동으로, 정확도와 강하게 연관되지만, 사고 훈련에서 거의 증폭되지 않음.
이러한 격차는 추론 훈련 목표가 단순히 추론 트레이스의 길이나 외형적 복잡도를 증가시키는 데 머무르고 있음을 시사.

기술적 접근법

주요 결과

의의 및 한계

이 연구는 추론 모델의 훈련 목표가 단순히 추론 트레이스의 길이나 외형적 복잡도를 증가시키는 데 머무르고 있음을 지적하며, **정확도와 강하게 연관된 행동**(예: Confidence Calibration, Knowledge Alignment)을 강화하는 새로운 훈련 전략이 필요함을 제안.
또한, 추론 트레이스는 모델의 내부 계산 과정을 완전히 반영하지 않을 수 있으며, 어노테이션은 자동화된 판단기(예: GPT-4o)에 의존하므로 시스템적 편향 가능성 있음.
**Behavioral Lift**는 설명적 지표이기 때문에, 행동이 성능 향상의 원인인지, 결과인지, 또는 다른 요인과 상호작용하는지 구분하기 어렵다는 한계가 있음.

실용적 활용

이 연구는 추론 모델의 훈련 및 평가 시, 단순히 추론 트레이스의 길이나 외형적 복잡도를 기준으로 하지 말고, **Confidence Calibration**, **Knowledge Alignment**, **Self-awareness**와 같은 행동을 강화하는 목표를 설정해야 함을 제안.
이러한 접근은 의료, 법률, 금융 등 정확한 추론이 필수적인 분야에서 모델 신뢰도를 높이는 데 활용될 수 있음.