Large Language Models Can Learn Temporal Reasoning

Siheng Xiong, Ali Payani, R. Kompella, F. Fekri

arXiv:2401.06853 · 2026-07-27 공개 · arXiv · PDF

chain-of-thought llm-finetuning temporal-reasoning synthetic-dataset graph-data-augmentation tg-llm temporal-graph tgqa

Abstract

While large language models (LLMs) have demonstrated remarkable reasoning capabilities, they are not without their flaws and inaccuracies. Recent studies have introduced various methods to mitigate these limitations. Temporal reasoning (TR), in particular, presents a significant challenge for LLMs due to its reliance on diverse temporal concepts and intricate temporal logic. In this paper, we propose TG-LLM, a novel framework towards language-based TR. Instead of reasoning over the original context, we adopt a latent representation, temporal graph (TG) that enhances the learning of TR. A synthetic dataset (TGQA), which is fully controllable and requires minimal supervision, is constructed for fine-tuning LLMs on this text-to-TG translation task. We confirmed in experiments that the capability of TG translation learned on our dataset can be transferred to other TR tasks and benchmarks. On top of that, we teach LLM to perform deliberate reasoning over the TGs via Chain-of-Thought (CoT) bootstrapping and graph data augmentation. We observed that those strategies, which maintain a balance between usefulness and diversity, bring more reliable CoTs and final results than the vanilla CoT distillation.

한국어 요약

한 줄 요약

TG-LLM을 제안하여 언어 기반 시공간 추론 성능을 향상시킨다.

핵심 기여도

핵심 아이디어

기존 LLM은 시공간 추론(TR)에서 다양한 시간 개념과 복잡한 논리 처리에 어려움을 겪는다. 본 연구는 TR을 개선하기 위해 텍스트를 시공간 그래프(TG)로 변환한 후, 이 그래프 상에서 추론을 수행하는 새로운 프레임워크 TG-LLM을 제안한다. TG는 시간 순서, 지속 시간, 빈도 등 TR에 필요한 핵심 개념을 구조화하여 표현하며, 이를 통해 LLM이 더 직관적이고 구조화된 정보를 처리할 수 있도록 돕는다. 또한, 합성 데이터셋 TGQA를 통해 텍스트-그래프 변환 능력을 학습시켜 다른 TR 작업으로 전이 학습이 가능하도록 한다. CoT 부트스트랩핑과 그래프 데이터 증강을 통해 생성된 추론 경로의 신뢰도를 높이며, 기존 CoT 디스틸레이션 방식보다 더 정확한 결과를 도출한다.

기술적 접근법

주요 결과

의의 및 한계

TG-LLM은 TR 작업에서 LLM의 추론 능력을 구조화된 방식으로 향상시키는 새로운 패러다임을 제시한다. 특히, TG라는 중간 표현을 도입함으로써 추론 과정의 해석성과 신뢰도를 높인다. 또한, 합성 데이터셋 TGQA를 통해 최소한의 감독으로도 훈련이 가능하다는 점에서 실용적 가치가 있다. 그러나, TGQA는 합성 데이터셋이므로 실제 세계 데이터와의 차이가 있을 수 있으며, 이는 일반화 능력에 영향을 줄 수 있다. 또한, CoT 부트스트랩핑은 훈련 데이터의 질에 크게 의존하므로, 데이터 생성 과정에서 오류가 발생할 경우 성능 저하가 나타날 수 있다.

실용적 활용

TG-LLM은 일정 관리, 의료 진단, 법적 문서 분석 등 시간 기반 정보를 처리해야 하는 산업 분야에 적용 가능하다. 또한, 복잡한 추론이 필요한 인공지능 시스템 개발에도 활용될 수 있다.