reinforcement-learning vision-language-model robotic-manipulation task-planning failure-detection dense-reward failgen aha-dataset
Abstract
Robotic manipulation in open-world settings requires not only task execution but also the ability to detect and learn from failures. While recent advances in vision-language models (VLMs) and large language models (LLMs) have improved robots' spatial reasoning and problem-solving abilities, they still struggle with failure recognition, limiting their real-world applicability. We introduce AHA, an open-source VLM designed to detect and reason about failures in robotic manipulation using natural language. By framing failure detection as a free-form reasoning task, AHA identifies failures and provides detailed, adaptable explanations across different robots, tasks, and environments. We fine-tuned AHA using FailGen, a scalable framework that generates the first large-scale dataset of robotic failure trajectories, the AHA dataset. FailGen achieves this by procedurally perturbing successful demonstrations from simulation. Despite being trained solely on the AHA dataset, AHA generalizes effectively to real-world failure datasets, robotic systems, and unseen tasks. It surpasses the second-best model (GPT-4o in-context learning) by 10.3% and exceeds the average performance of six compared models including five state-of-the-art VLMs by 35.3% across multiple metrics and datasets. We integrate AHA into three manipulation frameworks that utilize LLMs/VLMs for reinforcement learning, task and motion planning, and zero-shot trajectory generation. AHA's failure feedback enhances these policies' performances by refining dense reward functions, optimizing task planning, and improving sub-task verification, boosting task success rates by an average of 21.4% across all three tasks compared to GPT-4 models.
한국어 요약
한 줄 요약
AHA는 시각-언어 모델을 활용해 로봇 조작 실패를 탐지하고 추론하는 시스템으로, 실패 데이터셋 AHA와 FailGen을 통해 21.4% 성공률 향상.
핵심 기여도
- 실패 데이터셋 AHA를 생성하기 위한 FailGen 프레임워크 도입 (49K+ 이미지-쿼리 쌍, 79개 시뮬레이션 작업).
- AHA는 GPT-4o 대비 10.3%, LLaVA-v1.5-13B 대비 43.0% 성능 개선.
- 로봇 정책 성능을 21.4% 향상시키는 실패 피드백 기능 구현 (강화 학습, 작업 계획, 제로샷 경로 생성).
핵심 아이디어
AHA는 로봇 조작 실패를 이진 분류가 아닌 **자유 형식 추론**(free-form reasoning) 문제로 접근하여, 실패 원인에 대한 **자세한 설명**을 생성한다. 이는 기존 모델들이 실패 탐지 능력에 한계를 보이는 문제를 해결하기 위한 핵심 전략이다. AHA는 실패를 탐지할 뿐 아니라, 다양한 로봇, 작업, 환경에 **적응 가능한 설명**을 제공함으로써, 실패를 학습하고 수정하는 데 기여한다. 실패 데이터셋 AHA는 FailGen이라는 **시뮬레이션 기반 실패 생성 프레임워크**를 통해 생성되며, 성공 시연에 **절차적 변형**(procedural perturbation)을 적용해 실패 시나리오를 대규모로 생성한다. 이는 기존 실패 데이터셋의 부족한 양과 다양성을 보완한다.
기술적 접근법
- **FailGen**: 시뮬레이션에서 성공 시연을 기반으로 실패 시나리오를 생성하는 데이터 생성 파이프라인.
- **AHA 모델**: LLaVA-v1.5-13B 기반으로 **instruction fine-tuning**을 통해 실패 추론 능력 향상.
- **실험 설정**: 3개의 로봇 조작 프레임워크(강화 학습, 작업 및 운동 계획, 제로샷 경로 생성)에 통합.
- **성능 평가**: 4개 메트릭, 3개 평가 데이터셋에서 6개의 최신 VLM과 비교.
주요 결과
- AHA는 GPT-4o 대비 10.3%, 6개 모델 평균 대비 35.3% 성능 개선.
- 실패 피드백을 통한 로봇 정책 성능 향상: 21.4% 평균 성공률 증가.
- 실패 탐지 성능: LLaVA-v1.5-13B 대비 43.0% 개선.
- AHA 데이터셋: 49K+ 이미지-쿼리 쌍, 79개 시뮬레이션 작업.
의의 및 한계
AHA는 로봇이 실패를 자연어로 탐지하고 추론할 수 있게 하여, **오픈 월드 환경에서의 실용성**을 높인다. 실패 데이터셋 AHA와 FailGen은 실패 시나리오 생성의 **규모와 다양성**을 확보해, 기존 실패 탐지 연구의 한계를 극복한다. 그러나 AHA는 **훈련 데이터에 포함된 실패 유형에 제한**되어, 더 개방적인 실패 추론 능력 확보가 필요하다. 또한, FailGen은 시뮬레이션 기반 실패 생성에 의존하므로, 실제 환경에서의 실패를 포괄적으로 다루는 데 한계가 있을 수 있다.
실용적 활용
AHA는 로봇이 실패를 인식하고 수정하는 능력을 향상시켜, **자율 로봇 시스템**, **제조업**, **서비스 로봇** 등에서 실용적 활용이 가능하다. 특히, 강화 학습, 작업 계획, 제로샷 경로 생성과 같은 **로봇 정책 최적화**에 실패 피드백을 적용할 수 있어, 실시간 오류 수정과 작업 성공률 향상에 기여할 수 있다.