autonomous-agents embodied-agents vlms progress-estimation sequence-tracking recurrence-disambiguation state-recall context-dependent
Abstract
Embodied agents now take on ever longer tasks. For long tasks, knowing only whether a task finally succeeds or fails says little; the steps along the way matter. Progress Reward Models (PRMs) score how far a task has come at every step, and serve as dense rewards, verifiers and monitors. Yet in long tasks the current frame alone often cannot tell how far the task has come, because progress depends on what happened before. We call this problem context-dependent progress estimation. Existing benchmarks on progress estimation mostly focus on short tasks whose progress can be read from the current observation, and whether PRMs can estimate progress when context is needed remains underexplored. We therefore build ContextProgress-Bench, with 24 manipulation tasks for 120 episodes. The benchmark covers three settings: (i) State Recall, where information needed for progress appeared earlier but is not in the current frame; (ii) Sequence Tracking, where steps follow a fixed order, so progress requires knowing which steps are done and which comes next; and (iii) Recurrence Disambiguation, where look-alike frames sit at very different progress. We then run a paired diagnosis: each PRM keeps the same input format in both runs, and in one run its instruction integrates the right context. Even PRMs that read the entire history get lost in estimating progress, yet with the right context the same five models cut their progress error by 77-82%. Embodied PRMs are thus not incapable of progress estimation, but lost without the right context. We therefore propose ProgressCompass, an autonomous agentic loop that reorients an existing PRM and uses current general-purpose VLMs to supply the context the PRM needs. Wrapped in the loop, the same frozen PRM cuts its progress error by 63% and raises its rank agreement by 76%. With such a compass, PRMs estimate progress far better on longer, more complex tasks.
한국어 요약
한 줄 요약
ProgressCompass는 적절한 컨텍스트를 제공함으로써, 복잡한 작업에서 Progress Reward Model(PRMs)의 정확도를 63% 개선하는 자율 에이전트 루프를 제안한다.
핵심 기여도
- ContextProgress-Bench라는 24개 조작 작업, 120개 에피소드를 포함한 새로운 벤치마크를 제안.
- 기존 PRMs가 전체 히스토리를 사용하더라도 컨텍스트 없이는 77–82%의 오류를 보임.
- ProgressCompass는 기존 PRM에 컨텍스트를 제공함으로써 오류를 63% 감소시키고 순위 일치도를 76% 향상.
- ProgressCompass는 추가 학습 없이 기존 VLM을 활용해 컨텍스트를 제공.
핵심 아이디어
기존 PRMs는 단일 프레임만 기반으로 작업 진행도를 평가하지만, 긴 작업에서는 이전 정보(컨텍스트)가 없으면 정확한 평가가 어렵다. 예를 들어, State Recall에서는 이전 상태를 기억해야 하며, Sequence Tracking에서는 단계 순서를 추적해야 하고, Recurrence Disambiguation에서는 유사한 프레임이 다른 진행도를 가질 수 있다. 이에 따라, 컨텍스트 없이 PRM이 정확한 평가를 하기 어렵다는 점이 드러났다. ProgressCompass는 PRM이 필요로 하는 컨텍스트를 Vision-Language Model(VLM)을 통해 자동으로 제공하는 에이전트 루프를 설계함으로써, 기존 PRM의 한계를 극복한다.
기술적 접근법
- **ContextProgress-Bench**: 24개 조작 작업, 120개 에피소드, 552개의 서브태스크 구간을 포함.
- **세 가지 설정**: (i) State Recall, (ii) Sequence Tracking, (iii) Recurrence Disambiguation.
- **ProgressCompass 구성**: Orienter(현재 단계 식별), Verifier(단계 완료 여부 확인), Navigator(루프 실행).
- **모델 구성**: Qwen3.5-27B(Orienter), Qwen3.5-9B(Verifier), RoboMeter-4B(동결 PRM).
- **실험 설정**: 동일한 PRM이 동일한 입력 형식을 사용하며, 하나는 올바른 컨텍스트를 포함한 지시문을 받음.
- **성능 지표**: 평균 절대 오차(MAE), 순위 일치도(rank agreement).
주요 결과
- 기존 PRMs는 전체 히스토리를 사용하더라도 컨텍스트 없이는 MAE가 15.6–31.0 (0–100 스케일).
- 올바른 컨텍스트가 주어지면 MAE가 77–82% 감소하여 3.4–6.7로 감소.
- ProgressCompass를 적용한 동일 PRM은 MAE를 63% 감소시키고, 순위 일치도를 76% 향상.
- ProgressCompass는 조기 종료(Early stop), 추가 단계(Extra steps), 태스크 불일치(Mismatch)에도 강건.
의의 및 한계
ProgressCompass는 복잡한 작업에서 PRM의 정확도를 크게 향상시키며, 컨텍스트가 필요한 작업에서 기존 PRM의 한계를 보완하는 새로운 접근법을 제시한다. 특히, VLM을 활용해 학습 없이 컨텍스트를 제공함으로써, 기존 모델을 재사용하면서도 성능을 향상시킬 수 있다는 점에서 실용적 가치가 있다. 그러나, ProgressCompass는 VLM의 질에 의존하며, 특정 작업에서는 컨텍스트 추출이 여전히 어려울 수 있다. 또한, 복잡한 작업에서의 일반화 능력은 추가 실험을 통해 검증되어야 한다.
실용적 활용
ProgressCompass는 로봇 자율 실행, 작업 모니터링, 복잡한 작업의 단계별 검증 등에서 활용 가능하다. 특히, 컨텍스트가 필요한 장기 작업에서 PRM의 정확도를 향상시켜, 작업 완료 여부 판단, 진행도 추적, 에러 감지 등에 유용하게 사용될 수 있다.