OGBench: Benchmarking Offline Goal-Conditioned RL

Seohong Park, Kevin Frans, Benjamin Eysenbach, Sergey Levine

arXiv:2410.20092 · 2026-07-27 공개 · arXiv · PDF

long-horizon-reasoning offline-rl ogbench stochasticity goal-conditioned-rl stitching high-dimensional-inputs

Abstract

Offline goal-conditioned reinforcement learning (GCRL) is a major problem in reinforcement learning (RL) because it provides a simple, unsupervised, and domain-agnostic way to acquire diverse behaviors and representations from unlabeled data without rewards. Despite the importance of this setting, we lack a standard benchmark that can systematically evaluate the capabilities of offline GCRL algorithms. In this work, we propose OGBench, a new, high-quality benchmark for algorithms research in offline goal-conditioned RL. OGBench consists of 8 types of environments, 85 datasets, and reference implementations of 6 representative offline GCRL algorithms. We have designed these challenging and realistic environments and datasets to directly probe different capabilities of algorithms, such as stitching, long-horizon reasoning, and the ability to handle high-dimensional inputs and stochasticity. While representative algorithms may rank similarly on prior benchmarks, our experiments reveal stark strengths and weaknesses in these different capabilities, providing a strong foundation for building new algorithms. Project page: https://seohong.me/projects/ogbench

한국어 요약

한 줄 요약

OGBench는 오프라인 목표 조건화 강화학습 연구를 위한 8개 환경, 85개 데이터셋, 6개 알고리즘 구현을 포함한 새로운 벤치마크이다.

핵심 기여도

핵심 아이디어

오프라인 GCRL은 보상 없이 다양한 행동과 표현을 학습할 수 있는 중요한 문제이지만, 이를 체계적으로 평가할 수 있는 표준 벤치마크가 부족했다. OGBench는 이 문제를 해결하기 위해 설계된 벤치마크로, stitching, long-horizon reasoning, 고차원 입력 처리 능력, 불확실성 대응 능력 등 GCRL 알고리즘의 핵심 능력을 직접적으로 평가할 수 있도록 환경과 데이터셋을 구성했다. 기존 알고리즘은 기존 벤치마크에서는 유사한 성능을 보일 수 있지만, OGBench에서는 그들의 실제 한계가 드러난다.

기술적 접근법

주요 결과

의의 및 한계

OGBench는 오프라인 GCRL 연구에서 체계적인 알고리즘 평가를 가능하게 하며, 새로운 알고리즘 개발의 기반을 제공한다. 특히, 기존 벤치마크에서는 감지되지 않았던 알고리즘의 핵심 능력 차이를 명확히 드러내는 것이 학술적 가치이다. 한계로는 특정 환경이나 데이터셋에 대한 일반화 가능성, 또는 추가 알고리즘의 포함 여부가 언급되지 않았다.

실용적 활용

OGBench는 강화학습 연구자들이 오프라인 GCRL 알고리즘의 실제 성능을 체계적으로 평가하고, 새로운 알고리즘을 설계하는 데 활용할 수 있다. 특히, 다양한 환경과 데이터셋을 통해 실제 적용 시의 복잡성을 반영한 평가가 가능하다.