Video2Skill: From Streaming Experience to Reusable Embodied Skills

Jianshu Zhang, Ce Zhang, Xiyuan Yang, Chenwei Xu, Haoran Lu, Yijiang Li, Yaqi Xie, Katia P. Sycara, Han Liu

arXiv:2609.36691 · 2026-10-06 공개 · arXiv · PDF

vision-language-models robot-manipulation embodied-agents reusable-skills skill-library skill-discovery video2skill streaming-embodied

Abstract

Manipulation behaviors vary widely across objects and scenes, but they share a small set of reusable skills, and planning with these skills helps embodied agents generalize to new tasks. Yet an agent can only plan with skills it knows. Recovering skills from observed experience, the inverse of planning, builds this knowledge over time and yields skill data for training future agents. Vision-Language Models (VLMs) describe individual manipulation events well, but can they organize a stream of events into reusable skills? We formulate this problem as Streaming Embodied Skill Discovery (SESD): a model watches videos in sequence and maintains a persistent skill library that shapes its later decisions. To systematically measure this ability, we introduce Video2Skill, a benchmark that covers robot tabletop manipulation and human kitchen activity and tests three core capabilities: (i) locating manipulation events in time, (ii) grouping events of the same transformation, and (iii) deciding when to reuse an existing skill or create a new one. Across 19 open-source VLMs, many models group events at near-chance level, and scale does not consistently help. Their errors depend on how perception and library updates are coupled: joint models merge distinct transformations into one skill, while models that update the library from text descriptions duplicate recurring ones. Supervised fine-tuning, including our counterfactual library-state rebalancing (CLaRe), improves grouping but exposes a deeper bottleneck: trained models consolidate familiar skills yet rarely expand the library. Their libraries stall below half the reference size, and transformations unseen in training are located in time but almost never given a new skill. Recognizing when existing skills are insufficient thus emerges as the central challenge.

한국어 요약

한 줄 요약

Video2Skill은 스트리밍 영상에서 재사용 가능한 신체 기반 스킬을 발견하는 기준 평가 프레임워크로, 19개의 VLM 모델을 통해 스킬 그룹화 및 재사용 결정의 한계를 분석한다.

핵심 기여도

핵심 아이디어

기존 VLM은 개별 조작 행위를 설명하는 데 효과적이지만, **연속적인 영상 스트림을 재사용 가능한 스킬 라이브러리로 조직하는 능력은 부족**하다. 이 연구는 **Streaming Embodied Skill Discovery (SESD)** 문제를 제안하며, 모델이 **비디오를 순차적으로 처리하면서 지속적인 스킬 라이브러리를 유지**하고, 이를 바탕으로 **새로운 스킬을 생성하거나 기존 스킬을 재사용하는 결정을 내리는 능력**을 평가한다. 핵심 통찰은 **스킬 추출과 라이브러리 업데이트가 결합된 방식에 따라 성능이 크게 달라진다는 점**이다. 예를 들어, **통합 모델**(unified processing)은 서로 다른 변형을 하나의 스킬로 병합하는 경향이 있고, **분리 모델**(factorized processing)은 반복되는 변형에 대해 중복된 스킬을 생성한다.

기술적 접근법

주요 결과

의의 및 한계

실용적 활용