ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation

Jiawen Yu, Hairuo Liu, Qiaojun Yu, Jieji Ren, Ce Hao, Haitong Ding, Guan Huang, Guofan Huang, Yan Song, Panpan Cai, Cewu Lu, Wenqiang Zhang

arXiv:2505.22159 · 2026-07-27 공개 · arXiv · PDF

vision-language-action mixture-of-experts robotic-manipulation dexterous-manipulation proprioception multimodal-integration force-sensing contact-rich-tasks

Abstract

Vision-Language-Action (VLA) models have advanced general-purpose robotic manipulation by leveraging pretrained visual and linguistic representations. However, they struggle with contact-rich tasks that require fine-grained control involving force, especially under visual occlusion or dynamic uncertainty. To address these limitations, we propose ForceVLA, a novel end-to-end manipulation framework that treats external force sensing as a first-class modality within VLA systems. ForceVLA introduces FVLMoE, a force-aware Mixture-of-Experts fusion module that dynamically integrates pretrained visual-language embeddings with real-time 6-axis force feedback during action decoding. This enables context-aware routing across modality-specific experts, enhancing the robot's ability to adapt to subtle contact dynamics. We also introduce \textbf{ForceVLA-Data}, a new dataset comprising synchronized vision, proprioception, and force-torque signals across five contact-rich manipulation tasks. ForceVLA improves average task success by 23.2% over strong pi_0-based baselines, achieving up to 80% success in tasks such as plug insertion. Our approach highlights the importance of multimodal integration for dexterous manipulation and sets a new benchmark for physically intelligent robotic control. Code and data will be released at https://sites.google.com/view/forcevla2025.

한국어 요약

한 줄 요약

ForceVLA는 외부 힘을 1등급 모달로 통합한 VLA 모델로, 접촉이 많은 조작 작업 성능을 23.2% 개선한다.

핵심 기여도

핵심 아이디어

기존 VLA 모델은 시각과 언어 정보를 기반으로 조작을 수행하지만, 힘 정보를 충분히 활용하지 못해 접촉이 많은 작업에서 한계가 있었다. ForceVLA는 외부 힘을 1등급 모달로 취급하고, 이를 VLA 시스템에 통합하는 새로운 프레임워크를 제안한다. 핵심 기술은 FVLMoE라는 Mixture-of-Experts 모듈로, 이는 VLM 기반 시각-언어 임베딩과 실시간 6축 힘 피드백을 동적으로 통합한다. FVLMoE는 게이팅 메커니즘을 통해 작업 단계에 따라 전문가 서브네트워크를 적응적으로 활성화하여, 정교한 힘 제어를 가능하게 한다. 이는 작업 상황에 따라 다른 힘 조절이 필요한 상황에서 특히 효과적이다.

기술적 접근법

주요 결과

의의 및 한계

ForceVLA는 VLA 모델에 힘 정보를 체계적으로 통합함으로써, 접촉이 많은 작업에서의 정밀성과 안정성을 크게 향상시켰다. FVLMoE 모듈은 힘 정보를 기반으로 한 상황 인식 및 조작 전략의 적응성을 강화하며, ForceVLA-Data는 이 분야의 연구를 위한 새로운 기준을 제시한다. 그러나 현재는 추정된 외부 와트치 값을 사용하고 있어, 고정밀 힘 측정이 필요한 작업에서는 한계가 있다. 또한, 고가의 힘-토크 센서가 내장된 로봇 플랫폼에서만 검증되었기 때문에, 저비용 플랫폼에서의 확장 가능성도 한계로 제기된다.

실용적 활용

ForceVLA는 산업 로봇에서의 정밀 조립, 의료 로봇의 부드러운 조작, 서비스 로봇의 물체 처리 등 다양한 분야에서 적용 가능하다. 특히, 시야가 제한되거나 동적 불확실성이 높은 환경에서 힘 정보를 기반으로 한 정밀한 제어가 필요한 작업에 유용하다.