CARE: Certifying Acceleration for Vision-Language-Action Inference

Rui Liu, Tong Zheng, Jindong Gu, Zhipeng Wang

arXiv:2610.08917 · 2026-10-11 공개 · arXiv · PDF

vision-language-action qwen3-5-9b llama-3-1-8b sequential-testing rollout-evaluation compute-budgeting acceleration-certification policy-failure

Abstract

While vision-language-action (VLA) models have advanced rapidly, running them at every control step remains expensive. Prior work accelerates VLA inference using techniques like action chunking and visual-token pruning, typically evaluating based on latency and average task success. However, acceleration may discard information and break tasks the original policy would solve, a risk hidden by average metrics. Measuring these failures is challenging because action deviations compound over closed-loop trajectories, meaning task failure is only observable across full episodes. We therefore define an acceleration-induced failure via paired rollouts from identical initial conditions, tracking when the reference succeeds but the accelerated policy fails. To manage this, we introduce CARE, an approach for certified accelerator selection. CARE uses paired rollouts on a calibration set to provide finite-sample guarantees that acceleration-induced failure risk stays below a user-specified budget. It deploys the fastest certified candidate, falling back to the reference if none qualify. By relying only on terminal outcomes and measured compute, CARE applies unchanged across diverse acceleration mechanisms, while sequential testing and failure-triggered reference rollouts keep certification affordable. On four LIBERO suites with OpenVLA-OFT, CARE certifies 9.0--10.8times speedups while guaranteeing (at 95% confidence) that at least 85.8% of reference-solved episodes are preserved. Under tight budgets, selectors without guarantees exceed the budget in up to 75% of trials, whereas CARE stays within budget and its sequential form uses 78.9% fewer rollouts than exhaustive evaluation. CARE further generalizes to flow-step reduction for π_{0.5}, and to Qwen3.5-9B and Llama-3.1-8B agents in Crafter.

한국어 요약

한 줄 요약

CARE는 VLA 모델 가속화 시 발생할 수 있는 실패를 인증하면서 최대 10.8× 가속을 보장하는 정량적 접근법이다.

핵심 기여도

핵심 아이디어

CARE는 VLA 모델의 가속화가 실제 작업 성공률에 미치는 영향을 정량적으로 평가하기 위해 "가속화 유발 실패"라는 개념을 도입한다. 이는 참조 정책이 성공했지만, 가속화된 정책이 실패한 경우를 의미하며, 이는 단일 에피소드에서만 발생하는 것이 아니라 전체 클로즈드-루프 롤아웃을 통해 측정되어야 한다. CARE는 이 이벤트를 인증하기 위해 참조 정책과 가속화 정책의 쌍을 사용한 롤아웃을 수행하며, 유한 표본 기반의 통계적 보장을 통해 실패율이 사용자 지정 예산 이하로 유지되도록 보장한다. 핵심 아이디어는 "Learn-then-Test" 전략과 순차적 테스트를 결합하여 인증 비용을 최소화하면서도 정확도를 유지하는 것이다.

기술적 접근법

주요 결과

의의 및 한계

CARE는 VLA 모델의 가속화가 단순히 속도-정확도 트레이드오프가 아닌, 실패율을 명시적으로 제어할 수 있는 새로운 인증 프레임워크를 제시한다. 기존 평균 성공률 기반 평가 방식의 한계를 극복하고, 실제 작업 실패를 측정하는 새로운 메트릭을 도입한 점에서 학술적 의의가 있다. 또한, 다양한 가속화 메커니즘(액션 청킹, 시각 토큰 제거, 피처 캐싱 등)에 적용 가능한 모델-agnostic 접근법이라는 점에서 실용적 가치가 있다. 그러나 CARE는 참조 정책의 클로즈드-루프 롤아웃이 필수적이며, 이는 여전히 높은 계산 비용을 수반한다.

실용적 활용

CARE는 로봇 제어, 자율 주행, 멀티모달 인터페이스 등에서 VLA 모델의 실시간 실행이 필요한 상황에 적용 가능하다. 특히, 정책의 성능 저하를 최소화하면서도 가속화를 보장해야 하는 산업 현장에서 유용하다. 또한, 다양한 가속화 기법을 통합적으로 평가할 수 있어 연구 개발 과정에서도 활용 가능하다.