VibeVoice-ASR-Streaming Technical Report

Yujie Tu, Zhiliang Peng, Jianwei Yu, Li Dong, Songchen Xu, Yaoyao Chang, Wenhui Wang, Zilong Wang, Zehua Wang, Yan Xia, Jiajun Zhang, Xie Chen, Furu Wei

arXiv:2609.02812 · 2026-09-06 공개 · arXiv · PDF

low-latency end-to-end llm-based inference-code streaming-asr speaker-attribution vibevoice-asr wer-cer

Abstract

Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-to-end models such as VibeVoice-ASR have unified the two tasks within a single model. However, existing unified models still mainly support offline recognition, making it difficult to meet the low-latency requirements of real-time voice assistants and agents. To tackle this issue, we present VibeVoice-ASR-Streaming, one of the first LLM-based end-to-end approaches to streaming speaker-attributed ASR. It interleaves fixed-size audio chunks, a small amount of lookahead audio and previous text. This allows the model to produce ''who said what'' as speech arrives, without a separate diarization stage. For transcription accuracy, our 7B model achieves the lowest average WER/CER across five evaluation sets. For speaker attribution, it achieves the best or tied-best on 12 of 13 evaluation settings. We release the 1.5B and 7B model weights together with inference code.

한국어 요약

한 줄 요약

VibeVoice-ASR-Streaming은 실시간 화자 속성 ASR을 위한 최초의 LLM 기반 스트리밍 모델로, 단일 모델로 ASR과 화자 인식을 통합하고 2.0초 평균 지연을 달성한다.

핵심 기여도

핵심 아이디어

VibeVoice-ASR-Streaming은 기존의 ASR과 화자 인식을 별도 단계로 처리하는 방식을 대체하여, 단일 모델에서 실시간으로 “who said what”을 생성한다. 이는 고정 크기의 오디오 청크와 이전 텍스트, 그리고 제한된 로우룩 오디오를 번갈아 처리하는 방식을 통해 가능하다. 모델은 각 청크에 대해 4개의 레이턴트 프레임(0.5초)을 로우룩으로 사용하며, 이전의 오디오, 텍스트, 화자 정보를 유지해 일관된 화자 인식을 보장한다. 이는 기존의 오프라인 방식과 달리, 실시간 스트리밍 환경에서 화자 인식을 즉시 수행할 수 있게 한다.

기술적 접근법

주요 결과

의의 및 한계

VibeVoice-ASR-Streaming은 실시간 화자 속성 ASR을 위한 단일 모델 접근법을 제시하며, 기존의 별도 단계(예: 화자 클러스터링)를 제거함으로써 시스템 복잡도를 줄이고 지연을 최소화한다. 특히, 7B 모델은 다양한 언어와 환경에서 뛰어난 성능을 보이며, 실용적 적용 가능성을 높인다. 그러나 다국어 지원은 Qwen3-ForcedAligner-0.6B의 언어 범위에 제한되며, 오래 지속되는 다중 화자 중첩 상황에서는 성능이 저하된다. 또한, 최대 8분 길이의 녹음만 지원하며, 긴 녹음은 계산 비용 증가로 인해 제한된다.

실용적 활용

VibeVoice-ASR-Streaming은 실시간 음성 비서, 회의 자동 기록, 멀티유저 음성 인터페이스 등에서 활용 가능하다. 특히, 화자 인식과 텍스트 생성을 동시에 처리하는 능력은 대화형 응용에서 응답 속도와 정확도를 동시에 향상시킬 수 있다. 공개된 1.5B 및 7B 모델 가중치와 vLLM 지원 추론 코드는 연구 및 산업적 활용을 촉진할 것으로 기대된다.