language-models cross-lingual evaluation-methods domain-specific factoid-qa speech-audio spoken-question-answering telugu
Abstract
Question answering has advanced rapidly with large language models, but predominantly for high-resource languages, in both text and spoken settings. Spoken question answering (SQA) benchmark for Telugu remains unexplored, and the reliability of automatic evaluation in this setting remains unquantified. We introduce VākQA, a Telugu SQA benchmark of 2,001 factoid question-answer pairs across six domains, with 2.53 hours of speech audio, bilingual transcriptions, and human-verified reference answers. We first validate evaluation methods against human judgements: Gemini-as-a-judge best approximates human ratings but is non-uniformly strict, while open-weight judges systematically penalize correct Telugu answers that differ in surface form from the reference. Using this validated setup, we benchmark proprietary and open-weight models across input modality, language, and domain. We observe that Telugu phrasing retains cultural specificity that is lost in translation, speech input introduces phonetic confusions that alter question meaning, and cascaded ASR-MT errors compound progressively. VākQA is publicly released.
한국어 요약
한 줄 요약
텔루구어 음성 기반 팩트오이드 QA 벤치마크 VākQA를 제안하고, 평가 방법과 모델 성능을 분석한다.
핵심 기여도
- VākQA: 2,001개의 팩트오이드 QA 쌍과 2.53시간의 음성 데이터를 포함한 텔루구어 SQA 벤치마크 제안.
- Gemini-as-a-judge가 인간 평가를 가장 잘 반영하지만, 일관성 부족.
- 오픈웨이트 평가기는 표면 형태가 다른 정답을 체계적으로 감점.
- ASR-MT 오류가 누적되어 성능 저하 현상 관찰.
핵심 아이디어
기존 QA 연구는 고자원 언어에 집중되어 있으며, 텔루구어와 같은 저자원 언어의 음성 QA 연구는 미비한 상태였다. 본 연구는 텔루구어 음성 QA의 문화적 특수성을 유지하면서도 자동 평가의 신뢰도를 검증하는 데 초점을 맞춘다. 특히, Gemini-as-a-judge와 오픈웨이트 평가기의 한계를 실증적으로 분석함으로써, 향후 평가 프레임워크 개선에 기여한다. 텔루구어 음성 입력에서 발생하는 발음 혼동과 번역-음성 인식 오류의 누적 효과를 명확히 규명한 점이 핵심 통찰이다.
기술적 접근법
- VākQA: 6개 도메인, 2,001개 QA 쌍, 2.53시간 음성, 이중 언어 전사, 인간 검증 정답 포함.
- 평가 방법 검증: Gemini-as-a-judge, 여러 오픈웨이트 모델 사용.
- 입력 모달: 음성, 텍스트.
- 언어: 텔루구어, 영어.
- 도메인: 6개.
주요 결과
- Gemini-as-a-judge가 인간 평가와 가장 유사하지만, 일관된 엄격도를 보이지 않음.
- 오픈웨이트 평가기는 표면 형태가 다른 정답을 체계적으로 감점 (예: 정답률 감소).
- ASR-MT 오류가 누적되어 QA 성능이 점진적으로 저하됨.
의의 및 한계
VākQA는 저자원 언어의 음성 QA 연구에 중요한 기초 자료를 제공하며, 평가 방법의 신뢰도를 실증적으로 검증한 점에서 학술적 의의가 있다. 또한, 텔루구어의 문화적 특수성이 번역 과정에서 손실되는 문제를 드러내어, 다문화적 QA 연구에 기여한다. 그러나 2,001개의 QA 쌍은 대규모 벤치마크로는 한계가 있으며, 보다 다양한 도메인과 사용자 그룹이 포함될 필요가 있다.
실용적 활용
텔루구어 음성 QA는 인도 지역의 저자원 언어 사용자 대상의 음성 비서, 교육 도구, 고객 지원 시스템 등에 활용 가능하다. 특히, 문화적 특수성을 반영한 QA는 지역 사회 맞춤형 서비스 개발에 기여할 수 있다.