IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages

Sahil Deepak Gawande, Mayank Singh

arXiv:2607.23242 · 2026-07-28 공개 · arXiv · PDF

automated-pipeline conversational-ai indic-languages code-mixing multilingual-corpus llm-generation persona-based dialogue-corpus

Abstract

Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms. We present IndicTalk, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversations across 18 language varieties covering 9 Indic languages. The corpus is generated through a fully automated pipeline that combines real-world news grounding, persona-conditioned dialogue generation using multilingual LLMs, and automatic quality validation. Extensive linguistic, automatic, and human evaluations demonstrate that IndicTalk produces fluent, coherent, and naturally code-mixed conversations across both script variants. We will release IndicTalk to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages. The dataset is available at: https://huggingface.co/datasets/LingoIITGN/IndicTalk .

한국어 요약

한 줄 요약

IndicTalk은 인도 언어 대화 자료의 부족을 해결하기 위한 대규모 다국어 코드-믹스 대화 코퍼스이다.

핵심 기여도

핵심 아이디어

인도어 사용자들은 대화 중 자연스럽게 영어와 인도어를 번갈아 사용하며, 이는 로마자와 네이티브 스크립트 두 형태로 나타난다. 기존의 대화 자료는 이러한 코드-믹스 현상을 반영하지 못해, 인도어 대화 모델 개발에 한계가 있었다. IndicTalk은 이 문제를 해결하기 위해, 실제 뉴스 기반의 이벤트를 바탕으로 퍼소나 조건을 적용한 대화를 생성함으로써, 자연스럽고 일관된 코드-믹스 대화를 구축한다. 이는 다국어 LLM을 활용한 자동 생성과 자동 품질 검증을 결합한 방식으로 이루어진다.

기술적 접근법

주요 결과

의의 및 한계

IndicTalk은 인도 언어 대화 자료의 부족을 해결하고, 코드-믹스 대화 모델 개발에 기초 자료를 제공한다. 특히, 로마자 및 네이티브 스크립트를 모두 포함한 대화는 인도어 사용 환경에 적합한 모델 학습에 기여할 수 있다. 그러나 생성된 대화의 문화적, 지역적 다양성은 명시되지 않았으며, 실제 사용자 대화와의 유사도는 추가 연구가 필요하다.

실용적 활용

IndicTalk은 인도 언어를 사용하는 지역에서 대화형 AI, 번역 시스템, 챗봇 개발에 활용 가능하다. 특히, 코드-믹스 대화를 다루는 언어 모델의 성능 향상 및 평가에 기여할 수 있다.