IndicTalk: 인도 언어를 위한 대규모 페르소나 기반 다국어 대화 말뭉치
IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages
July 25, 2026
저자: Sahil Deepak Gawande, Mayank Singh
cs.AI
초록
대규모 언어 모델(LLMs)은 대화형 AI를 혁신적으로 변화시켰지만, 특히 화자들이 영어와 모국어를 자국어 문자 및 로마자 표기 형태로 자연스럽게 혼용하는 인도 언어의 경우 고품질 다국어 코드 혼합 대화 자원은 여전히 부족한 실정이다. 본 논문에서는 9개의 인도 언어를 아우르는 18개 언어 변종에 걸쳐 1,328,604개 이상의 사건 기반 다중 턴 대화로 구성된, 가장 큰 규모의 다국어 인도 코드 혼합 대화 코퍼스 중 하나인 IndicTalk을 제시한다. 이 코퍼스는 실제 뉴스 기반의 근거 제공, 다국어 LLM을 활용한 페르소나 조건부 대화 생성, 그리고 자동 품질 검증을 결합한 완전 자동화 파이프라인을 통해 생성되었다. 광범위한 언어학적, 자동 및 인간 평가를 통해 IndicTalk이 두 문자 변종 모두에서 유창하고 일관되며 자연스러운 코드 혼합 대화를 생성함을 입증한다. 과소 대표된 인도 언어를 위한 다국어 대화형 AI의 개발 및 평가를 지원하기 위해 IndicTalk을 공개할 예정이다. 데이터셋은 다음 링크에서 확인할 수 있다: https://huggingface.co/datasets/LingoIITGN/IndicTalk
English
Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms. We present IndicTalk, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversations across 18 language varieties covering 9 Indic languages. The corpus is generated through a fully automated pipeline that combines real-world news grounding, persona-conditioned dialogue generation using multilingual LLMs, and automatic quality validation. Extensive linguistic, automatic, and human evaluations demonstrate that IndicTalk produces fluent, coherent, and naturally code-mixed conversations across both script variants. We will release IndicTalk to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages. The dataset is available at: https://huggingface.co/datasets/LingoIITGN/IndicTalk .