IndicTalk: インド諸言語向けの大規模パーソナベース多言語会話コーパス
IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages
July 25, 2026
著者: Sahil Deepak Gawande, Mayank Singh
cs.AI
要旨
大規模言語モデル(LLMs)は対話型AIを変革したが、高品質な多言語コードミックス対話リソースは依然として不足しており、特に話者が原語表記とローマ字表記の両方で自然に英語と母語を切り替えるインド系言語では顕著である。本稿では、9つのインド系言語にわたる18の言語変種において、132万8604件以上のイベント基盤マルチターン対話からなる、最大級の多言語インド系コードミックス対話コーパスであるIndicTalkを紹介する。このコーパスは、実世界のニュース基盤、多言語LLMを用いたペルソナ条件付き対話生成、および自動品質検証を組み合わせた完全自動化パイプラインによって生成される。広範な言語学的評価、自動評価、人間による評価により、IndicTalkが両方の表記変種において流暢で一貫性があり自然なコードミックス対話を生成することが示された。IndicTalkを公開し、過小評価されているインド系言語向けの多言語対話型AIの開発と評価を支援する。データセットは https://huggingface.co/datasets/LingoIITGN/IndicTalk で入手可能である。
English
Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms. We present IndicTalk, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversations across 18 language varieties covering 9 Indic languages. The corpus is generated through a fully automated pipeline that combines real-world news grounding, persona-conditioned dialogue generation using multilingual LLMs, and automatic quality validation. Extensive linguistic, automatic, and human evaluations demonstrate that IndicTalk produces fluent, coherent, and naturally code-mixed conversations across both script variants. We will release IndicTalk to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages. The dataset is available at: https://huggingface.co/datasets/LingoIITGN/IndicTalk .