ChatPaper.aiChatPaper

IndicTalk:一個大規模的、基於人物角色的多語言對話語料庫,適用於印度語言

IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages

July 25, 2026
作者: Sahil Deepak Gawande, Mayank Singh
cs.AI

摘要

大型語言模型(LLMs)已徹底改變了對話式 AI,然而高品質的多語種語碼混合對話資源仍然稀缺,特別是在印度語言中,使用者無論使用原生文字還是羅馬化形式,都經常在英語與母語之間自然切換。我們提出 IndicTalk,這是一個大型多語種印度語碼混合對話語料庫,涵蓋 18 種語言變體(涵蓋 9 種印度語言),包含超過 1,328,604 輪以事件為基礎的多輪對話。該語料庫透過全自動化流程生成,結合真實新聞事件作為基礎、使用多語種 LLM 進行基於個性的對話生成,以及自動化品質驗證。廣泛的語言學、自動化及人工評估顯示,IndicTalk 在兩種文字變體中均能產出流暢、連貫且自然語碼混合的對話。我們將公開 IndicTalk,以支援代表性不足的印度語系多語種對話式 AI 的開發與評估。該資料集可於以下網址取得:https://huggingface.co/datasets/LingoIITGN/IndicTalk。
English
Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms. We present IndicTalk, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversations across 18 language varieties covering 9 Indic languages. The corpus is generated through a fully automated pipeline that combines real-world news grounding, persona-conditioned dialogue generation using multilingual LLMs, and automatic quality validation. Extensive linguistic, automatic, and human evaluations demonstrate that IndicTalk produces fluent, coherent, and naturally code-mixed conversations across both script variants. We will release IndicTalk to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages. The dataset is available at: https://huggingface.co/datasets/LingoIITGN/IndicTalk .