ChatPaper.aiChatPaper

IndicTalk:面向印度语言的大规模基于角色的多语言对话语料库

IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages

July 25, 2026
作者: Sahil Deepak Gawande, Mayank Singh
cs.AI

摘要

大型语言模型(LLMs)已革新了对话式人工智能,但高质量的多语言代码混合对话资源仍然稀缺,尤其是在印度语言场景中——说话者自然地在英语与母语之间切换,且同时使用本土文字和罗马化形式。我们提出IndicTalk,这是目前规模最大的多语言印度语言代码混合对话语料库之一,包含超过1,328,604个基于事件驱动的多轮对话,覆盖9种印度语言的18种语言变体。该语料库通过全自动流水线生成,结合了真实世界新闻事件驱动、基于多语言LLM的角色条件化对话生成以及自动化质量验证。广泛的语言学评估、自动评估及人工评估结果表明,IndicTalk在两种文字变体下均能生成流畅、连贯且自然的代码混合对话。我们将开放IndicTalk数据集,以支持代表性不足的印度语言多语言对话人工智能的开发与评估。数据集获取地址:https://huggingface.co/datasets/LingoIITGN/IndicTalk。
English
Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms. We present IndicTalk, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversations across 18 language varieties covering 9 Indic languages. The corpus is generated through a fully automated pipeline that combines real-world news grounding, persona-conditioned dialogue generation using multilingual LLMs, and automatic quality validation. Extensive linguistic, automatic, and human evaluations demonstrate that IndicTalk produces fluent, coherent, and naturally code-mixed conversations across both script variants. We will release IndicTalk to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages. The dataset is available at: https://huggingface.co/datasets/LingoIITGN/IndicTalk .