ChatPaper.aiChatPaper

任务条件流匹配用于平衡多语言文本嵌入适配

Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation

August 6, 2026
作者: Tirth Bhatt, Naren Kumar S, Mayank Singh
cs.AI

摘要

多语言文本嵌入模型通常在不同任务上使用单一训练目标进行适配,尽管不同任务需要截然不同的优化策略。我们提出了任务条件流匹配(Task-Conditional Flow Matching, TCFM),这是一种多语言嵌入适配框架,它选择性地将流匹配应用于翻译任务,同时使用与检索、分类和成对分类任务学习动态更契合的目标来优化这些任务。TCFM 进一步结合了教师引导的表示保持与三阶段课程,以实现稳定的适配。在印度语系大规模文本嵌入基准(Indic Massive Text Embedding Benchmark)上的评估表明,TCFM 取得了新的最先进成果,在多种多语言任务中持续提升嵌入质量,并能跨嵌入模型家族进行泛化。论文被接收后,我们将公开发布代码库和数据集。
English
Multilingual text embedding models are commonly adapted using a single training objective across diverse tasks, despite different tasks requiring fundamentally different optimization strategies. We introduce Task-Conditional Flow Matching (TCFM), a multilingual embedding adaptation framework that selectively applies Flow Matching to translation tasks while optimizing retrieval, classification, and pair-classification tasks with objectives better aligned to their learning dynamics. TCFM further combines teacher-guided representation preservation with a three-stage curriculum to enable stable adaptation. Evaluated on the Indic Massive Text Embedding Benchmark, TCFM establishes a new state-of-the-art, consistently improving embedding quality across a diverse set of multilingual tasks and generalizing across embedding model families. We will publicly release the codebase and datasets upon acceptance of the paper.