ChatPaper.aiChatPaper

任務條件式流匹配以實現均衡的多語言文本嵌入適應

Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation

August 6, 2026
作者: Tirth Bhatt, Naren Kumar S, Mayank Singh
cs.AI

摘要

多語言文本嵌入模型通常以單一訓練目標在多樣化任務中進行適配,儘管不同任務需要根本不同的最佳化策略。我們提出任務條件流匹配(TCFM),這是一個多語言嵌入適配框架,它選擇性地將流匹配應用於翻譯任務,同時以與其學習動態更為對齊的目標來最佳化檢索、分類及配對分類任務。TCFM 進一步結合教師引導的表示保留與三階段課程,以實現穩定適配。在 Indic 大規模文本嵌入基準上的評估中,TCFM 達到了新的最先進水準,在多樣化的多語言任務中持續提升嵌入品質,並跨嵌入模型家族進行泛化。我們將在論文被接受後公開釋出程式碼庫與資料集。
English
Multilingual text embedding models are commonly adapted using a single training objective across diverse tasks, despite different tasks requiring fundamentally different optimization strategies. We introduce Task-Conditional Flow Matching (TCFM), a multilingual embedding adaptation framework that selectively applies Flow Matching to translation tasks while optimizing retrieval, classification, and pair-classification tasks with objectives better aligned to their learning dynamics. TCFM further combines teacher-guided representation preservation with a three-stage curriculum to enable stable adaptation. Evaluated on the Indic Massive Text Embedding Benchmark, TCFM establishes a new state-of-the-art, consistently improving embedding quality across a diverse set of multilingual tasks and generalizing across embedding model families. We will publicly release the codebase and datasets upon acceptance of the paper.