ChatPaper.aiChatPaper

タスク条件付きフローマッチングによるバランスの取れた多言語テキスト埋め込み適応

Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation

August 6, 2026
著者: Tirth Bhatt, Naren Kumar S, Mayank Singh
cs.AI

要旨

多言語テキスト埋め込みモデルは、タスクごとに根本的に異なる最適化戦略が必要であるにもかかわらず、多様なタスクにわたって単一の学習目的で適応されることが一般的である。我々は、タスク条件付きフローマッチング(TCFM)を導入する。これは、翻訳タスクにはフローマッチングを選択的に適用し、検索、分類、ペア分類タスクには各自の学習ダイナミクスにより適合した目的関数で最適化する多言語埋め込み適応フレームワークである。TCFMはさらに、教師誘導による表現保持と3段階カリキュラムを組み合わせることで、安定した適応を可能にする。Indic Massive Text Embedding Benchmark上で評価した結果、TCFMは新たな最先端を確立し、多様な多言語タスクにわたって埋め込み品質を一貫して向上させ、埋め込みモデルファミリー間で一般化することを示す。論文採録後、コードベースとデータセットを公開する予定である。
English
Multilingual text embedding models are commonly adapted using a single training objective across diverse tasks, despite different tasks requiring fundamentally different optimization strategies. We introduce Task-Conditional Flow Matching (TCFM), a multilingual embedding adaptation framework that selectively applies Flow Matching to translation tasks while optimizing retrieval, classification, and pair-classification tasks with objectives better aligned to their learning dynamics. TCFM further combines teacher-guided representation preservation with a three-stage curriculum to enable stable adaptation. Evaluated on the Indic Massive Text Embedding Benchmark, TCFM establishes a new state-of-the-art, consistently improving embedding quality across a diverse set of multilingual tasks and generalizing across embedding model families. We will publicly release the codebase and datasets upon acceptance of the paper.