균형 잡힌 다국어 텍스트 임베딩 적응을 위한 태스크 조건부 플로우 매칭
Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation
August 6, 2026
저자: Tirth Bhatt, Naren Kumar S, Mayank Singh
cs.AI
초록
다국어 텍스트 임베딩 모델은 다양한 작업이 근본적으로 다른 최적화 전략을 요구함에도 불구하고, 일반적으로 다양한 작업에 걸쳐 단일 학습 목적 함수를 사용하여 적응된다. 본 논문에서는 다국어 임베딩 적응 프레임워크인 Task-Conditional Flow Matching(TCFM)을 제안한다. TCFM은 번역 작업에는 Flow Matching을 선택적으로 적용하고, 검색, 분류, 쌍 분류 작업에는 해당 학습 역학에 더 적합한 목적 함수를 사용하여 최적화한다. 또한, TCFM은 교사 기반 표현 보존(teacher-guided representation preservation)을 3단계 커리큘럼과 결합하여 안정적인 적응을 가능하게 한다. Indic Massive Text Embedding Benchmark에서 평가한 결과, TCFM은 다양한 다국어 작업 전반에 걸쳐 임베딩 품질을 일관되게 개선하고 여러 임베딩 모델 군에 일반화되며 새로운 최첨단 성능을 달성한다. 본 논문이 채택되면 코드베이스와 데이터셋을 공개할 예정이다.
English
Multilingual text embedding models are commonly adapted using a single training objective across diverse tasks, despite different tasks requiring fundamentally different optimization strategies. We introduce Task-Conditional Flow Matching (TCFM), a multilingual embedding adaptation framework that selectively applies Flow Matching to translation tasks while optimizing retrieval, classification, and pair-classification tasks with objectives better aligned to their learning dynamics. TCFM further combines teacher-guided representation preservation with a three-stage curriculum to enable stable adaptation. Evaluated on the Indic Massive Text Embedding Benchmark, TCFM establishes a new state-of-the-art, consistently improving embedding quality across a diverse set of multilingual tasks and generalizing across embedding model families. We will publicly release the codebase and datasets upon acceptance of the paper.