내부 완화, 전역 균형: 비전-언어 전문가 혼합을 위한 기하학 기반 부하 분산
Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts
August 1, 2026
저자: Ziang Wu, Peng Jin, Qishen Yin, Munan Ning, Hao Li, Peizhen Zhang, Li Yuan
cs.AI
초록
비전-언어 MoE 배치(batch)는 서로 다른 수의 이미지 토큰과 텍스트 토큰을 포함한다. 이미지 해상도, 이미지 개수, 타일링, 프롬프트 길이는 모두 이 토큰 혼합 비율을 변화시킨다. 우리는 표준 토큰 수준 Switch 보조 손실을 Std-Aux라고 부른다. Std-Aux는 혼합 부하만 균형화하므로, 특정 혼합 비율에서는 이미지와 텍스트의 큰 부하 오류가 서로 상쇄될 수 있다. 우리의 주 모델에서 동일하게 학습된 라우터는 이미지 해상도에 따라 부하 불균형이 5배 이상 변화한다. 우리는 이미지 및 텍스트 부하 프로파일을 고정하고, 토큰 혼합 비율이 변함에 따른 정확한 부하 곡선을 유도한다. 이미지-텍스트 부하 차이는 토큰 혼합 비율에 대한 민감도를 결정한다. 물리적 전처리 또한 조건부 프로파일을 변경할 수 있으며, 고정 프로파일 하의 법칙은 이러한 변화를 배제한다. 해결책을 설계하기 위해 우리는 라우터 입력 구조를 살펴본다. 이미지와 텍스트는 서로 다른 영역을 차지하며, 시각 토큰은 출처 이미지별로 강하게 군집한다. 모달리티 경계는 이미지와 텍스트에 대한 별도의 손실 항을 도입하도록 동기를 부여하고, 이미지 경계는 이미지당 하나의 균등 가중치 라우팅 인스턴스를 도입하도록 동기를 부여한다. ReBA(Relax Within, Balance Across, 즉 내부는 완화하고 전체적으로 균형을 맞춤)는 이 두 가지 선택을 모두 구현한다. 네 개의 분할 백본에 걸쳐, ReBA는 보고된 모든 벤치마크 입력에서 부하를 낮추면서도 평균 태스크 정확도를 Std-Aux와 비슷한 수준으로 유지한다. 또한 ReBA는 테스트된 범위에서의 평균 부하와 해상도 및 타일링 변화 하에서의 최악의 물리적 부하를 낮춘다. 코드는 https://github.com/ZiangWu-77/ReBA에서 제공된다.
English
Vision-language MoE batches contain different numbers of image and text tokens. Image resolution, image count, tiling, and prompt length all change this token mix. We call the standard token-level Switch auxiliary loss Std-Aux. Std-Aux balances only the mixed load, so large image and text load errors can cancel at one mix. On our main model, the same trained router shows more than a fivefold change in load imbalance across image resolutions. We hold the image and text load profiles fixed and derive the exact load curve as the token mix varies. The image-text load gap controls sensitivity to the token mix. Physical preprocessing can also change the conditional profiles. The fixed-profile law excludes such changes. To design a remedy, we examine the router input structure. Image and text occupy distinct regions, while visual tokens group strongly by source image. The modality boundary motivates separate image and text terms. The image boundary motivates one equal-weight routing instance per image. ReBA, or Relax Within, Balance Across, implements both choices. Across four split backbones, ReBA lowers load on every reported benchmark input while keeping mean task accuracy comparable to Std-Aux. ReBA also lowers average load over the tested range and worst physical load under resolution and tiling shifts. Code is available at https://github.com/ZiangWu-77/ReBA.