Any-OPD: 表現空間ブリッジングによるフローマッチングモデル向け異種オンポリシー蒸留
Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging
August 4, 2026
著者: Siming Fu, Zheming Fu, Ruizhe He, Hualiang Wang, Jie Huang, Xiaoxiao Ma, Mingchen Zhong, Weihu Huang, Xiaoxuan He, Haojun Xu
cs.AI
要旨
オン方策蒸留は、教師が生徒自身が生成したサンプルを修正する手法であるが、そこでは二つのモデルが同じ言語を話すこと、すなわち同一のVAE潜在変数、一致するアーキテクチャ、共通のタイムステップグリッドを持つことが前提となる。我々は、これらが一つも成立しない場合、例えば利用可能な最強の教師と展開したい生徒が異なるモデルファミリーに由来する場合に何が起こるのかを問い、標準的な手法には答えがないことを見出す。教師の潜在変数は異なる座標系ではターゲットとして機能できず、局所的な詳細を確率的に再描画する教師に対する画素単位の損失はぼやけや発散に退化し、タイムステップインデックスは不一致のスケジュール間では意味を失う。我々はAny-OPDを提案する。これは、我々の知る限り、任意の潜在フローマッチング生成器ペア間でのオン方策蒸留のための最初のフレームワークである。Any-OPDは教師を純粋にブラックボックスのサンプラーとして扱い、二つのモデルをちょうど一点でのみ接続する。すなわち、凍結されモデル非依存な視覚表現の中で、独立に復号された出力同士を比較するのであり、潜在変数、特徴量、アーキテクチャに関するあらゆる仮定を回避する。軌道の対応関係は、ステップインデックスではなく連続的なノイズレベルをマッチングすることによって回復される。さらに、教師のサンプルを生徒自身のVAEを通して再符号化する短いアンカリングフェーズにより、オン方策勾配がドメインミスマッチではなくサンプル品質を測定することが保証される。12BのFLUX.1-devを2.5BのSD3.5-Mediumへ蒸留する際、Any-OPDは生徒のPickScoreを0.846から0.884へ、HPSv3を9.12から10.97へと引き上げ、直接的な潜在回帰ではまったく学習が進行しない状況において、教師の5分の1の規模でありながら教師に匹敵する性能を達成する。
English
On-policy distillation, in which a teacher corrects samples that the student itself generates, presupposes that the two models speak the same language: identical VAE latents, matching architectures, and a common timestep grid. We ask what happens when none of this holds, as when the strongest teacher available and the student one wishes to deploy come from different model families, and find that the standard recipes have no answer: teacher latents cannot serve as targets in a foreign coordinate system, per-pixel losses against a teacher that stochastically re-draws local detail degenerate into blur or divergence, and timestep indices lose their meaning across mismatched schedules. We present Any-OPD, to our knowledge the first framework for on-policy distillation between arbitrary pairs of latent flow-matching generators. Any-OPD treats the teacher purely as a black-box sampler and connects the two models at exactly one point: a frozen, model-agnostic vision representation in which their independently decoded outputs are compared, sidestepping every assumption about latents, features, or architecture. Trajectory correspondence is recovered by matching continuous noise levels instead of step indices, and a brief anchoring phase, in which teacher samples are re-encoded through the student's own VAE, ensures the on-policy gradient measures sample quality rather than domain mismatch. Distilling the 12B FLUX.1-dev into the 2.5B SD3.5-Medium, Any-OPD lifts the student's PickScore from 0.846 to 0.884 and HPSv3 from 9.12 to 10.97, rivaling the teacher at a fifth of its size, where direct latent regression fails to train at all.