往復一貫性:双方向拡散モデルは自身のロールアウト誤差を予測できる
Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors
August 1, 2026
著者: Alexander Scheinker
cs.AI
要旨
自己回帰モデルは長いロールアウトにわたって誤差を蓄積するが、実運用時にはそれを測定するための正解データが存在しない。我々は、方向フラグによって力学系を時間方向に進めたり逆行させたりする単一の条件付き潜在拡散モデルを学習し、この双方向性が測定不要のテスト時誤差信号を提供することを示す。すなわち、iステップ前進させてからiステップ後退させるとモデルは開始点に戻らなければならないため、往復不一致C_iは観測不可能なロールアウト誤差の自己教師あり代理指標となる。アンサンブルも、ホールドアウトデータも、支配方程式も不要であり、必要となるのは追加のロールアウト1回だけである。我々は、圧縮性磁気流体力学(MHD)、天体物理的乱流放射混合層、および自然顔動画(CelebV-HQ)で検証する。ホールドアウトされたMHD軌道では、C_iはロールアウト誤差をランク付けする(固定深度ではスピアマン0.91〜0.98、同一軌道内では0.69±0.16)。さらに、学習ロールアウトに適合させた単純なキャリブレータは、その誤差の大きさをほぼ名目どおりの被覆率のもとで68%の場合には1.14倍以内、95%の場合には1.29倍以内に予測し、これは深さのみの予測器を1ナット上回り、復号された6つすべての物理場に転移する。同じ信号は、サンプリング分散ベースラインが逆転するまさにその状況で、分布外のオルサグ・タング渦を検出し(AUROC 0.98、深さ10で1.0)、被覆率80%で発生する誤差を15%削減する。これは深さのみのベースラインの3倍の改善である。双方向学習は負のコストで実現され、両方向でそれぞれに特化したモデルを凌駕し、逆方向の推論は高速逆ソルバーを兼ねる。LE-PDE-UQの乱流ナビエ・ストークスベンチマークでは、単一の双方向モデルが、10モデルアンサンブルの精度の1.3倍以内を学習コスト10分の1で達成し、さらに学習不要のピクセルレベルキャリブレーションとしても最良である。往復整合性は、可逆性を生成モデルにとって実用的な信頼シグナルへと変える。
English
Autoregressive models accumulate error over long rollouts, yet at deployment there is no ground truth to measure it against. We train a single conditional latent diffusion model that steps a dynamical system forward or backward in time via a direction flag, and show that this bidirectionality supplies a measurement-free test-time error signal: rolling forward i steps and then backward i steps must return the model to its start, so the round-trip discrepancy C_i is a self-supervised proxy for the unobservable rollout error: no ensembles, no held-out data, no governing equations, for one extra rollout. We validate on compressible magnetohydrodynamics (MHD), an astrophysical turbulent radiative mixing layer, and natural face videos (CelebV-HQ). On held-out MHD trajectories, C_i ranks rollout error (Spearman 0.91-0.98 at fixed depth; 0.69 pm 0.16 within trajectories), and a simple calibrator fit on training rollouts predicts its magnitude to within 1.14times (68%) and 1.29times (95%) with near-nominal coverage - one nat beyond a depth-only predictor, transferring to all six decoded physical fields. The same signal flags the out-of-distribution Orszag-Tang vortex (AUROC 0.98; 1.0 by depth 10) exactly where sampling-dispersion baselines invert, and it cuts incurred error by 15% at 80% coverage - three times the depth-only baseline. Bidirectional training comes at negative cost, beating direction specialists in both directions, and the backward direction doubles as a fast inverse solver. On LE-PDE-UQ's turbulent Navier-Stokes benchmark, a single bidirectional model reaches accuracy within 1.3times of their ten-model ensemble at a tenth of the training cost, with the best training-free pixel-level calibration. Round-trip consistency turns reversibility into a practical trust signal for generative models.