ChatPaper.aiChatPaper

往返一致性:雙向擴散模型能預測自身的展開誤差

Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors

August 1, 2026
作者: Alexander Scheinker
cs.AI

摘要

自迴歸模型在長程展開中會累積誤差,然而部署時並無真實值可供比對衡量。我們訓練單一條件潛在擴散模型,透過方向標記使動力系統在時間上向前或向後步進,並證明此雙向性提供了無需量測的測試時誤差訊號:向前展開 i 步再向後展開 i 步必須將模型帶回起點,因此往返差異 C_i 即為不可觀測展開誤差的自監督代理指標:無需集成、無需留出數據、無需控制方程,只需一次額外展開。我們在可壓縮磁流體動力學(MHD)、天文物理湍流輻射混合層以及自然人臉影片(CelebV-HQ)上進行驗證。在留出的 MHD 軌跡上,C_i 能對展開誤差排序(固定深度下 Spearman 0.91–0.98;軌跡內 0.69 ± 0.16);在訓練展開上擬合的簡單校正器,能預測其量級,68% 與 95% 的情況下分別落在 1.14 倍與 1.29 倍以內,且覆蓋率接近名義值——比僅以深度為輸入的預測器多出一個自然對數單位(nat),並可遷移至全部六個解碼物理場。同一訊號在分布外的 Orszag-Tang 渦流上發出預警(AUROC 0.98;深度 10 時達 1.0),而正是在此處採樣分散度基線的表現發生反轉;該訊號在 80% 覆蓋率下將產生的誤差降低 15%——為僅深度基線的三倍。雙向訓練具有負成本,在兩個方向上均勝過單向專家模型,且反向過程可兼作快速逆向求解器。在 LE-PDE-UQ 的湍流 Navier-Stokes 基準上,單一雙向模型以十分之一的訓練成本達到其十模型集成精度的 1.3 倍以內,並提供最佳的免訓練像素級校準。往返一致性將可逆性轉化為生成模型的實用信任訊號。
English
Autoregressive models accumulate error over long rollouts, yet at deployment there is no ground truth to measure it against. We train a single conditional latent diffusion model that steps a dynamical system forward or backward in time via a direction flag, and show that this bidirectionality supplies a measurement-free test-time error signal: rolling forward i steps and then backward i steps must return the model to its start, so the round-trip discrepancy C_i is a self-supervised proxy for the unobservable rollout error: no ensembles, no held-out data, no governing equations, for one extra rollout. We validate on compressible magnetohydrodynamics (MHD), an astrophysical turbulent radiative mixing layer, and natural face videos (CelebV-HQ). On held-out MHD trajectories, C_i ranks rollout error (Spearman 0.91-0.98 at fixed depth; 0.69 pm 0.16 within trajectories), and a simple calibrator fit on training rollouts predicts its magnitude to within 1.14times (68%) and 1.29times (95%) with near-nominal coverage - one nat beyond a depth-only predictor, transferring to all six decoded physical fields. The same signal flags the out-of-distribution Orszag-Tang vortex (AUROC 0.98; 1.0 by depth 10) exactly where sampling-dispersion baselines invert, and it cuts incurred error by 15% at 80% coverage - three times the depth-only baseline. Bidirectional training comes at negative cost, beating direction specialists in both directions, and the backward direction doubles as a fast inverse solver. On LE-PDE-UQ's turbulent Navier-Stokes benchmark, a single bidirectional model reaches accuracy within 1.3times of their ten-model ensemble at a tenth of the training cost, with the best training-free pixel-level calibration. Round-trip consistency turns reversibility into a practical trust signal for generative models.