ChatPaper.aiChatPaper

DiFA: 拡散モデルのための推論時前方過程アライメント

DiFA: Inference-Time Forward-Process Alignment for Diffusion Models

July 20, 2026
著者: Shigui Li, Delu Zeng
cs.AI

要旨

拡散モデルの一般的な推論フレームワークは、生成を基本的に数値積分の問題として定式化している。この見方は、モデルを正確な推定量と見なし、ノイズ除去プロセスに内在する統計的不確実性を無視している。本研究では、Forward-Process Aligned Diffusion prediction(DiFA)を提案する。これは訓練不要のフレームワークであり、推論時のデータ予測の精緻化を逐次状態推定問題として再定義する。過去の出力を単に数値積分に再利用するのではなく、DiFAは逆方向軌道に沿った反復的なデータ予測を相関のある観測として扱い、前方方向に整合した時間的コンセンサスを構築する。カルマンフィルタリングに着想を得て、このコンセンサスは構造的一貫性とノイズレベルの互換性に従って過去の予測を集約する。時間的コンセンサスが過度の平滑化に陥る傾向に対抗するため、偏差ガイダンスメカニズムを導入し、残差詳細を適応的に保持する。実験的に、DiFAはCIFAR-10とImageNetにおいて、FID、IS、FD-DINOv2を含む評価指標全体で顕著な改善を示し、推論を前方の統計的構造と整合させることで生成の忠実度が大幅に向上することを実証している。
English
The prevailing inference framework for diffusion models formulates generation fundamentally as a problem of numerical integration. This perspective casts the model as an exact estimator, neglecting the inherent statistical uncertainty of the denoising process. In this work, we propose Forward-Process Aligned Diffusion prediction (DiFA), a training-free framework that reframes inference-time data prediction refinement as a sequential state estimation problem. Rather than reusing past outputs solely for numerical integration, DiFA treats iterative data predictions along the reverse trajectory as correlated observations to build a forward-aligned temporal consensus. Inspired by Kalman filtering, this consensus aggregates historical predictions according to structural consistency and noise-level compatibility. To counteract the over-smoothing tendency of temporal consensus, we introduce a deviation guidance mechanism to adaptively preserve residual details. Empirically, DiFA yields significant improvements on CIFAR-10 and ImageNet across the evaluated metrics, including FID, IS, and FD-DINOv2, demonstrating that aligning inference with the forward statistical structure substantially improves generative fidelity.