ChatPaper.aiChatPaper

データスケーリングによる高解像度気象予測の限界への挑戦

Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling

July 31, 2026
著者: Yang Zhao, Peisong Niu, Tian Zhou, Ziqing Ma, Guanlong Ma, Rong Jin, Huiling Yuan, Liang Sun
cs.AI

要旨

機械学習(ML)に基づく0.1度全球天気予報モデルの開発は、高解像度データの利用可能性が限られているために制約されている。というのも、数十年分の再解析データは0.25度解像度でしか入手できないからである。既存の手法は限られた0.1度サンプルを用いて0.25度予報モデルをファインチューニングするが、本稿では、この転移が粗解像度予報に内在する不可逆的な情報損失によって妨げられることを示す。そこで我々は、モデルの転移からデータの転移へと焦点を移すフレームワークであるBaguanHRを提案する。まず、超解像(SR)が予報よりも低い条件付きエントロピーと入力の増幅を有し、解像度転移のためのよりロバストな手段であることを示す。この利点を変数別のSRを通じて活用することにより、ERA5から大量の0.1度データを合成する。合成データと実データを組み合わせたデータセットにおけるBaguanHRの性能は、MLベースの手法とIFS-HRESの両方を上回り、72時間以内のリードタイムの85%以上で優れた性能を達成する。さらに、我々の発見は冪乗則に従うスケーリング効果を浮き彫りにする。データを2倍に増やすと、72時間予報ではRMSEが4.6%、120時間予報では4.9%減少する。これらの結果は、高解像度MLベース予報のスケーリングが主にデータのボトルネックに制約されること、そして変数別の超解像が、高解像度トレーニングのために長期間の粗解像度再解析データを活用可能にする簡潔かつ一般的な解決策を提供することを示している。
English
The development of 0.1^{circ} global weather forecasting models based on machine learning (ML) is constrained by the limited availability of high-resolution data, as decades of reanalysis are only available at 0.25^{circ} resolution. While existing approaches fine-tune 0.25^{circ} forecast models on limited 0.1^{circ} samples, we show that this transfer is hindered by the irreversible information loss inherent in coarse-resolution forecasting. Therefore, we propose BaguanHR, a framework that shifts the focus from transferring models to transferring data. We first show that super-resolution (SR) has lower conditional entropy and input amplification than forecasting, making it a more robust vehicle for resolution transfer. By leveraging this advantage through variable-wise SR, we synthesize extensive 0.1^{circ} data from ERA5. BaguanHR's performance on the synthetic-plus-real dataset exceeds both ML-based methods and IFS-HRES, achieving superior performance across over 85% of the lead times within 72 hours. Furthermore, our findings highlight a power-law scaling effect, as a twofold increase in data reduces RMSE by 4.6% for 72-hour forecasting and 4.9% for 120-hour forecasting. Our results demonstrate that scaling high resolution ML-based forecasting is primarily a data bottleneck, and that variable-wise super-resolution provides a simple yet general solution to unlock long coarse-resolution reanalyses for high-resolution training.