透過資料擴展突破高解析度天氣預報的極限
Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling
July 31, 2026
作者: Yang Zhao, Peisong Niu, Tian Zhou, Ziqing Ma, Guanlong Ma, Rong Jin, Huiling Yuan, Liang Sun
cs.AI
摘要
基於機器學習(ML)的0.1°全球天氣預報模型的發展受到高解析度資料可用性有限的制約,因為數十年的再分析資料僅提供0.25°解析度。現有方法在有限的0.1°樣本上微調0.25°預報模型,但我們表明,此種遷移受到粗解析度預報固有的不可逆資訊損失所阻礙。為此,我們提出BaguanHR,一個將重點從遷移模型轉向遷移資料的框架。我們首先證明,超解析度(SR)的條件熵和輸入放大均低於預報,因此是更穩健的解析度遷移工具。藉由逐變量SR利用此優勢,我們從ERA5合成大量0.1°資料。BaguanHR在合成加真實資料集上的表現優於基於ML的方法及IFS-HRES,在72小時內的預報時效中,超過85%的時效實現優越表現。此外,我們的結果凸顯了冪律縮放效應:資料量增加一倍,72小時預報的RMSE降低4.6%,120小時預報的RMSE降低4.9%。我們的結果表明,擴展高解析度ML預報的主要瓶頸在於資料,而逐變量超解析度提供了一個簡單且通用的解決方案,可釋放長期粗解析度再分析資料以供高解析度訓練使用。
English
The development of 0.1^{circ} global weather forecasting models based on machine learning (ML) is constrained by the limited availability of high-resolution data, as decades of reanalysis are only available at 0.25^{circ} resolution. While existing approaches fine-tune 0.25^{circ} forecast models on limited 0.1^{circ} samples, we show that this transfer is hindered by the irreversible information loss inherent in coarse-resolution forecasting. Therefore, we propose BaguanHR, a framework that shifts the focus from transferring models to transferring data. We first show that super-resolution (SR) has lower conditional entropy and input amplification than forecasting, making it a more robust vehicle for resolution transfer. By leveraging this advantage through variable-wise SR, we synthesize extensive 0.1^{circ} data from ERA5. BaguanHR's performance on the synthetic-plus-real dataset exceeds both ML-based methods and IFS-HRES, achieving superior performance across over 85% of the lead times within 72 hours. Furthermore, our findings highlight a power-law scaling effect, as a twofold increase in data reduces RMSE by 4.6% for 72-hour forecasting and 4.9% for 120-hour forecasting. Our results demonstrate that scaling high resolution ML-based forecasting is primarily a data bottleneck, and that variable-wise super-resolution provides a simple yet general solution to unlock long coarse-resolution reanalyses for high-resolution training.