通过数据扩展突破高分辨率天气预报的极限
Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling
July 31, 2026
作者: Yang Zhao, Peisong Niu, Tian Zhou, Ziqing Ma, Guanlong Ma, Rong Jin, Huiling Yuan, Liang Sun
cs.AI
摘要
基于机器学习(ML)的0.1°全球天气预报模型的发展受到高分辨率数据可用性有限的制约,因为数十年的再分析数据仅能以0.25°分辨率获取。虽然现有方法在有限的0.1°样本上对0.25°预报模型进行微调,但我们表明,这种迁移受到粗分辨率预报中固有的不可逆信息损失的阻碍。为此,我们提出BaguanHR框架,将关注点从模型迁移转向数据迁移。我们首先证明,超分辨率(SR)比预报具有更低的条件熵和输入放大效应,因此是分辨率迁移的更稳健载体。通过利用这一优势,我们采用分变量超分辨率从ERA5合成了大量0.1°数据。BaguanHR在合成与真实数据集上的性能超越了基于ML的方法和IFS-HRES,在72小时内超过85%的预报时效上实现了更优表现。此外,我们的研究结果凸显了一种幂律尺度效应:数据量每增加一倍,72小时预报的RMSE降低4.6%,120小时预报降低4.9%。我们的结果表明,高分辨率基于ML预报的规模化扩展主要受制于数据瓶颈,而分变量超分辨率提供了一种简单且通用的解决方案,可释放长时程粗分辨率再分析数据在高分辨率训练中的潜力。
English
The development of 0.1^{circ} global weather forecasting models based on machine learning (ML) is constrained by the limited availability of high-resolution data, as decades of reanalysis are only available at 0.25^{circ} resolution. While existing approaches fine-tune 0.25^{circ} forecast models on limited 0.1^{circ} samples, we show that this transfer is hindered by the irreversible information loss inherent in coarse-resolution forecasting. Therefore, we propose BaguanHR, a framework that shifts the focus from transferring models to transferring data. We first show that super-resolution (SR) has lower conditional entropy and input amplification than forecasting, making it a more robust vehicle for resolution transfer. By leveraging this advantage through variable-wise SR, we synthesize extensive 0.1^{circ} data from ERA5. BaguanHR's performance on the synthetic-plus-real dataset exceeds both ML-based methods and IFS-HRES, achieving superior performance across over 85% of the lead times within 72 hours. Furthermore, our findings highlight a power-law scaling effect, as a twofold increase in data reduces RMSE by 4.6% for 72-hour forecasting and 4.9% for 120-hour forecasting. Our results demonstrate that scaling high resolution ML-based forecasting is primarily a data bottleneck, and that variable-wise super-resolution provides a simple yet general solution to unlock long coarse-resolution reanalyses for high-resolution training.