데이터 스케일링을 통한 고해상도 기상 예보의 한계 극복
Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling
July 31, 2026
저자: Yang Zhao, Peisong Niu, Tian Zhou, Ziqing Ma, Guanlong Ma, Rong Jin, Huiling Yuan, Liang Sun
cs.AI
초록
기계학습(ML) 기반 0.1° 전지구 기상 예보 모델의 개발은 고해상도 데이터의 제한된 가용성으로 인해 제약을 받는다. 수십 년간의 재분석 자료가 0.25° 해상도로만 제공되기 때문이다. 기존 접근법은 제한된 0.1° 샘플에 대해 0.25° 예보 모델을 미세 조정하지만, 우리는 이러한 전이가 저해상도 예보에 내재된 비가역적 정보 손실로 인해 방해받는다는 것을 보여준다. 따라서 우리는 모델 전이에서 데이터 전이로 초점을 전환하는 프레임워크인 BaguanHR을 제안한다. 우리는 먼저 초해상도(SR)가 예보보다 더 낮은 조건부 엔트로피와 입력 증폭을 가지므로 해상도 전이를 위한 더 견고한 수단임을 보여준다. 우리는 이러한 이점을 변수별 SR을 통해 활용하여 ERA5로부터 방대한 0.1° 데이터를 합성한다. 합성-실측 데이터셋에서 BaguanHR의 성능은 ML 기반 방법과 IFS-HRES를 모두 능가하며, 72시간 이내의 예측 시한 중 85% 이상에서 우수한 성능을 달성한다. 또한, 우리의 연구 결과는 데이터가 두 배로 증가할 때 72시간 예보의 RMSE가 4.6%, 120시간 예보의 RMSE가 4.9% 감소한다는 멱법칙 스케일링 효과를 강조한다. 우리의 결과는 고해상도 ML 기반 예보의 확장이 주로 데이터 병목 현상에 기인함을 보여주며, 변수별 초해상도가 장기간의 저해상도 재분석 자료를 고해상도 훈련에 활용할 수 있게 하는 단순하면서도 일반적인 해법을 제공함을 입증한다.
English
The development of 0.1^{circ} global weather forecasting models based on machine learning (ML) is constrained by the limited availability of high-resolution data, as decades of reanalysis are only available at 0.25^{circ} resolution. While existing approaches fine-tune 0.25^{circ} forecast models on limited 0.1^{circ} samples, we show that this transfer is hindered by the irreversible information loss inherent in coarse-resolution forecasting. Therefore, we propose BaguanHR, a framework that shifts the focus from transferring models to transferring data. We first show that super-resolution (SR) has lower conditional entropy and input amplification than forecasting, making it a more robust vehicle for resolution transfer. By leveraging this advantage through variable-wise SR, we synthesize extensive 0.1^{circ} data from ERA5. BaguanHR's performance on the synthetic-plus-real dataset exceeds both ML-based methods and IFS-HRES, achieving superior performance across over 85% of the lead times within 72 hours. Furthermore, our findings highlight a power-law scaling effect, as a twofold increase in data reduces RMSE by 4.6% for 72-hour forecasting and 4.9% for 120-hour forecasting. Our results demonstrate that scaling high resolution ML-based forecasting is primarily a data bottleneck, and that variable-wise super-resolution provides a simple yet general solution to unlock long coarse-resolution reanalyses for high-resolution training.