시계열 파운데이션 모델에서의 예측 붕괴
Forecast Collapse in Time-Series Foundation Models
August 14, 2026
저자: Shu Wan, Miles Ma, Hank Zhu, Guangqi Liu, Stephen Wang, Qingsong Wen, Huan Liu
cs.AI
초록
1,000개 미국 주식의 시간별 수익률을 예측할 때, 우리는 예상치 못한 현상을 관찰한다: 예측값이 거의 평평해지고, 횡단면 상관관계로 측정한 주식 순위 성능이 저조해진다. 우리는 이를 예측 붕괴(forecast collapse)라고 명명한다. 놀랍게도, 동일한 설정에서 거래량을 예측할 때는 이러한 현상이 대부분 사라진다. 우리는 시계열 기반 모델(TSFM), 12개의 딥러닝 예측 모델, 97개의 공개 벤치마크 구성에 걸쳐 예측 붕괴를 조사했으며, 이것이 목표 변수의 예측 가능성과 밀접하게 연관되어 있음을 발견했다. 우리는 그 이면에 두 가지 서로 다른 원인이 있음을 확인했다: 낮은 예측 가능성은 보정된 점 예측의 진폭을 제한하고, 계열별 목적 함수는 계열 간 구조를 식별하지 못하게 한다. 이러한 발견은 보정-순위 트레이드오프를 드러낸다: 제곱 오차를 최적화하면 평평한 예측이 생성되는 반면, 횡단면 상관관계를 직접 최적화하면 순위는 개선되지만 예측 진폭이 한 자릿수 이상 부풀려질 수 있다. 이러한 트레이드오프를 해결하기 위해 우리는 보정과 순위의 균형을 맞추는 간단한 목적 함수인 CalibRank를 도입한다. Finance1K에서 CalibRank는 진폭을 목표값에 가깝게 유지하면서 횡단면 상관관계를 거의 3배로 높이며, 테스트한 모든 모델에서 상관관계를 개선한다. 우리의 결과는 기존 시계열 평가의 사각지대를 드러낸다: 계열별 지표는 후속 의사결정에 필요한 계열 간 구조의 실패를 숨길 수 있다.
English
When forecasting hourly returns for 1,000 US equities, we observe an unexpected phenomenon: predictions become nearly flat and show poor stock ranking, as measured by cross-sectional correlation. We call this forecast collapse. Surprisingly, the phenomenon largely disappears when forecasting trading volume under the same setting. We investigate forecast collapse across time-series foundation models (TSFMs), twelve deep-learning forecasting models, and 97 public benchmark configurations, and find that it is closely tied to target predictability. We identify two distinct reasons behind it: low predictability limits the amplitude of calibrated point forecasts, while per-series objectives leave cross-series structure unidentified. These findings reveal a calibration-ranking tradeoff: optimizing squared error leads to flat predictions, whereas directly optimizing cross-sectional correlation improves ranking but can inflate forecast amplitude by more than an order of magnitude. To address this tradeoff, we introduce CalibRank, a simple objective that balances calibration and ranking. On Finance1K, CalibRank nearly triples cross-sectional correlation while keeping amplitude close to the target, and improves correlation on all tested models. Our results reveal a blind spot in conventional time-series evaluation: per-series metrics can hide failures in cross-series structure needed by downstream decisions.