ChatPaper.aiChatPaper

TESSERA v2: 픽셀 단위 지구 기초 모델의 확장

TESSERA v2: Scaling Pixel-wise Earth Foundation Models

July 4, 2026
저자: Zhengpeng Feng, Sadiq Jaffer, Ira Shokar, Jovana Knezevic, Mark Elvers, Clement Atzberger, Robin Young, Aneesh Naik, Niall Robinson, Andrew Blake, David Coomes, Anil Madhavapeddy, Srinivasan Keshav
cs.AI

초록

픽셀 단위 지구 관측(EO) 기반 모델은 현재 생성된 공간 임베딩을 통해 최첨단 성능을 달성하고 있다. 그러나 이러한 모델이 어떻게 확장되며 사전 학습 예산을 가장 효율적으로 사용하는 방법에 대한 이해는 여전히 부족하다. 본 연구는 지금까지 EO 분야에서 수행된 가장 큰 규모의 통제된 확장 연구를 제시한다: 고정된 픽셀 단위 Barlow Twins 계열 내에서 1,024개의 GH200 슈퍼칩에서 395회의 훈련 실행을 수행하고, 각각을 15개의 하위 작업에서 평가했다. 그 결과, 사전 학습 손실은 하위 작업 성능을 거의 예측하지 못하며(|피어슨 r| < 0.2), 손실 기준으로 모델을 선택하는 것은 상당한 계산 자원을 낭비함을 발견했다. 또한, 훈련 예산이 증가함에 따라 인코더와 데이터는 함께 확장되어야 하지만 투영기는 고정된 상태로 유지되어야 하며, 이는 계산 자원 할당을 위한 간단한 규칙을 제공한다. 이 규칙을 사용하여 픽셀 단위 모델군(0.5B, 1B, 훈련 중인 2B 모델)을 훈련하고 이를 소형 학생 모델로 증류하여 임베딩-데이터-배포 방식에 활용한다. 2,100만 개의 파라미터를 가진 증류된 TESSERA v2-1B-M은 테스트된 모든 공개 및 독점 모델(일부는 수십 배 더 큼)을 종합적으로 능가한다. 이 학생 모델들은 서비스 비용이 저렴한 마트료시카 표현을 생성하며, 16차원 접두사가 전체 128차원 성능의 92%를 유지하면서 저장 공간은 1/8만 사용한다. 훈련 완료 후 2017-2025년을 포괄하는 v2 글로벌 임베딩을 공개할 예정이다. 이러한 결과는 픽셀 단위 EO 기반 모델 확장을 위한 구체적이고 경험적으로 근거한 방법론을 제시한다: 대형 인코더를 훈련하고, 하위 작업 성능으로 선택하며, 유연한 학생 모델로 증류하는 것이다. 전체 코드는 https://github.com/ucam-eo/tessera에서 공개될 예정이다.
English
Pixel-wise Earth-observation (EO) foundation models are now achieving state-of-the-art performance via generated spatial embeddings. However, how these models scale and how best to spend a pretraining budget remain poorly understood. We present the largest controlled scaling study for EO to date: 395 training runs on 1,024 GH200 superchips within a fixed pixel-wise Barlow Twins family, each evaluated on 15 downstream tasks. We find that pretraining loss barely predicts downstream performance (|Pearson r| < 0.2), so selecting models by loss wastes a large share of the compute. We also find that, as the training budget grows, the encoder and the data should grow together while the projector stays fixed, which gives a simple rule for allocating compute. Using this rule, we train a family of pixel-wise models (0.5B and 1B, with a 2B model in training) and distill them into compact students for embeddings-as-data deployment. The 21-million-parameter distilled TESSERA v2-1B-M in aggregate outperforms all open and proprietary models tested, some of which are orders of magnitude larger. These students produce Matryoshka representations that are inexpensive to serve: a 16-dimensional prefix keeps 92% of the full 128-dimensional performance at 1/8 of the storage. Upon completion of training we plan to release v2 global embeddings covering 2017-2025. Together, these results give a concrete, empirically grounded recipe for scaling pixel-wise EO foundation models: train large encoders, select by downstream performance, and distil into flexible student models. All code will be released at https://github.com/ucam-eo/tessera.