ChatPaper.aiChatPaper

자연 이미지를 이용한 대조 학습 이론

A Theory of Contrastive Learning with Natural Images

July 8, 2026
저자: Antonio Torralba, Yair Weiss
cs.AI

초록

단순 이미지와 증강을 사용한 대조 학습이 왜 하위 작업에 유용한 표현을 생성하는가? 우리는 이 질문에 답하기 위해 정상 통계를 따르는 모든 이미지 데이터셋과 다양한 기본 증강에 대해 대조 손실 측면에서 최적 표현을 해석적으로 계산한다. 특정 증강의 경우, 최적해는 첫 번째 계층의 필터가 정현파이고, 이후 점별 비선형성, 전역 평균 풀링, 그리고 부분 백색화를 수행하는 최종 선형 계층이 뒤따르는 CNN에 의해 달성될 수 있음을 보인다. 또한 더 복잡한 증강에 대한 이러한 CNN의 최적 가중치 역시 정현파임을 보인다. 정현파의 주파수와 그 가중치는 데이터셋의 기대 전력 스펙트럼이 주어졌을 때 간단한 워터필링 알고리즘을 사용하여 계산할 수 있다. 다양한 이미지 데이터셋과 증강을 사용한 실험은 이러한 CNN이 SGD로 훈련될 때 경험적으로 첫 번째 계층에서 정현파를 학습하고 부분 백색화를 수행함을 보여준다.
English
Why does contrastive learning with simple images and augmentations yield useful representations for downstream tasks? We address this question by analytically computing the optimal representation in terms of a contrastive loss for a range of basic augmentations and any image dataset with stationary statistics. We show that for certain augmentations the optimum can be attained by a CNN whose first layer filters are sinusoids, followed by a pointwise nonlinearity, global average pooling, and a final linear layer that performs partial whitening. We also show that the optimal weights in such CNNs for more complicated augmentations are still sinusoids. The frequencies of the sinusoids and their weights can be computed using a simple waterfilling algorithm given the dataset's expected power spectrum. Experiments with different image datasets and augmentations show that such CNNs trained with SGD empirically learn sinusoids in their first layer and to perform partial whitening