ChatPaper.aiChatPaper

Een theorie over contrastief leren met natuurlijke beelden

A Theory of Contrastive Learning with Natural Images

July 8, 2026
Auteurs: Antonio Torralba, Yair Weiss
cs.AI

Samenvatting

Waarom levert contrastief leren met eenvoudige afbeeldingen en augmentaties bruikbare representaties op voor downstream taken? We behandelen deze vraag door de optimale representatie analytisch te berekenen in termen van een contrastief verlies voor een reeks basisaugmentaties en elke beelddataset met stationaire statistiek. We tonen aan dat voor bepaalde augmentaties het optimum kan worden bereikt door een CNN waarvan de eerste laag filters sinusoïden zijn, gevolgd door een puntgewijze niet-lineariteit, globale gemiddelde pooling en een laatste lineaire laag die gedeeltelijke witmaking uitvoert. We laten ook zien dat de optimale gewichten in dergelijke CNN's voor complexere augmentaties nog steeds sinusoïden zijn. De frequenties van de sinusoïden en hun gewichten kunnen worden berekend met behulp van een eenvoudig wateropvulalgoritme, gegeven het verwachte vermogensspectrum van de dataset. Experimenten met verschillende beelddatasets en augmentaties tonen aan dat dergelijke CNN's, getraind met SGD, empirisch sinusoïden leren in hun eerste laag en gedeeltelijke witmaking uitvoeren.
English
Why does contrastive learning with simple images and augmentations yield useful representations for downstream tasks? We address this question by analytically computing the optimal representation in terms of a contrastive loss for a range of basic augmentations and any image dataset with stationary statistics. We show that for certain augmentations the optimum can be attained by a CNN whose first layer filters are sinusoids, followed by a pointwise nonlinearity, global average pooling, and a final linear layer that performs partial whitening. We also show that the optimal weights in such CNNs for more complicated augmentations are still sinusoids. The frequencies of the sinusoids and their weights can be computed using a simple waterfilling algorithm given the dataset's expected power spectrum. Experiments with different image datasets and augmentations show that such CNNs trained with SGD empirically learn sinusoids in their first layer and to perform partial whitening