ZooClaw-FashionSigLIP2: Gedistilleerde fijnafstemming voor robuuste mode retrieval
ZooClaw-FashionSigLIP2: Distilled Fine-tuning for Robust Fashion Retrieval
June 26, 2026
Auteurs: Siqiao Xue, Chunxue Xu
cs.AI
Samenvatting
Het aanpassen van een fundamentele visie-taalencoder aan een gespecialiseerde zoektaak creëert een fundamentele afweging: winst op de doeldistributie gaat ten koste van de brede generalisatie van het fundamentmodel, en mode-retrieval is een stringent voorbeeld van dit probleem. We presenteren ZooClaw-FashionSigLIP2, een op mode gespecialiseerd SigLIP2-basismodel dat deze afweging oplost met een eenvoudig recept — volledige fijnafstelling met kennisdistillatie op samengestelde domeinspecifieke gegevens, gevolgd door wiseft~wortsman2022wiseft gewichtsinterpolatie met het basismodel — en presteert beter dan LoRA, grotere backbones (tot 1B parameters) en externe trainingsgegevens. Onder eerlijke evaluatie presteert ZooClaw-FashionSigLIP2 beter dan alle baselines op elke benchmark in onze suite. Daarnaast brengen we ZooClaw-Fashion uit, een nieuwe hoogwaardige mode-retrieval benchmark, en een systematische kwaliteitsanalyse van veelgebruikte benchmarks die structurele vertekeningen in hun openbare grondwaarheid blootlegt en mitigeert. We open-sourcen de modelgewichten en alle evaluatieartefacten om toekomstig onderzoek te ondersteunen.
English
Adapting a foundation vision-language encoder to a specialized retrieval task creates a fundamental tradeoff: gains on the target distribution come at the cost of the foundation model's broad generalization, and fashion retrieval is a stringent instance of this problem. We present ZooClaw-FashionSigLIP2, a fashion-specialized SigLIP2-base model that resolves this tradeoff with a simple recipe -- full fine-tuning with knowledge distillation on curated in-domain data, followed by \wiseft~wortsman2022wiseft weight interpolation with the base model -- and outperforms LoRA, larger backbones (up to 1B parameters), and external training data. Under fair evaluation, ZooClaw-FashionSigLIP2 outperforms all baselines on every benchmark in our suite. In addition, we release ZooClaw-Fashion, a new high-quality fashion retrieval benchmark, and a systematic quality analysis of widely-used benchmarks that exposes and mitigates structural biases in their public ground truth. We open-source the model weights and all evaluation artifacts to facilitate future research.