ChatPaper.aiChatPaper

BioVITA: Biologische Dataset, Model en Benchmark voor Visueel-Textueel-Akoestische Afstemming

BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment

March 25, 2026
Auteurs: Risa Shinoda, Kaede Shiohara, Nakamasa Inoue, Kuniaki Saito, Hiroaki Santo, Fumio Okura
cs.AI

Samenvatting

Het begrijpen van diersoorten op basis van multimodale data vormt een opkomende uitdaging op het snijvlak van computervisie en ecologie. Hoewel recente biologische modellen, zoals BioCLIP, een sterke afstemming hebben aangetoond tussen afbeeldingen en tekstuele taxonomische informatie voor soortidentificatie, blijft de integratie van de audiomodule een onopgelost probleem. Wij stellen BioVITA voor, een nieuw visueel-textueel-akoestisch afstemmingsraamwerk voor biologische toepassingen. BioVITA omvat (i) een traindataset, (ii) een representatiemodel en (iii) een retrievalbenchmark. Ten eerste construeren we een grootschalige traindataset bestaande uit 1,3 miljoen audioclips en 2,3 miljoen afbeeldingen, die 14.133 soorten bestrijkt, geannoteerd met 34 ecologische kenmerklabels. Ten tweede introduceren we, voortbouwend op BioCLIP2, een tweefasen-trainingsraamwerk om audiorepresentaties effectief af te stemmen op visuele en tekstuele representaties. Ten derde ontwikkelen we een cross-modale retrievalbenchmark die alle mogelijke directionele retrieval tussen de drie modaliteiten dekt (d.w.z. beeld-naar-audio, audio-naar-tekst, tekst-naar-beeld en hun omgekeerde richtingen), met drie taxonomische niveaus: Familie, Geslacht en Soort. Uitgebreide experimenten tonen aan dat ons model een verenigde representatieruimte leert die semantiek op soortniveau vastlegt die verder gaat dan taxonomie, waardoor het multimodale begrip van biodiversiteit wordt bevorderd. De projectpagina is beschikbaar op: https://dahlian00.github.io/BioVITA_Page/
English
Understanding animal species from multimodal data poses an emerging challenge at the intersection of computer vision and ecology. While recent biological models, such as BioCLIP, have demonstrated strong alignment between images and textual taxonomic information for species identification, the integration of the audio modality remains an open problem. We propose BioVITA, a novel visual-textual-acoustic alignment framework for biological applications. BioVITA involves (i) a training dataset, (ii) a representation model, and (iii) a retrieval benchmark. First, we construct a large-scale training dataset comprising 1.3 million audio clips and 2.3 million images, covering 14,133 species annotated with 34 ecological trait labels. Second, building upon BioCLIP2, we introduce a two-stage training framework to effectively align audio representations with visual and textual representations. Third, we develop a cross-modal retrieval benchmark that covers all possible directional retrieval across the three modalities (i.e., image-to-audio, audio-to-text, text-to-image, and their reverse directions), with three taxonomic levels: Family, Genus, and Species. Extensive experiments demonstrate that our model learns a unified representation space that captures species-level semantics beyond taxonomy, advancing multimodal biodiversity understanding. The project page is available at: https://dahlian00.github.io/BioVITA_Page/