ChatPaper.aiChatPaper

Scènebegrip op Pixelniveau in Één Token: Visuele Toestanden Vereisen een Wat-is-Waar-Compositie

Pixel-level Scene Understanding in One Token: Visual States Need What-is-Where Composition

March 14, 2026
Auteurs: Seokmin Lee, Yunghee Lee, Byeonghyun Pak, Byeongju Woo
cs.AI

Samenvatting

Voor robotagenten die opereren in dynamische omgevingen is het leren van visuele toestandsrepresentaties uit streamende videoobservaties essentieel voor sequentieel besluitvorming. Recente zelfgesuperviseerde leermethoden hebben sterke overdraagbaarheid tussen visietaken getoond, maar ze behandelen niet expliciet wat een goede visuele toestand zou moeten coderen. Wij stellen dat effectieve visuele toestanden wat-waar-is moeten vastleggen door zowel de semantische identiteiten van scène-elementen als hun ruimtelijke posities gezamenlijk te coderen, waardoor betrouwbare detectie van subtiele dynamiek tussen observaties mogelijk wordt. Hiertoe stellen we CroBo voor, een raamwerk voor het leren van visuele toestandsrepresentaties gebaseerd op een globaal-naar-lokaal reconstructiedoel. Gegeven een referentieobservatie gecomprimeerd tot een compacte bottleneck-token, leert CroBo zwaar gemaskeerde patches in een lokaal doelgebied te reconstrueren op basis van schaars zichtbare aanwijzingen, waarbij de globale bottleneck-token als context wordt gebruikt. Dit leerdoel moedigt de bottleneck-token aan om een fijnkorrelige representatie te coderen van semantische entiteiten in de gehele scène, inclusief hun identiteiten, ruimtelijke posities en configuraties. Hierdoor onthullen de geleerde visuele toestanden hoe scène-elementen in de tijd bewegen en interacteren, wat sequentiële besluitvorming ondersteunt. We evalueren CroBo op diverse visiegebaseerde benchmarks voor robotbeleidsleren, waar het state-of-the-art prestaties behaalt. Reconstructieanalyses en perceptual straightness-experimenten tonen verder aan dat de geleerde representaties pixel-level scènesamenstelling behouden en coderen wat-waar-naartoe-beweegt tussen observaties. Projectpagina beschikbaar op: https://seokminlee-chris.github.io/CroBo-ProjectPage.
English
For robotic agents operating in dynamic environments, learning visual state representations from streaming video observations is essential for sequential decision making. Recent self-supervised learning methods have shown strong transferability across vision tasks, but they do not explicitly address what a good visual state should encode. We argue that effective visual states must capture what-is-where by jointly encoding the semantic identities of scene elements and their spatial locations, enabling reliable detection of subtle dynamics across observations. To this end, we propose CroBo, a visual state representation learning framework based on a global-to-local reconstruction objective. Given a reference observation compressed into a compact bottleneck token, CroBo learns to reconstruct heavily masked patches in a local target crop from sparse visible cues, using the global bottleneck token as context. This learning objective encourages the bottleneck token to encode a fine-grained representation of scene-wide semantic entities, including their identities, spatial locations, and configurations. As a result, the learned visual states reveal how scene elements move and interact over time, supporting sequential decision making. We evaluate CroBo on diverse vision-based robot policy learning benchmarks, where it achieves state-of-the-art performance. Reconstruction analyses and perceptual straightness experiments further show that the learned representations preserve pixel-level scene composition and encode what-moves-where across observations. Project page available at: https://seokminlee-chris.github.io/CroBo-ProjectPage.