Scènes als objecten, niet als primitieven: instantie-gestructureerde 3D-tokenisatie vanuit niet-geposeerde aanzichten
Scenes as Objects, Not Primitives: Instance-Structured 3D Tokenization from Unposed Views
June 28, 2026
Auteurs: Mijin Yoo, In Cho, Subin Jeon, Jiwoo Lee, Eunbyung Park, Seon Joo Kim
cs.AI
Samenvatting
Een 3D-scène wordt begrepen via zijn objecten, niet via de primitieven waaruit ze zijn opgebouwd. Toch genereren feed-forward-reconstructiemethoden dichte, ongestructureerde verzamelingen punten of Gaussianen, waardoor de objectniveaustructuur achteraf moet worden hersteld. Wij stellen een feed-forward-raamwerk voor dat een scène rechtstreeks vanuit ongeposeerde multi-view-beelden ontleedt in instantie-gestructureerde 3D-tokengroepen — compacte objectgecentreerde eenheden waaruit reconstructie, segmentatie en manipulatie allemaal volgen. Elke tokengroep koppelt een instantie-token dat de identiteit op entiteitsniveau vastlegt aan ankertokens die de lokale geometrie en het uiterlijk coderen, en die worden gedecodeerd tot een reeks 3D-Gaussianen. Deze tweeledige factorisatie ontkoppelt objectidentiteit van lokaal uiterlijk, waardoor objectinstanties een native interface van de representatie worden in plaats van een afgeleid product. De tokengroepen worden geleerd via differentieerbare rendering met gezamenlijke reconstructie- en segmentatiesupervisie, zonder dat er 3D-annotaties nodig zijn. Ons feed-forward-model overtreft per-scène-optimalisatiebaselines in klasse-agnostische instantiesegmentatie, terwijl het concurrerend blijft in nieuw-gezichtssynthese. Naast deze metrieken maken dezelfde tokengroepen direct instantieniveau-scènebewerking mogelijk — verwijderen, verplaatsen of invoegen van objecten door op hun groepen te manipuleren — evenals efficiënt open-vocabulary 3D-instantie-opvragen, waarbij de opvraagcomplexiteit schaalt met het aantal instanties in plaats van primitieven.
English
A 3D scene is understood through its objects, not the primitives that compose them. Yet feed-forward reconstruction methods output dense, unstructured sets of points or Gaussians, leaving object-level structure to be recovered after the fact. We propose a feed-forward framework that decomposes a scene into instance-structured 3D token groups directly from unposed multi-view images -- compact object-centric units from which reconstruction, segmentation, and manipulation all follow. Each token group pairs an instance token capturing entity-level identity with anchor tokens that encode local geometry and appearance, which are decoded into a set of 3D Gaussians. This two-level factorization decouples object identity from local appearance, making object instances a native interface of the representation rather than a derived product. The token groups are learned through differentiable rendering with joint reconstruction and segmentation supervision, requiring no 3D annotations. Our feed-forward model surpasses per-scene optimization baselines in class-agnostic instance segmentation while remaining competitive in novel view synthesis. Beyond these metrics, the same token groups directly unlock instance-level scene editing -- removing, translating, or inserting objects by operating on their groups -- as well as efficient open-vocabulary 3D instance retrieval, where retrieval complexity scales with the number of instances rather than primitives.