GaussianSelector:3D高斯潑濺中結合圖最佳化的輕量級人類引導物件選擇
GaussianSelector: Lightweight Human-Guided Object Selection in 3D Gaussian Splatting with Graph Optimization
August 2, 2026
作者: Baihan Yang, Tiexin Li, Yuheng Liu, Xin Lin, Xinke Li, Xiaohui Xie, Truong Nguyen
cs.AI
摘要
從重建場景中選取完整的3D物體,且僅需最少的使用者操作,對於實際場景編輯與具身互動至關重要。現有的基於3DGS方法,要不是重新訓練高斯表示以嵌入逐物體標籤,就是建構密集的多視圖SAM觀測,兩者皆需要大量計算與密集的視角覆蓋,而這在實務上很少能取得。我們提出GaussianSelector,這是一個免訓練框架,可從稀疏視圖與稀疏塗鴉引導中進行互動式3D物體選取。此方法直接操作於原生高斯基元,將密集高斯點粗化為幾何一致的超點,並利用外觀與空間線索構建連續性加權圖。我們透過可見性感知的透射率覆蓋,將稀疏的使用者塗鴉映射至3D空間,並將選取問題求解為全域圖割能量最小化,藉此將稀疏證據傳播至完整的3D物體。此設計自然支援多輪精化,使用者可從額外視圖迭代修正選取結果,以逐步提升成果。實驗顯示,GaussianSelector在選取品質上可與最先進的多視圖SAM方法媲美,同時需要顯著較少的互動視圖與大幅更低的計算開銷。這些特性使其非常適合在真實世界部署場景中,用於人在迴路的3D場景編輯與3D資產提取。
English
Selecting a complete 3D object from a reconstructed scene with minimal user effort is essential for practical scene editing and embodied interaction. Existing 3DGS-based methods either retrain the Gaussian representation to embed per-object labels, or build dense multi-view SAM observations, both requiring heavy computation and dense viewpoint coverage that is rarely available in practice. We present GaussianSelector, a training-free framework for interactive 3D object selection from sparse views and sparse scribble guidance. Operating directly on native Gaussian primitives, we coarsen dense Gaussians into geometrically coherent superpoints and construct a continuity-weighted graph using appearance and spatial cues. Sparse user scribbles are lifted into 3D via visibility-aware transmittance coverage, and selection is solved as a global graph-cut energy minimization that propagates sparse evidence to a complete 3D object. This design naturally supports multi-round refinement, where users iteratively correct the selection from additional viewpoints to progressively improve the result. Experiments demonstrate that GaussianSelector achieves competitive selection quality against state-of-the-art multi-view SAM-based methods, while requiring significantly fewer interaction views and substantially lower computational overhead. These properties make it well suited for human-in-the-loop 3D scene editing and 3D asset extraction in real-world deployment scenarios.