CoVeR: VLM에서의 다시점 3D 추론을 위한 커버리지 기반 토큰 프루닝
CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs
September 8, 2026
저자: Nhat-Tan Bui, Varshini Elangovan, Arun Reddy Anugu, Sreyas Mohan, Wei Ye, Dilin Wang, JQ Huang, Rakesh Ranjan, Aviral Chharia, Fernando De la Torre
cs.AI
초록
3D 장면을 다시점 이미지로 표현하면 2D VLM들이 사전 학습에서 얻은 사전 정보를 재사용하여 3D에서 추론할 수 있게 하므로, 주석이 달린 3D 데이터의 부족 문제를 우회할 수 있다. 그러나 이는 수천 개의 중복 시각 토큰을 생성하며, 그 비용은 뷰가 추가될 때마다 증가한다. 기존 시각 토큰 프루너는 두 계열로 나뉘지만, 각각은 3D 다시점 설정에서 한계를 지닌다. 학습 기반 중요도 방법은 어텐션 또는 인코더 특징을 기준으로 토큰의 순위를 매기는데, 여기서 중복성은 근본적으로 공간적이기 때문에 몇몇 두드러진 영역에서 거의 중복된 토큰들을 유지하고 장면 대부분을 표현하지 못한 채 남긴다. 복셀화 방법은 공간적 커버리지를 향상시키지만 정확한 토큰 예산을 강제할 수 없고, 3D에서 다시점 관측이 중첩되면서 포화되어 유지율을 목표치보다 훨씬 낮게 제한한다. 우리는 공간적 커버리지가 3D 추론 성능과 관련이 있음을 보이고, 학습된 신호 없이 토큰 좌표만을 사용하는 결정론적이며 학습이 필요 없는 선택기인 CoVeR를 소개한다. CoVeR는 장면의 모든 영역을 집합적으로 커버하는 토큰을 선택하며, 두 계열의 한계를 모두 해결한다. 즉, 장면별 정확한 예산을 강제하고, 복셀화의 포화 정체 구간을 깨뜨리며, 학습 기반 중요도 방법의 거의 중복된 선택을 피한다. 광범위한 실험은 CoVeR가 세 가지 3D 추론 벤치마크 모두에서 기존 SOTA들을 능가하고 네 가지 VLM 전반에서 테스트된 플러그 앤 플레이 모듈로서 일반화됨을 보여준다. 주목할 점은, 시각 토큰의 약 8%만으로 전체 토큰 성능의 93.5%를 유지하여, 벤치마크 전반에서 평균적으로 SOTA를 3.9%p 능가한다는 것이다.
English
Representing a 3D scene as multi-view images allows 2D VLMs to reason in 3D by reusing priors from pre-training, sidestepping the scarcity of annotated 3D data. However, it produces thousands of redundant visual tokens whose cost grows with every view. Existing visual token pruners fall into two families, each limited in the 3D multi-view setting. Learned importance methods rank tokens by attention or encoder features; because redundancy here is fundamentally spatial, they keep near-duplicate tokens from a few prominent regions and leave most of the scene unrepresented. Voxelization methods improve spatial coverage but cannot enforce an exact token budget and saturate as multi-view observations overlap in 3D, capping retention well below the target. We show that spatial coverage is associated with 3D reasoning performance and introduce CoVeR, a deterministic, training-free selector that uses only token coordinates, with no learned signals. CoVeR selects tokens that collectively cover every region of the scene, and solves the limitations of both families: it enforces an exact per-scene budget, breaks the voxelization saturation plateau, and avoids the near-duplicate selections of learned importance. Extensive experiments show CoVeR outperforms prior SOTAs on all three 3D reasoning benchmarks and generalizes as a plug-and-play module tested across four VLMs. Notably, with only approx8% of visual tokens, it preserves 93.5% of full-token performance, surpassing SOTA by 3.9 percentage points on average across benchmarks.