ChatPaper.aiChatPaper

REBASE:參考-背景子空間消除技術應用於無需訓練的上下文內分割

REBASE: Reference-Background Subspace Elimination for Training-Free In-Context Segmentation

July 10, 2026
作者: Mantha Sai Gopal, Jaison Saji Chacko, Harsh Nandwana, Sandesh Hegde, Debarshi Banerjee, Uma Mahesh
cs.AI

摘要

免訓練上下文分割允許在推理時透過單張標註參考圖像引入新的物體類別,從而消除類增量學習所需的重新訓練與記憶體開銷。當前方法透過結合用於語義對應的視覺基礎模型與如 SAM 這類可提示分割網路來實現此目標。然而,其效能根本受限於跨圖像相似度圖的品質;參考圖像與查詢圖像之間共享的上下文背景會系統性地提升非目標區域的相似度,從而降低提示定位的準確性。我們提出 REBASE,一個明確抑制這些虛假上下文對應的免訓練框架。我們的方法從參考圖像中識別出低秩背景特徵子空間,並以封閉形式將參考與查詢特徵投影至其正交補餘上,從而得到更乾淨的語義匹配。接著,我們使用相似度加權的最遠點採樣生成正點提示,並搭配改良的密集相似度先驗。在無需任何訓練或參數更新的情況下,我們的方法在 PACO-Part、FSS-1000 以及跨領域數據集(如 ISIC2018)上,於免訓練方法中達到最新最佳表現,證明明確的背景子空間移除是一種高效的一次性定位原則。
English
Training-free in-context segmentation enables new object categories to be introduced at inference time from a single annotated reference image, eliminating the retraining and memory overhead of class-incremental learning. Recent approaches achieve this by combining vision foundation models for semantic correspondence with promptable segmentation networks like SAM. However, their performance is fundamentally limited by the quality of the cross-image similarity map; shared contextual backgrounds between the reference and query systematically elevate similarity in non-target regions, degrading prompt localization. We present REBASE, a training-free framework that explicitly suppresses these spurious contextual correspondences. Our method identifies the low-rank background feature subspace from the reference image and project the reference and query features onto its orthogonal complement in closed form, yielding cleaner semantic matching. We then generate positive point prompts using similarity-weighted farthest-point sampling, paired with a refined dense similarity prior. Without any training or parameter updates, our approach establishes a new state of the art among training-free methods on PACO-Part, FSS-1000, and cross-domain datasets such as ISIC2018, demonstrating that explicit background subspace removal is a highly effective principle for one-shot localization.