ChatPaper.aiChatPaper

REBASE: 学習不要な文脈内セグメンテーションのための参照背景部分空間除去

REBASE: Reference-Background Subspace Elimination for Training-Free In-Context Segmentation

July 10, 2026
著者: Mantha Sai Gopal, Jaison Saji Chacko, Harsh Nandwana, Sandesh Hegde, Debarshi Banerjee, Uma Mahesh
cs.AI

要旨

訓練不要のコンテキスト内セグメンテーションは、推論時に単一の注釈付き参照画像から新たな物体カテゴリを導入することを可能にし、クラス増分学習における再訓練やメモリオーバーヘッドを排除する。近年のアプローチでは、意味対応のための視覚基盤モデルと、SAMのようなプロンプト可能なセグメンテーションネットワークを組み合わせることでこれを実現している。しかし、その性能は本質的に画像間類似度マップの品質に制約される。参照画像とクエリ画像の間で共有される文脈的背景が、非対象領域の類似度を系統的に高め、プロンプトの位置特定を劣化させる。本稿では、これらの偽の文脈的対応を明示的に抑制する訓練不要フレームワークREBASEを提案する。提案手法は、参照画像から低ランク背景特徴部分空間を特定し、参照特徴とクエリ特徴をその直交補空間へ閉形式で投影することで、よりクリーンな意味対応を実現する。その後、類似度重み付き最遠点サンプリングを用いて正の点プロンプトを生成し、洗練された密な類似度事前分布と組み合わせる。一切の訓練やパラメータ更新を行わずに、本手法はPACO-Part、FSS-1000、およびISIC2018などのクロスドメインデータセットにおいて、訓練不要手法の中で新たな最先端を確立する。これは、明示的な背景部分空間除去がワンショット位置特定に対して極めて効果的な原理であることを示している。
English
Training-free in-context segmentation enables new object categories to be introduced at inference time from a single annotated reference image, eliminating the retraining and memory overhead of class-incremental learning. Recent approaches achieve this by combining vision foundation models for semantic correspondence with promptable segmentation networks like SAM. However, their performance is fundamentally limited by the quality of the cross-image similarity map; shared contextual backgrounds between the reference and query systematically elevate similarity in non-target regions, degrading prompt localization. We present REBASE, a training-free framework that explicitly suppresses these spurious contextual correspondences. Our method identifies the low-rank background feature subspace from the reference image and project the reference and query features onto its orthogonal complement in closed form, yielding cleaner semantic matching. We then generate positive point prompts using similarity-weighted farthest-point sampling, paired with a refined dense similarity prior. Without any training or parameter updates, our approach establishes a new state of the art among training-free methods on PACO-Part, FSS-1000, and cross-domain datasets such as ISIC2018, demonstrating that explicit background subspace removal is a highly effective principle for one-shot localization.