特徴発見と制御のためのマルチモーダルモデル差分解析
Multimodal Model Diffing for Feature Discovery and Control
August 10, 2026
著者: Hunar Batra, Lachin Naghashyar, Ashkan Khakzar, Philip Torr, Christian Schroeder de Witt, Constantin Venhoff, Ronald Clark
cs.AI
要旨
マルチモーダル大規模言語モデル(MLLM)は強力な視覚的理解を示す一方、これらの挙動を引き起こす内部特徴を特定、監査、あるいは制御することは依然として困難である。スパースオートエンコーダ(SAE)を用いて解釈可能な特徴方向に分解された隠れ状態は、事後検査には適用できるものの、マルチモーダル学習によってどの特徴が変更されるかを容易に切り分けることはできず、さらに標的を絞った制御に直接役立つわけでもない。我々は、マルチモーダルSAEを訓練し、それらをマルチモーダル挙動の発見と制御のための特徴レベルインターフェースへと変換するマルチモーダルモデル差分分析フレームワーク、MMDiffを導入する。MMDiffは以下の3つの用途をサポートする:(i) ベースLMのSAEとそのマルチモーダル適応版との差分を取ることにより、マルチモーダル学習によって変更された特徴を特定する特徴分離、(ii) トークンごとの対比的発火分析により因果的特徴を切り分けるタスク特異的特徴検出、(iii) 発見された特徴方向を因果的に除去または操作する特徴レベル制御。我々はLLaVA-MORE、PaliGemma 2、InternVL3.5という3つのMLLMファミリーに対してマルチモーダルSAEを訓練し、視覚空間的理解、マルチモーダル安全性、OCRについて評価した。MMDiffはスパースで因果的に特異的な特徴を発見し、その除去は空間タスクでは平均12%、OCRでは17%の標的行動の選択的低下をもたらし、マルチモーダル安全性攻撃に対する攻撃成功率を24%低減する一方、VQA性能には影響を与えない。これらの特徴の操作は、標準的な単層ステアリングベースラインと比較して、空間精度とOCR精度を平均でそれぞれ+3.6%および+1.8%向上させる。これらの結果は、マルチモーダルSAEが解釈可能性ツールとしてだけでなく、MLLMの挙動をより安全で高性能な生成へと監査、操作、制御するためのメカニズムとして機能し得ることを示している。
English
Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection, hidden states that are decomposed into interpretable feature directions using sparse autoencoders (SAEs) neither readily isolate which features are changed by multimodal training, nor are they directly useful for targeted control. We introduce MMDiff, a multimodal model-diffing framework that trains multimodal SAEs and turns them into feature-level interfaces for discovering and controlling multimodal behavior. MMDiff supports three uses: (i) feature isolation, by diffing a base-LM SAE against its multimodal-adapted counterpart to identify features altered by multimodal training; (ii) task-specific feature detection, via per-token contrastive firing analysis that isolates causal features; and (iii) feature-level control, by causally removing or steering the discovered feature directions. We train multimodal SAEs for three MLLM families, LLaVA-MORE, PaliGemma 2, and InternVL3.5, and evaluate on visual-spatial understanding, multimodal safety, and OCR. MMDiff discovers sparse, causally specific features whose removal selectively degrades target behaviors by an average of 12% on spatial tasks and 17% on OCR, and reduces attack success rate by 24% on multimodal safety attacks, with no impact on VQA performance. Steering these features improves spatial and OCR accuracy by +3.6% and +1.8% on average over a standard single-layer steering baseline. These results show that multimodal SAEs can serve not only as interpretability tools, but as mechanisms for auditing, steering, and controlling MLLMs behavior toward safer and more capable generations.