多模態模型差分:特徵發現與控制
Multimodal Model Diffing for Feature Discovery and Control
August 10, 2026
作者: Hunar Batra, Lachin Naghashyar, Ashkan Khakzar, Philip Torr, Christian Schroeder de Witt, Constantin Venhoff, Ronald Clark
cs.AI
摘要
多模態大型語言模型(MLLMs)展現出強大的視覺理解能力,然而導致這些行為的內部特徵仍難以識別、審計或控制。雖然稀疏自編碼器(SAEs)可將隱藏狀態分解為可解釋的特徵方向,適用於事後檢查,但此類分解既難以直接隔離多模態訓練所改變的特徵,也無法直接用於目標性控制。我們提出MMDiff,一個多模態模型差異分析框架,透過訓練多模態SAEs並將其轉化為特徵層級的介面,以發現和控制多模態行為。MMDiff支援三種用途:(i)特徵隔離,透過將基礎語言模型SAE與其多模態適配版本進行差異分析,以識別多模態訓練所改變的特徵;(ii)任務特定特徵偵測,透過逐詞元對比激發分析,隔離因果特徵;(iii)特徵層級控制,透過因果性地移除或引導所發現的特徵方向。我們為三個MLLM系列——LLaVA-MORE、PaliGemma 2和InternVL3.5——訓練多模態SAEs,並在視覺空間理解、多模態安全和OCR任務上進行評估。MMDiff發現稀疏且具因果特異性的特徵,移除這些特徵會選擇性地降低目標行為表現,在空間任務上平均降低12%,在OCR上降低17%,並將多模態安全攻擊的攻擊成功率降低24%,且對VQA效能沒有影響。引導這些特徵相較於標準的單層引導基準,在空間和OCR準確率上平均分別提升+3.6%和+1.8%。這些結果表明,多模態SAEs不僅可作為可解釋性工具,更可作為審計、引導和控制MLLMs行為的機制,以朝向更安全且更強大的生成結果。
English
Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection, hidden states that are decomposed into interpretable feature directions using sparse autoencoders (SAEs) neither readily isolate which features are changed by multimodal training, nor are they directly useful for targeted control. We introduce MMDiff, a multimodal model-diffing framework that trains multimodal SAEs and turns them into feature-level interfaces for discovering and controlling multimodal behavior. MMDiff supports three uses: (i) feature isolation, by diffing a base-LM SAE against its multimodal-adapted counterpart to identify features altered by multimodal training; (ii) task-specific feature detection, via per-token contrastive firing analysis that isolates causal features; and (iii) feature-level control, by causally removing or steering the discovered feature directions. We train multimodal SAEs for three MLLM families, LLaVA-MORE, PaliGemma 2, and InternVL3.5, and evaluate on visual-spatial understanding, multimodal safety, and OCR. MMDiff discovers sparse, causally specific features whose removal selectively degrades target behaviors by an average of 12% on spatial tasks and 17% on OCR, and reduces attack success rate by 24% on multimodal safety attacks, with no impact on VQA performance. Steering these features improves spatial and OCR accuracy by +3.6% and +1.8% on average over a standard single-layer steering baseline. These results show that multimodal SAEs can serve not only as interpretability tools, but as mechanisms for auditing, steering, and controlling MLLMs behavior toward safer and more capable generations.