触覚に目覚めよ!MLLMsにおけるマスク分離触覚アライメント学習
Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs
July 1, 2026
著者: Yoonhyung Park, Minji Kim, Sungwon Moon, Jiyoung Lee
cs.AI
要旨
触覚は、摩擦やコンプライアンスといった本質的な物体の材料特性を知覚するために必要な物理的な基盤を提供するが、視覚だけではしばしばこれらの特性を識別できない。しかし、マルチモーダルLLMにこの触覚を付与する最近の試みでは、ゼロサム的なトレードオフが露呈している。すなわち、コンパクトなモデルの限られたパラメータ予算では、新たな感覚モダリティを獲得することと、既存の視覚・言語推論能力を維持することの間で選択を迫られるのである。本稿では、MLLM向けのマスク分離型触覚アライメント学習フレームワークであるSplashを提案する。Splashは、学習済みパラメータの重要度を定量化し、パラメータ空間を休眠部分空間と重要部分空間に分割する。重要な部分空間は凍結され、一般的な視覚知識を保護する安定したアンカーとして機能する一方、Splashは分離された休眠部分空間を更新し、LLMへの触覚アライメントを内在化させる。この選択的かつ非破壊的な拡張は、破滅的忘却を効果的に防止し、モダリティの非破壊的拡張を保証する。広範な実験により、SplashはLLM部分に追加の推論オーバーヘッドを発生させることなく触覚推論を実現し、SSVTP、TVL、TacQuadなどの視覚・触覚ベンチマークにおいて最先端の性能を示すと同時に、本来の汎用能力を維持することを実証した。
English
Touch supplies the physical grounding needed to perceive intrinsic material properties, such as friction and compliance, that vision alone often cannot resolve. Recent efforts for equipping multimodal LLMs with this tactile sense, however, expose a zero-sum trade-off: the limited parameter budget of compact models forces a choice between acquiring the new sensory modality and preserving the established vision-language reasoning. We present Splash, a mask-isolated tactile alignment learning framework for MLLMs. Splash quantifies the significance of each pretrained parameter, and partitions the parameter space into a dormant and critical subspace. While the frozen critical subspace acts as a stable anchor to safeguard general visual knowledge, Splash updates the isolated dormant subspace to internalize tactile alignment towards LLMs. This selective, non-destructive expansion effectively prevents catastrophic forgetting and ensures non-destructive modality expansion. Extensive experiments show that Splash effectively achieves tactile reasoning without additional inference overhead in the LLM part, demonstrating state-of-the-art performance on visuo-tactile benchmarks, including SSVTP, TVL, and TacQuad, while preserving its original general-purpose capabilities.