ChatPaper.aiChatPaper

ステアリング幾何学:LLMステアリング空間における人間価値幾何学の検証

Steering Geometry: Validating Human Value Geometry in LLM Steering Space

September 5, 2026
著者: Mohammad Mahdi Abootorabi, Armin Saghafian, Ali Bazshoushtari, Hamid Rezaei, EunJeong Hwang, Vered Shwartz, Parvin Mousavi, Purang Abolmaesumi
cs.AI

要旨

大規模言語モデル(LLM)がアラインメントに敏感な文脈でますます展開されるにつれ、活性化ステアリングは、行動制御のための微調整手法(例:RLHF、DPO)に対する軽量な推論時代替手段として台頭してきた。しかし、既存研究は通常、ステアリングを個別の行動で検証しており、ステアリングベクトルが一貫した意味構造を符号化しているのか、それとも単に行動固有のショートカットを利用しているのかは不明確なままである。本稿では、LLMステアリングベクトルの潜在幾何構造が、人間の価値観と道徳性における理論上規定された構造を反映しているかどうかを調査する。主要な細粒度フレームワークとしてシュワルツの基本的人間価値理論を用い、20種類の人間価値を網羅する26Kサンプルのベンチマークを導入し、多様なモデルファミリーとサイズにわたり、分布駆動型手法(例:CAA、SphericalSteer、ODESteer)と行動中心型アプローチ(例:COLD-Steer、BiPO)を分析する。分布駆動型手法は、理論的予測と整合する人間価値のトポロジーを復元することが分かった(スピアマンρは最大0.51、p < 10^{-13})。対照的に、行動中心型手法は同等のステアリング性能を達成するが、期待される価値の幾何構造とはほとんど相関を示さない。幾何学的忠実性はモデル規模とともに向上するが、指示チューニング後には低下する。最後に、より良い幾何学的整合性は、価値間のより人間と整合的な転移にもつながる。すなわち、ある価値をステアリングすると、適合する価値が正しく高まり、対立する価値が抑制される。コードとデータは以下で利用可能である:https://github.com/DeepRCL/Steering_Geometry。
English
As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inference-time alternative to fine-tuning methods (e.g., RLHF, DPO) for behavioral control. However, existing work typically validates steering on isolated behaviors, leaving it unclear whether steering vectors encode coherent semantic structure or merely exploit behavior-specific shortcuts. We investigate whether the latent geometry of LLM steering vectors reflects theory-specified structure in human values and morality. Using Schwartz's Theory of Basic Human Values as our primary fine-grained framework, we introduce a 26K-sample benchmark covering 20 human values and analyze distribution-driven methods (e.g., CAA, SphericalSteer, ODESteer) and behavior-centric approaches (e.g., COLD-Steer, BiPO) across diverse model families and sizes. We find that distribution-driven methods recover human value topologies aligned with theoretical predictions (Spearman ρ up to 0.51, p < 10^{-13}). In contrast, behavior-centric methods achieve comparable steering performance but show little correlation with the expected value geometry. Geometric fidelity improves with model scale but drops after instruction tuning. Finally, better geometric alignment also leads to more human-consistent transfer across values: steering one value correctly lifts compatible values and suppresses opposing ones. Code and data are available at: https://github.com/DeepRCL/Steering_Geometry.