誰のための安全性なのか?制御された大規模言語モデルの安全拒否のための境界認識型自己蒸留
Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal
September 3, 2026
著者: Alejo López-Ávila, Iker García-Ferrero, Jezabel Garcia, Antonio Tiene, Román Orús
cs.AI
要旨
安全性アライメントは通常、トピックレベルの問いとして定式化される。すなわち、この主題は有害か、という問いである。しかし、実運用ではより狭い問いが課される。公民教育のチューターと公共部門向けアシスタントは、基盤モデルを共有しつつも、同一トピック内で異なる境界を必要とする。すなわち、標的とされた政治的誘導は拒否しつつ、同一の選挙に関する事実的な質問には回答する、という具合である。本稿では、これを狭境界安全性(narrow-boundary safety)として定式化し、制御付きトピック生成、カバレッジ修復、分布内補償データ、および訓練と評価のための有害–無害ペアを組み合わせた、オフラインの自己生成フレームワークを導入する。単回生成では受理可能な拒否トレースが得られないプロンプトが19.88%残るのに対し、段階的再試行(エスカレーション)では0.20%まで低減する。Qwen3-8Bを用いた政治的説得タスクにおいて、Escalateを通じて完成させた拒否データによる訓練は、対象ドメインの拒否率を9.47%から84.75%へと向上させ、3つのより広範な有害性ベンチマークにおける平均不安全応答率を26.26%から0.14%へと低減する一方、XSTestにおける過剰拒否率は2.00%から74.00%へと増加する。別のマッチド比較では、外部応答を検証済みの対象モデル応答に置き換えることで、過剰拒否率は15.20%から5.20%へと低減する。境界ペアデータは、保持ペアにおけるコンプライアンス側の過剰拒否率を32.94%から4.16%へと低減する一方、有害側の拒否率は91.88%から87.72%への僅かな低下にとどまる。これらの結果は、データ構成が安全性と実用性のトレードオフを制御すること、そして安全性アライメントは意図された拒否境界の両側で評価されるべきであることを示している。
English
Safety alignment is usually posed as a topic-level question: is this subject harmful? Deployments ask a narrower one. A civics tutor and a public-sector assistant may share a base model yet need different boundaries inside the same topic, refusing targeted political manipulation while still answering factual questions about the same election. We formulate this as narrow-boundary safety and introduce an offline self-generated framework combining controlled topic generation, coverage repair, in-distribution compensation data, and harmful-benign pairs for training and evaluation. Single-shot generation leaves 19.88% of prompts without accepted refusal traces, whereas escalating retries leave 0.20%. On political persuasion with Qwen3-8B, training on refusal data completed through Escalate increases target-domain refusal from 9.47% to 84.75% and reduces the mean unsafe-response rate across three broader harmfulness benchmarks from 26.26% to 0.14%, but increases XSTest over-refusal from 2.00% to 74.00%. In a separate matched comparison, replacing external responses with verified target-model responses reduces over-refusal from 15.20% to 5.20%. Boundary-pair data reduces comply-side over-refusal on held-out pairs from 32.94% to 4.16%, while harmful-side refusal decreases only from 91.88% to 87.72%. These results show that data composition controls the safety and usability trade-off, and that safety alignment should be evaluated on both sides of the intended refusal boundary.