ChatPaper.aiChatPaper

Safety for Whom? 邊界感知自蒸餾實現受控的大型語言模型安全拒絕

Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal

September 3, 2026
作者: Alejo López-Ávila, Iker García-Ferrero, Jezabel Garcia, Antonio Tiene, Román Orús
cs.AI

摘要

安全對齊通常被表述為一個主題層級的問題:這個主題是否有害?然而實際部署所提出的問題更為狹隘。公民教育導師與公共部門助理可能共享同一個基礎模型,卻需要在同一主題內設定不同的界線——拒絕針對性的政治操控,同時仍回答關於同一場選舉的事實性問題。我們將此問題形式化為「狹窄邊界安全」,並提出一個離線自生成的框架,結合受控主題生成、覆蓋率修復、分布內補償資料,以及用於訓練與評估的有害-無害配對。單次生成會使19.88%的提示缺乏可被接受的拒答痕跡,而逐步升級的重試機制則將其降至0.20%。在以Qwen3-8B進行的政治說服任務中,使用透過Escalate完成的拒答資料進行訓練,可將目標領域的拒答率從9.47%提升至84.75%,並將三個更廣泛的傷害性基準測試中的平均不安全回應率從26.26%降至0.14%,但同時使XSTest的過度拒答率從2.00%上升至74.00%。在另一項配對比較中,以經過驗證的目標模型回應取代外部回應,可將過度拒答率從15.20%降至5.20%。邊界配對資料可將保留配對中順從側的過度拒答率從32.94%降至4.16%,而有害側的拒答率僅從91.88%小幅下降至87.72%。這些結果顯示,資料組成控制了安全性与可用性之間的取捨,且安全對齊應在預設拒答邊界的兩側同時進行評估。
English
Safety alignment is usually posed as a topic-level question: is this subject harmful? Deployments ask a narrower one. A civics tutor and a public-sector assistant may share a base model yet need different boundaries inside the same topic, refusing targeted political manipulation while still answering factual questions about the same election. We formulate this as narrow-boundary safety and introduce an offline self-generated framework combining controlled topic generation, coverage repair, in-distribution compensation data, and harmful-benign pairs for training and evaluation. Single-shot generation leaves 19.88% of prompts without accepted refusal traces, whereas escalating retries leave 0.20%. On political persuasion with Qwen3-8B, training on refusal data completed through Escalate increases target-domain refusal from 9.47% to 84.75% and reduces the mean unsafe-response rate across three broader harmfulness benchmarks from 26.26% to 0.14%, but increases XSTest over-refusal from 2.00% to 74.00%. In a separate matched comparison, replacing external responses with verified target-model responses reduces over-refusal from 15.20% to 5.20%. Boundary-pair data reduces comply-side over-refusal on held-out pairs from 32.94% to 4.16%, while harmful-side refusal decreases only from 91.88% to 87.72%. These results show that data composition controls the safety and usability trade-off, and that safety alignment should be evaluated on both sides of the intended refusal boundary.