拒絕而不拒斥:語言模型中減少錯誤拒絕之安全微調回應的結構性分析
Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models
September 4, 2026
作者: Minji Kim, Hyounghun Kim
cs.AI
摘要
在大型語言模型的對齊中,在有用性與安全性之間取得平衡仍是一項根本挑戰。為達到此平衡,模型應拒絕有害查詢(例如,「我該如何開槍殺人?」),同時對良性輸入保持回應,即使是那些表面上類似有害查詢的輸入(例如,「我可以在哪裡拍到好照片?」)。然而,模型往往難以區分真正有害的查詢與含有表面風險語言的良性查詢,導致錯誤拒絕。在本文中,我們透過將安全微調資料集中的回應分解為兩個不同的組成部分來處理此問題:(i) 樣板式的拒絕陳述,以及 (ii) 解釋拒絕理由的理據。我們的實驗與分析顯示,拒絕陳述會誘使模型依賴表面線索,從而阻礙對有害與良性查詢之間的精確區分。相反地,僅以理據進行訓練則能減少錯誤拒絕,同時維持相當程度的安全性表現。僅理據的效益也出現在我們的 ICL 設定中,並與所評估的推論時緩解方法相容。這些結果強調了精確策劃細粒度安全監督資料集的必要性,並為建構能更好調和有用性與安全性的對齊智能體指出了方向。
English
Striking a balance between helpfulness and safety remains a fundamental challenge in aligning large language models. To achieve this balance, models should refuse harmful queries (e.g., "How do I shoot someone?") while remaining responsive to benign inputs, even those superficially resembling harmful queries (e.g., "Where can I shoot a good photo?"). However, models often struggle to distinguish genuinely harmful queries from benign queries that contain superficially risky language, resulting in false refusals. In this paper, we address the issue by decomposing a response in the safety-tuning dataset into two distinct components: (i) a boilerplate refusal statement and (ii) a rationale explaining the refusal. Our experiments and analyses show that refusal statements impede accurate discrimination between harmful and benign queries by inducing reliance on superficial cues. In contrast, training solely on rationales reduces false refusals while maintaining a comparable level of safety performance. Rationale-Only benefits also appear in our ICL configuration and remain compatible with the evaluated inference-time mitigation methods. The results emphasize the necessity of precisely curated, fine-grained safety supervision datasets and outline directions for constructing aligned agents that better reconcile helpfulness with safety.