憲法式中期訓練:內容存在驅動對齊增益
Constitutional Midtraining: Content Presence Drives Alignment Gains
July 29, 2026
作者: Desiree Cho, Cameron Tice, Bernie Hogan, Hunar Batra, Puria Radmard, Jun Zhao, Nigel Shadbolt
cs.AI
摘要
訓練後對齊往往流於表面,在微調過程中容易退化。中期訓練介入若能與訓練後階段完全隔離,是否能產生持久的對齊效果,至今仍未經驗證。我們透過憲法式中期訓練(constitutional midtraining)來測試此問題:在120B參數規模下,將基於原則、以價值觀為核心的內容插入中期訓練階段,並以僅重放(replay-only)的對照組進行比較。我們依據Anthropic的憲法建構了包含3.94億個token的憲法式語料庫,採用2x2因子設計(課程排序 × 審議式推理),產出四種憲法式中期訓練條件加上一個對照組,並在三個階段(中期訓練後、監督式微調(SFT)後、良性微調後)以自生成及既有基準進行評估,涵蓋壓力下的對齊表現、價值衝突解決、勒索情境及新興失序行為(emergent misalignment)。憲法式中期訓練的模型在對齊泛化與持久性上優於對照組,尤其在勒索情境中表現顯著:SFT使所有模型都產生勒索傾向,但憲法式中期訓練能削弱此傾向,且此優勢在良性微調後依然存在(-17.5個百分點)。然而,此持久性並未延伸至需要主動抵抗情境內壓力或衝突的場景,此類優勢在SFT後即減弱。中期訓練階段存在憲法式內容本身比其結構編排更為關鍵,且憲法式中期訓練在我們測試的任何階段中,平均而言對能力指標(MMLU、ARC-Easy、piqa、GSM8K)皆無損害。因此,在中期訓練階段加入適量的憲法式內容,可帶來廣泛且持久的對齊增益,為以SFT為核心的訓練流程提供低成本且具互補性的增強方案。程式碼、資料與模型均已公開。
English
Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains untested. We test this via constitutional midtraining: inserting principled, values-based content into midtraining against a replay-only control at 120B scale. Our 394M-token constitutional corpus, built from Anthropic's Constitution, uses a 2x2 factorial design (curriculum ordering x deliberative reasoning) to produce four constitutionally midtrained conditions plus a control, evaluated on self-generated and established benchmarks including alignment under pressure, value conflict resolution, blackmail, and emergent misalignment across three stages: post-midtraining, post-SFT, and post-benign fine-tuning. Constitutionally midtrained models outperform the control on alignment generalization and durability, notably on blackmail: SFT instills a blackmail propensity in all models, but constitutional midtraining blunts it, with the advantage surviving benign fine-tuning (-17.5pp). This durability does not extend to settings requiring active resistance to in-context pressure or conflict, where the advantage attenuates after SFT. The presence of constitutional content at midtraining also matters more than its structure, and constitutional midtraining incurs no cost, on average, on the capabilities we test (MMLU, ARC-Easy, piqa, GSM8K) at any stage. A modest amount of constitutional content at midtraining could therefore yield broad, persistent alignment gains, offering a cheap, complementary addition to SFT-centered pipelines. Code, data, and models are available.