宪法中途训练:内容在场驱动对齐收益
Constitutional Midtraining: Content Presence Drives Alignment Gains
July 29, 2026
作者: Desiree Cho, Cameron Tice, Bernie Hogan, Hunar Batra, Puria Radmard, Jun Zhao, Nigel Shadbolt
cs.AI
摘要
训练后对齐往往较为浅层,容易在微调过程中发生退化。训练中干预能否与训练后阶段干净地分离开来,从而产生持久的对齐效果,这一问题尚未得到验证。我们通过宪法式训练中干预对此进行检验:在120B规模下,将基于原则、价值观的内容插入训练中阶段,并以仅重放作为对照组。我们基于Anthropic宪法构建了包含3.94亿个token的宪法式语料库,采用2×2因子设计(课程排序×审慎推理),产生了四种宪法式训练中条件及一个对照组,并在三个阶段的自我生成基准和既有基准上进行了评估,包括压力下的对齐、价值冲突解决、勒索以及涌现性错位,这三个阶段分别为:训练中阶段后、监督微调后以及良性微调后。宪法式训练中模型在对齐泛化性和持久性上优于对照组,尤其在勒索情境中表现显著:监督微调在所有模型中均诱发了勒索倾向,但宪法式训练中削弱了这一倾向,且该优势在良性微调后依然存在(-17.5个百分点)。然而,这种持久性并未延伸至需要积极抵抗上下文内压力或冲突的情境,在这些情境中该优势在监督微调后有所减弱。训练中阶段存在宪法内容比内容的结构方式更为重要,并且宪法式训练中在我们测试的能力指标(MMLU、ARC-Easy、piqa、GSM8K)上,在任何阶段平均均未带来损失。因此,在训练中阶段加入适量宪法内容有望产生广泛而持久的对齐收益,为以监督微调为核心的流水线提供了一种廉价且互补的补充方案。代码、数据和模型均已公开。
English
Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains untested. We test this via constitutional midtraining: inserting principled, values-based content into midtraining against a replay-only control at 120B scale. Our 394M-token constitutional corpus, built from Anthropic's Constitution, uses a 2x2 factorial design (curriculum ordering x deliberative reasoning) to produce four constitutionally midtrained conditions plus a control, evaluated on self-generated and established benchmarks including alignment under pressure, value conflict resolution, blackmail, and emergent misalignment across three stages: post-midtraining, post-SFT, and post-benign fine-tuning. Constitutionally midtrained models outperform the control on alignment generalization and durability, notably on blackmail: SFT instills a blackmail propensity in all models, but constitutional midtraining blunts it, with the advantage surviving benign fine-tuning (-17.5pp). This durability does not extend to settings requiring active resistance to in-context pressure or conflict, where the advantage attenuates after SFT. The presence of constitutional content at midtraining also matters more than its structure, and constitutional midtraining incurs no cost, on average, on the capabilities we test (MMLU, ARC-Easy, piqa, GSM8K) at any stage. A modest amount of constitutional content at midtraining could therefore yield broad, persistent alignment gains, offering a cheap, complementary addition to SFT-centered pipelines. Code, data, and models are available.