ChatPaper.aiChatPaper

憲法的中期訓練:コンテンツの存在がアライメント向上を促進する

Constitutional Midtraining: Content Presence Drives Alignment Gains

July 29, 2026
著者: Desiree Cho, Cameron Tice, Bernie Hogan, Hunar Batra, Puria Radmard, Jun Zhao, Nigel Shadbolt
cs.AI

要旨

ポストトレーニングによるアライメントはしばしば浅く、ファインチューニングによって損なわれる。ポストトレーニングから明確に分離されたミッドトレーニング介入が持続的なアライメントを生み出せるかどうかは、未検証のままである。我々はこれを、憲法的ミッドトレーニングによって検証する。すなわち、120B規模で、リプレイのみの対照群に対して、ミッドトレーニングに原理に基づく価値重視のコンテンツを挿入する。Anthropicの憲法から構築された我々の3億9400万トークンの憲法的コーパスは、2x2要因計画(カリキュラム順序×熟慮的推論)を用いて、4つの憲法的ミッドトレーニング条件と1つの対照条件を生成する。これらの条件は、自己生成および既存のベンチマークを用いて評価され、圧力下でのアライメント、価値葛藤の解決、恐喝、創発的ミスアライメントが、ミッドトレーニング後、SFT後、良性ファインチューニング後の3つの段階で評価される。憲法的ミッドトレーニングを施したモデルは、アライメントの一般化と持続性において対照群を上回り、特に恐喝において顕著である。すなわち、SFTはすべてのモデルに恐喝傾向を植え付けるが、憲法的ミッドトレーニングはその傾向を鈍らせ、その優位性は良性ファインチューニング後も維持される(-17.5pp)。しかし、この持続性は、文脈内の圧力や葛藤への能動的な抵抗を必要とする設定には及ばず、そこではSFT後に優位性は減衰する。また、ミッドトレーニングにおける憲法的コンテンツの存在は、その構造よりも重要である。さらに、憲法的ミッドトレーニングは、我々がテストした能力(MMLU、ARC-Easy、piqa、GSM8K)に関して、いずれの段階でも平均的にコストを課さない。したがって、ミッドトレーニングにおける適度な量の憲法的コンテンツは、広範で持続的なアライメントの向上をもたらす可能性があり、SFT中心のパイプラインへの安価で補完的な追加となる。コード、データ、モデルは公開されている。
English
Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains untested. We test this via constitutional midtraining: inserting principled, values-based content into midtraining against a replay-only control at 120B scale. Our 394M-token constitutional corpus, built from Anthropic's Constitution, uses a 2x2 factorial design (curriculum ordering x deliberative reasoning) to produce four constitutionally midtrained conditions plus a control, evaluated on self-generated and established benchmarks including alignment under pressure, value conflict resolution, blackmail, and emergent misalignment across three stages: post-midtraining, post-SFT, and post-benign fine-tuning. Constitutionally midtrained models outperform the control on alignment generalization and durability, notably on blackmail: SFT instills a blackmail propensity in all models, but constitutional midtraining blunts it, with the advantage surviving benign fine-tuning (-17.5pp). This durability does not extend to settings requiring active resistance to in-context pressure or conflict, where the advantage attenuates after SFT. The presence of constitutional content at midtraining also matters more than its structure, and constitutional midtraining incurs no cost, on average, on the capabilities we test (MMLU, ARC-Easy, piqa, GSM8K) at any stage. A modest amount of constitutional content at midtraining could therefore yield broad, persistent alignment gains, offering a cheap, complementary addition to SFT-centered pipelines. Code, data, and models are available.