分割、提示、聚合:語言模型中的統計自一致性
Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models
July 16, 2026
作者: Patrik Wolf, Thomas Kleine Buening, Andreas Krause, Celestine Mendler-Dünner
cs.AI
摘要
上下文學習通常被解讀為一種條件推論形式,其中提示指定了背景脈絡,而模型的輸出則被視為對應條件分布之估計。若此解讀成立,則大型語言模型的估計應滿足基本的機率恆等式。具體而言,全機率定律主張先驗加權的條件分布會在任何有效的族群分割上匯總成族群層級的邊際分布。在本研究中,我們探討大型語言模型的估計在多大程度上遵循此自洽性原則。我們使用二元樹作為評估框架,藉以遞迴分割族群,形成日益細緻的子群體。接著,我們透過提示向大型語言模型提供語言化的子群體描述,將由此得出的估計結果匯總回族群層級,並比較不同粒度分割下的估計值。將此方法應用於多個問題領域及當前最先進的前沿模型後,我們發現基本一致性性質普遍遭到違反。針對角色提示的深入分析揭示了一種我們稱之為「宏觀謬誤」的模式:從更細緻子群體回應中重建的估計,往往比直接取得的族群層級估計更貼近人類參考數據。此效應在樹狀結構與估計任務的變化中持續存在,並可經由隱式提示部分復現。綜合而言,這些發現顯示模型具備相關的子群體知識,但無法可靠地將其傳播至匯總估計中。此差距使得統計自洽性成為一項未飽和、且無需參考點的評估大型語言模型之準則。
English
In-context learning is commonly interpreted as a form of conditional inference, in which the prompt specifies a context and the model's output is treated as an estimate of the corresponding conditional distribution. If this interpretation holds, then LLM estimates should satisfy basic probabilistic identities. In particular, the law of total probability asserts that prior-weighted conditional distributions aggregate into population-level marginals over any valid partition of the population. In this work, we investigate to what extent LLM estimates adhere to this self-consistency principle. We use binary trees as an evaluation scaffold to recursively partition a population into increasingly fine-grained subpopulations. We then prompt LLMs with verbalized subpopulation descriptions in context, aggregate the resulting estimates back into population-level estimates, and compare them across partitions of varying granularity. Applying this protocol across problem domains and state-of-the-art frontier models, we show widespread violations of basic consistency properties. An in-depth study of persona prompting reveals a pattern we call the macro fallacy: estimates reconstructed from more fine-grained subpopulation responses are often better aligned with human reference data than direct population-level estimates. This effect persists across variations in tree structure and estimation task, and can be partially recovered through implicit prompting. Together, these findings suggest that models possess relevant subpopulation knowledge but do not reliably propagate it into aggregate estimates. This gap establishes statistical self-consistency as an unsaturated, reference-free criterion for evaluating LLMs.