ChatPaper.aiChatPaper

分割、提示、聚合:语言模型中的统计自洽性

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models

July 16, 2026
作者: Patrik Wolf, Thomas Kleine Buening, Andreas Krause, Celestine Mendler-Dünner
cs.AI

摘要

上下文学习通常被解释为一种条件推断形式,其中提示指定了上下文,而模型的输出被视为对应条件分布的估计。若此解释成立,则大语言模型的估计应满足基本概率恒等式。特别地,全概率定律表明,先验加权条件分布可在总体任意有效划分下聚合为总体边际分布。本研究旨在探究大语言模型估计在多大程度上遵循这一自洽性原则。我们以二叉树作为评估框架,将总体递归划分为粒度渐细的子群体。随后,我们向大语言模型提供以自然语言表述的子群体描述作为上下文,将所得估计重新聚合为总体级估计,并在不同粒度划分间进行比较。将该方法应用于多个问题领域及最先进的前沿模型,我们发现基本一致性属性的广泛违反情况。对角色提示的深入研究表明一种称为“宏观谬误”的模式:由更细粒度子群体响应重建的估计通常比直接总体级估计更符合人类参考数据。该效应在树结构变化与估计任务变体中持续存在,且可通过隐式提示部分恢复。综上,这些发现表明模型具备相关子群体知识,但未能可靠地将其传播至聚合估计中。这一差距将统计自洽性确立为评估大语言模型的一个未饱和、无参考标准。
English
In-context learning is commonly interpreted as a form of conditional inference, in which the prompt specifies a context and the model's output is treated as an estimate of the corresponding conditional distribution. If this interpretation holds, then LLM estimates should satisfy basic probabilistic identities. In particular, the law of total probability asserts that prior-weighted conditional distributions aggregate into population-level marginals over any valid partition of the population. In this work, we investigate to what extent LLM estimates adhere to this self-consistency principle. We use binary trees as an evaluation scaffold to recursively partition a population into increasingly fine-grained subpopulations. We then prompt LLMs with verbalized subpopulation descriptions in context, aggregate the resulting estimates back into population-level estimates, and compare them across partitions of varying granularity. Applying this protocol across problem domains and state-of-the-art frontier models, we show widespread violations of basic consistency properties. An in-depth study of persona prompting reveals a pattern we call the macro fallacy: estimates reconstructed from more fine-grained subpopulation responses are often better aligned with human reference data than direct population-level estimates. This effect persists across variations in tree structure and estimation task, and can be partially recovered through implicit prompting. Together, these findings suggest that models possess relevant subpopulation knowledge but do not reliably propagate it into aggregate estimates. This gap establishes statistical self-consistency as an unsaturated, reference-free criterion for evaluating LLMs.