ChatPaper.aiChatPaper

Partitioneren, Prompten, Aggregeren: Statistische Zelfconsistentie in Taalmodellen

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models

July 16, 2026
Auteurs: Patrik Wolf, Thomas Kleine Buening, Andreas Krause, Celestine Mendler-Dünner
cs.AI

Samenvatting

In-context learning wordt vaak geïnterpreteerd als een vorm van conditionele inferentie, waarbij de prompt een context specificeert en de output van het model wordt beschouwd als een schatting van de bijbehorende conditionele verdeling. Als deze interpretatie geldig is, dan zouden LLM-schattingen moeten voldoen aan elementaire probabilistische identiteiten. In het bijzonder stelt de wet van de totale kans dat a priori gewogen conditionele verdelingen aggregeren tot marginalen op populatieniveau over elke geldige partitie van de populatie. In dit werk onderzoeken we in hoeverre LLM-schattingen zich aan dit zelfconsistentieprincipe houden. We gebruiken binaire bomen als evaluatie-raamwerk om een populatie recursief te partitioneren in steeds fijnmazigere subpopulaties. Vervolgens geven we LLM's prompts met verwoorde subpopulatiebeschrijvingen in context, aggregeren we de resulterende schattingen terug naar schattingen op populatieniveau, en vergelijken we ze over partities van verschillende granulariteit. Door dit protocol toe te passen op verschillende probleemdomeinen en geavanceerde voorhoedemodellen, tonen we wijdverbreide schendingen van elementaire consistentie-eigenschappen aan. Een diepgaand onderzoek naar persona-prompting onthult een patroon dat we de macro-dwaling noemen: schattingen gereconstrueerd uit fijnmazigere subpopulatieresponsen komen vaak beter overeen met menselijke referentiegegevens dan directe schattingen op populatieniveau. Dit effect blijft bestaan over variaties in boomstructuur en schattingstaak, en kan gedeeltelijk worden hersteld door impliciete prompting. Samen suggereren deze bevindingen dat modellen relevante subpopulatiekennis bezitten maar deze niet betrouwbaar doorgeven in geaggregeerde schattingen. Deze kloof vestigt statistische zelfconsistentie als een onverzadigd, referentie-vrij criterium voor het evalueren van LLM's.
English
In-context learning is commonly interpreted as a form of conditional inference, in which the prompt specifies a context and the model's output is treated as an estimate of the corresponding conditional distribution. If this interpretation holds, then LLM estimates should satisfy basic probabilistic identities. In particular, the law of total probability asserts that prior-weighted conditional distributions aggregate into population-level marginals over any valid partition of the population. In this work, we investigate to what extent LLM estimates adhere to this self-consistency principle. We use binary trees as an evaluation scaffold to recursively partition a population into increasingly fine-grained subpopulations. We then prompt LLMs with verbalized subpopulation descriptions in context, aggregate the resulting estimates back into population-level estimates, and compare them across partitions of varying granularity. Applying this protocol across problem domains and state-of-the-art frontier models, we show widespread violations of basic consistency properties. An in-depth study of persona prompting reveals a pattern we call the macro fallacy: estimates reconstructed from more fine-grained subpopulation responses are often better aligned with human reference data than direct population-level estimates. This effect persists across variations in tree structure and estimation task, and can be partially recovered through implicit prompting. Together, these findings suggest that models possess relevant subpopulation knowledge but do not reliably propagate it into aggregate estimates. This gap establishes statistical self-consistency as an unsaturated, reference-free criterion for evaluating LLMs.