同儕投票式LLM智能體壓力測試發現饋送誘發之詞彙趨同,但分散式來源未呈現可靠的匹配曝光優勢
Peer-Voted LLM-Agent Stress Tests Find Feed-Induced Lexical Convergence but No Reliable Matched-Exposure Advantage for Distributed Sources
August 20, 2026
作者: Rana Muhammad Usman, Dominic Williamson
cs.AI
摘要
大型語言模型(LLM)智能體在群體層面的行為無法以單一智能體基準來表徵。我們引入PV-SST(同儕投票社群平台測試平台),並報告一項分別凍結、預先註冊的匹配暴露實驗,涵蓋四個主題、四個未使用的隨機種子、四個開放權重模型系列,以及三個預先指定的較大變體。該實驗包含448次試驗與112個完整的模型×主題×種子區塊。相對於僅含主題的對照組,以前一輪同儕貼文(按同儕產生的讚數排名)組成的訊息流提高了最終回合的詞彙相似度:在四系列核心面板中(配對平均差異+0.0082個TF-IDF餘弦單位,95%區塊拔靴信賴區間 [0.0043, 0.0121],隨機化p=0.000105,n=64個區塊),以及在三變體規模擴展中(+0.0109 [0.0069, 0.0151],p=0.000001,n=48)。此對比同時涵蓋同儕貼文暴露與排名因素,因此無法識別僅由排名造成的效應。對立立場存續率在核心面板中下降(-3.9個百分點 [-6.8, -1.6],p=0.0068),但在較大變體中無確切結論(-1.0個百分點 [-3.1, 0.4],p=0.50)。在對抗性印象保持不變的情況下,四個分散來源對誠實智能體立場的影響,並不穩定地大於單一來源。預先註冊的分散來源減去單一來源對比在核心面板中為正向但無確切結論(+0.057 [-0.009, 0.125],p=0.112),在較大變體中則為負向(-0.040 [-0.113, 0.035],p=0.332),未通過預先指定的跨模型與跨主題一致性標準。因此,穩健的結果是在所測試的同儕排名訊息流下產生詞彙趨同,而非普遍的意見捕獲或普遍的協調優勢。本研究評估的是合成LLM智能體群體;其並不估計對人類或實際生產平台的效應。
English
Population-level behavior in large-language-model (LLM) agents cannot be characterized by single-agent benchmarks. We introduce PV-SST, a peer-voted social-platform testbed, and report a separately frozen, preregistered matched-exposure experiment spanning four topics, four unused seeds, four open-weight model families, and three prespecified larger variants. The experiment comprises 448 trials and 112 complete model-by-topic-by-seed blocks. Relative to a topic-only control, a feed of previous-round peer posts ranked by peer-generated likes increases final-round lexical similarity in both the four-family core panel (paired mean difference +0.0082 TF-IDF cosine units, 95% block-bootstrap CI [0.0043, 0.0121], randomization p=0.000105, n=64 blocks) and the three-variant size extension (+0.0109 [0.0069, 0.0151], p=0.000001, n=48). This contrast bundles peer-post exposure with ranking and therefore does not identify a ranking-only effect. Opposite-side survival falls in the core panel (-3.9 percentage points [-6.8, -1.6], p=0.0068) but not conclusively in the larger variants (-1.0 pp [-3.1, 0.4], p=0.50). Holding adversarial impressions fixed, four distributed sources do not reliably move honest-agent stance more than one source. The preregistered distributed-minus-single contrast is positive but inconclusive in the core panel (+0.057 [-0.009, 0.125], p=0.112) and negative in the larger variants (-0.040 [-0.113, 0.035], p=0.332), failing the prespecified cross-model and cross-topic consistency criterion. Thus the robust result is lexical convergence under the tested peer-ranked feed, not general opinion capture or a general coordination advantage. The study evaluates synthetic LLM-agent populations; it does not estimate effects on people or production platforms.