피어 투표 기반 LLM 에이전트 스트레스 테스트에서 피드 유발 어휘 수렴은 확인되었으나 분산 소스에 대한 신뢰할 수 있는 일치 노출 이점은 발견되지 않음
Peer-Voted LLM-Agent Stress Tests Find Feed-Induced Lexical Convergence but No Reliable Matched-Exposure Advantage for Distributed Sources
August 20, 2026
저자: Rana Muhammad Usman, Dominic Williamson
cs.AI
초록
대규모 언어 모델(LLM) 에이전트의 집단 수준 행동은 단일 에이전트 벤치마크로는 특성화될 수 없다. 우리는 동료 투표 기반 소셜 플랫폼 테스트베드인 PV-SST를 소개하고, 네 가지 주제, 네 개의 미사용 시드, 네 개의 오픈웨이트 모델 계열, 세 개의 사전 지정된 대형 변형을 포괄하는 별도로 동결된, 사전 등록된 짝지은 노출 실험을 보고한다. 이 실험은 448회의 시행과 112개의 완전한 모델-주제-시드 블록으로 구성된다. 주제만 있는 통제 조건에 비해, 동료가 생성한 좋아요로 순위가 매겨진 이전 라운드 동료 게시물의 피드는 최종 라운드 어휘 유사도를 증가시킨다. 이는 네 계열 핵심 패널(쌍별 평균 차이 +0.0082 TF-IDF 코사인 단위, 95% 블록 부트스트랩 신뢰구간 [0.0043, 0.0121], 무작위화 p=0.000105, n=64 블록)과 세 변형 규모 확장(+0.0109 [0.0069, 0.0151], p=0.000001, n=48) 모두에서 나타난다. 이 대비는 동료 게시물 노출과 순위를 결합하므로 순위만의 효과를 식별하지 못한다. 반대측 생존율은 핵심 패널에서 감소하지만(-3.9% 포인트 [-6.8, -1.6], p=0.0068), 대형 변형에서는 결정적으로 나타나지 않는다(-1.0% 포인트 [-3.1, 0.4], p=0.50). 적대적 노출량을 고정했을 때, 네 개의 분산된 출처는 하나의 출처보다 정직한 에이전트의 입장을 안정적으로 움직이지 못한다. 사전 등록된 분산-단일 대비는 핵심 패널에서 양의 값을 보이나 결론적이지 않으며(+0.057 [-0.009, 0.125], p=0.112), 대형 변형에서는 음의 값을 보인다(-0.040 [-0.113, 0.035], p=0.332). 이는 사전 지정된 교차 모델 및 교차 주제 일관성 기준을 충족하지 못한다. 따라서 강건한 결과는 테스트된 동료 순위 피드에서의 어휘 수렴이지, 일반적인 의견 포획이나 일반적인 조정 이점이 아니다. 이 연구는 합성 LLM 에이전트 집단을 평가하며, 사람이나 실제 운영 플랫폼에 대한 효과를 추정하지 않는다.
English
Population-level behavior in large-language-model (LLM) agents cannot be characterized by single-agent benchmarks. We introduce PV-SST, a peer-voted social-platform testbed, and report a separately frozen, preregistered matched-exposure experiment spanning four topics, four unused seeds, four open-weight model families, and three prespecified larger variants. The experiment comprises 448 trials and 112 complete model-by-topic-by-seed blocks. Relative to a topic-only control, a feed of previous-round peer posts ranked by peer-generated likes increases final-round lexical similarity in both the four-family core panel (paired mean difference +0.0082 TF-IDF cosine units, 95% block-bootstrap CI [0.0043, 0.0121], randomization p=0.000105, n=64 blocks) and the three-variant size extension (+0.0109 [0.0069, 0.0151], p=0.000001, n=48). This contrast bundles peer-post exposure with ranking and therefore does not identify a ranking-only effect. Opposite-side survival falls in the core panel (-3.9 percentage points [-6.8, -1.6], p=0.0068) but not conclusively in the larger variants (-1.0 pp [-3.1, 0.4], p=0.50). Holding adversarial impressions fixed, four distributed sources do not reliably move honest-agent stance more than one source. The preregistered distributed-minus-single contrast is positive but inconclusive in the core panel (+0.057 [-0.009, 0.125], p=0.112) and negative in the larger variants (-0.040 [-0.113, 0.035], p=0.332), failing the prespecified cross-model and cross-topic consistency criterion. Thus the robust result is lexical convergence under the tested peer-ranked feed, not general opinion capture or a general coordination advantage. The study evaluates synthetic LLM-agent populations; it does not estimate effects on people or production platforms.