같은 에이전트, 다른 답변: 검색 증강 QA에서 말뭉치 유발 답변 변동에 대한 반복 인지 감사
Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA
August 24, 2026
저자: Jingjie Ning, Xueqi Li
cs.AI
초록
검색 증강 질의응답 시스템은 요청된 모델 식별자, 프롬프트, 검색 정책, 증거 깊이, 렌더링, 노출된 생성 제어를 고정한 상태에서도 인덱스 확장 후 서로 다른 답변을 반환할 수 있다. 전체 정확도는 이득과 손실이 상쇄될 때 이러한 변화를 숨길 수 있으며, 일반적인 생성 변동성은 일회성 비교가 업데이트 효과를 과장하게 만든다. 우리는 이러한 숨은 현상을 정확도-비가시적 답변 변동(accuracy-blind answer churn)이라 부르며, 교차 스냅샷 불일치에서 동일 스냅샷 반복 불일치를 빼서 초과 답변 변동을 추정하는 스냅샷 호환성 감사(Snapshot Compatibility Audit)를 제안한다. 우리는 하나의 고정된 FineWeb 접두사를 1개에서 7개 샤드로 확장하여 이를 구현한다. 사전 등록된 400문항 Natural Questions 연구에서 정규화-정확 및 블라인드-의미 초과 변동은 각각 6.44 및 10.25 퍼센트 포인트인 반면, 정확 일치 정확도는 -1.50 포인트만 변화한다. 사후 분석에서는 40/400 문항에서 반복-안정 의미 반전이 발견된다. 별도로 사전 등록된 200문항 TriviaQA 연구에서는 더 작고 방향적으로 일관된 초과 변동이 나타나는 반면, 정확 일치 정확도는 반대 방향으로 움직인다. 두 번째 DeepSeek 생성기와 서빙 구성을 사용한 결과-눈가림 사후 100문항 하위 집합 복제에서는 정확 일치가 3.00 퍼센트 포인트 상승함에도 불구하고 8.75 퍼센트 포인트의 의미 초과 변동이 발견된다. 따라서 답변 수준 호환성은 눈에 띄거나 일관된 방향의 효용 변화 없이도 실패할 수 있다. 검색 증강 릴리스는 효용과 함께 호환성을 감사해야 한다.
English
A retrieval-augmented QA system can return different answers after an index expansion even when its requested model identifier, prompt, retrieval policy, evidence depth, rendering, and exposed generation controls are held fixed. Aggregate accuracy may hide these changes when gains and losses cancel, while ordinary generation variability makes one-shot comparisons overstate update effects. We call the hidden phenomenon accuracy-blind answer churn and introduce the Snapshot Compatibility Audit, which estimates excess answer churn by subtracting same-snapshot repeat disagreement from cross-snapshot disagreement. We instantiate it by expanding one frozen FineWeb prefix from one to seven shards. In a preregistered 400-question Natural Questions study, normalized-exact and blinded-semantic excess churn are 6.44 and 10.25 percentage points while exact-match accuracy changes by only -1.50 points. A post-hoc analysis finds repeat-stable semantic flips on 40/400 questions. A separately preregistered 200-question TriviaQA study yields smaller, directionally consistent excess churn while exact-match accuracy moves in the opposite direction. An outcome-blind post-hoc 100-question subset replication with a second DeepSeek generator and serving configuration finds 8.75 pp of semantic excess churn even as exact match rises by 3.00 percentage points. Answer-level compatibility can therefore fail without a conspicuous or consistently directed utility shift. Retrieval-augmented releases should audit compatibility alongside utility.