에이전트가 에이전트를 적발하다: 임상 멀티에이전트 시스템에서의 지름길 연쇄와 벤치마크 게이밍
Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
August 4, 2026
저자: Sebastián Andrés Cajas Ordóñez, Agastya Munnangi, Aldo Marzullo, Felipe Ocampo Osorio, Quang Bui, Mohammad Shahin, Armaan Grewal, Emmanuel Paul Kwesiga, Anqi Peter Li, Josephine Nanyonjo, Aaditya Panchal, Arshnoor Bhutani, Nikhil Jaiswal, Milit S. Patel, Maximin Lange, Leo Anthony Celi
cs.AI
초록
임상 의사결정 지원은 공유 작업 공간에서 숙의하는 언어 모델 에이전트 위원회 방식으로 나아가고 있다. 우리는 그러한 위원회가 편법, 즉 벤치마크가 보상하지만 임상의라면 무시할 단서에 의해 조작될 수 있는지를 묻는다. 텍스트(MedQA-USMLE, MedMCQA, MIMIC-CXR 판독문), 영상(NIH ChestX-ray14, MIMIC-CXR-JPG, CheXpert), 표형식 중환자실 기록(SUPPORT2)을 아우르는 여섯 개 공개 데이터셋의 일곱 코호트에 걸쳐, Gemini 위원회는 단독으로 제시된 이러한 단서에는 저항하지만(반전률 5~16%), 사회적으로 개연성 있는 편법은 확산된다. 두 동료가 동일한 오답을 주장하면 시험 대상 홀드아웃은 38%의 사례에서 그 답을 채택하며, 허위 "사전 선별" 시스템 플래그도 두 성능 계층 모두에서 동일한 효과를 보인다. 세 명의 감독 에이전트 중에서, 게이트는 채택을 정직한 동의와 구분하지 못한다(오탐률 100%). 동일 계열 판정자는 전사본만 읽고 텍스트에서 채택을 정확하게 식별하지만(정밀도 100%, 재현율 93%), 영상에서는 게이트와 동일한 수준으로 붕괴된다. 홀드아웃을 비공개로 재질의하는 심판은 영상에서도 전이 가능하며(정밀도 77~88%, 오탐률 13~21%). 단서의 시각적 현저성을 세 배로 늘려도 전염은 움직이지 않지만, 두 번째 동료 목소리는 전염을 다시 절반만큼 높인다. 숨겨진 루브릭을 조작하는 것은 거의 드러나지 않는다. 텍스트의 10분의 1, 영상의 134분의 1의 이탈 사례만이 자신이 이동한 루브릭을 언급할 뿐이다. 위원회를 조작하는 것은 사회적 개연성이며, 자기 보고와 독립적인 심판만이 이를 잡아낸다. 코드: https://github.com/criticaldata/benchmaxxing
English
Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We ask whether such committees can be gamed by shortcuts, cues a benchmark rewards but a clinician would ignore. Across seven cohorts on six public datasets spanning text (MedQA-USMLE, MedMCQA, MIMIC-CXR reports), imaging (NIH ChestX-ray14, MIMIC-CXR-JPG, CheXpert) and tabular ICU records (SUPPORT2), Gemini committees resist these cues in isolation (flip 5-16%), yet a socially plausible shortcut spreads: when two peers assert the same wrong answer, the holdout under test adopts it in 38% of cases, as does a false "pre-screen" system flag, on both capability tiers. Of three oversight agents, a gate cannot separate adoption from honest agreement (false-positive rate 100%); a same-lineage judge reading only the transcript flags adoption on text (precision 100%, recall 93%) but collapses onto the gate in imaging; a referee that privately re-queries the holdout transfers to imaging (77-88% precision, 13-21% false-positive rate). Tripling a cue's visual salience does not move contagion, whereas a second peer voice raises it by half again. Gaming a hidden rubric is near-silent: only 1/10 text and 1/134 imaging drifters name the rubric they moved toward. What games a committee is social plausibility, and only a referee independent of self-report catches it. Code: https://github.com/criticaldata/benchmaxxing