智能體互捕:臨床多智能體系統中的捷徑串聯與基準測試博弈
Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
August 4, 2026
作者: Sebastián Andrés Cajas Ordóñez, Agastya Munnangi, Aldo Marzullo, Felipe Ocampo Osorio, Quang Bui, Mohammad Shahin, Armaan Grewal, Emmanuel Paul Kwesiga, Anqi Peter Li, Josephine Nanyonjo, Aaditya Panchal, Arshnoor Bhutani, Nikhil Jaiswal, Milit S. Patel, Maximin Lange, Leo Anthony Celi
cs.AI
摘要
臨床決策支援正逐漸演變成由語言模型代理組成的委員會,在共享工作區上進行集體審議。我們要探討這類委員會是否可能被捷徑所欺騙,也就是那些基準測試會獎勵、但臨床醫師會忽略的線索。在涵蓋文字(MedQA-USMLE、MedMCQA、MIMIC-CXR報告)、影像(NIH ChestX-ray14、MIMIC-CXR-JPG、CheXpert)及表格型ICU紀錄(SUPPORT2)的六個公開資料集、七個隊列中,Gemini 委員會在單獨接觸這些線索時能抵抗(翻轉率5-16%);然而,一種社交上合理的捷徑會擴散:當兩個同儕斷言同一個錯誤答案時,受測的留出者有38%會採納該答案;虛假的「預篩選」系統標記也同樣會被採納,且兩個能力等級皆然。在三種監督代理中,閘門代理無法區分「採納」與「誠實同意」(假陽性率100%);僅閱讀轉錄稿的同源評判者能在文字任務中標記出採納行為(精確度100%,召回率93%),但在影像任務中則與閘門無異;而私下重新查詢留出者的裁判則可轉移至影像任務(精確度77-88%,假陽性率13-21%)。將線索的視覺顯著性增為三倍並不會改變傳染效應,但增加第二個同儕的聲音卻會使其再提高一半。利用隱藏評分標準進行欺騙幾乎不留痕跡:文字任務中僅有1/10、影像任務中僅有1/134的偏移者會說出他們所趨向的評分標準。能欺騙委員會的是社交上的合理性,而只有獨立於自我報告的裁判才能捕捉到它。程式碼:https://github.com/criticaldata/benchmaxxing
English
Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We ask whether such committees can be gamed by shortcuts, cues a benchmark rewards but a clinician would ignore. Across seven cohorts on six public datasets spanning text (MedQA-USMLE, MedMCQA, MIMIC-CXR reports), imaging (NIH ChestX-ray14, MIMIC-CXR-JPG, CheXpert) and tabular ICU records (SUPPORT2), Gemini committees resist these cues in isolation (flip 5-16%), yet a socially plausible shortcut spreads: when two peers assert the same wrong answer, the holdout under test adopts it in 38% of cases, as does a false "pre-screen" system flag, on both capability tiers. Of three oversight agents, a gate cannot separate adoption from honest agreement (false-positive rate 100%); a same-lineage judge reading only the transcript flags adoption on text (precision 100%, recall 93%) but collapses onto the gate in imaging; a referee that privately re-queries the holdout transfers to imaging (77-88% precision, 13-21% false-positive rate). Tripling a cue's visual salience does not move contagion, whereas a second peer voice raises it by half again. Gaming a hidden rubric is near-silent: only 1/10 text and 1/134 imaging drifters name the rubric they moved toward. What games a committee is social plausibility, and only a referee independent of self-report catches it. Code: https://github.com/criticaldata/benchmaxxing