智能体捕捉智能体:临床多智能体系统中的捷径级联与基准测试博弈
Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
August 4, 2026
作者: Sebastián Andrés Cajas Ordóñez, Agastya Munnangi, Aldo Marzullo, Felipe Ocampo Osorio, Quang Bui, Mohammad Shahin, Armaan Grewal, Emmanuel Paul Kwesiga, Anqi Peter Li, Josephine Nanyonjo, Aaditya Panchal, Arshnoor Bhutani, Nikhil Jaiswal, Milit S. Patel, Maximin Lange, Leo Anthony Celi
cs.AI
摘要
临床决策支持正朝着由语言模型智能体组成的委员会在共享工作空间中进行审议的方向发展。我们探究此类委员会是否会被捷径所操纵——即基准测试予以奖励、而临床医生会忽略的提示线索。在覆盖文本(MedQA-USMLE、MedMCQA、MIMIC-CXR报告)、影像(NIH ChestX-ray14、MIMIC-CXR-JPG、CheXpert)和表格型ICU记录(SUPPORT2)的六个公开数据集上的七个队列中,Gemini委员会在孤立呈现这些线索时能够予以抵抗(翻转率5-16%),然而一种社交合理的捷径依然会传播:当两个同侪断言同一错误答案时,被测保留模型在38%的案例中会采纳该答案;虚假的"预筛查"系统标记在两个能力层级上同样如此。在三种监督智能体中,门控器无法区分受操纵后的采纳与真诚的认同(假阳性率100%);仅读取对话记录的同源判官能在文本任务上识别受操纵采纳(精确率100%,召回率93%),但在影像任务上其表现退化为与门控器相同;而私下重新查询保留模型的裁判可迁移至影像任务(精确率77-88%,假阳性率13-21%)。将线索的视觉显著性提高至三倍并不能改变传染效应,而增加第二个同侪的声音则使传染率再提升一半。利用隐藏评分标准进行操纵几乎无声无息:文本任务中仅1/10、影像任务中仅1/134的发生漂移样本提及了它们所趋向的评分标准。真正能操纵委员会的是社交合理性,且只有独立于自我报告的裁判方能识别此类操纵。代码:https://github.com/criticaldata/benchmaxxing
English
Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We ask whether such committees can be gamed by shortcuts, cues a benchmark rewards but a clinician would ignore. Across seven cohorts on six public datasets spanning text (MedQA-USMLE, MedMCQA, MIMIC-CXR reports), imaging (NIH ChestX-ray14, MIMIC-CXR-JPG, CheXpert) and tabular ICU records (SUPPORT2), Gemini committees resist these cues in isolation (flip 5-16%), yet a socially plausible shortcut spreads: when two peers assert the same wrong answer, the holdout under test adopts it in 38% of cases, as does a false "pre-screen" system flag, on both capability tiers. Of three oversight agents, a gate cannot separate adoption from honest agreement (false-positive rate 100%); a same-lineage judge reading only the transcript flags adoption on text (precision 100%, recall 93%) but collapses onto the gate in imaging; a referee that privately re-queries the holdout transfers to imaging (77-88% precision, 13-21% false-positive rate). Tripling a cue's visual salience does not move contagion, whereas a second peer voice raises it by half again. Gaming a hidden rubric is near-silent: only 1/10 text and 1/134 imaging drifters name the rubric they moved toward. What games a committee is social plausibility, and only a referee independent of self-report catches it. Code: https://github.com/criticaldata/benchmaxxing