エージェントがエージェントを捉える:臨床マルチエージェントシステムにおけるショートカットカスケードとベンチマークゲーミング
Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
August 4, 2026
著者: Sebastián Andrés Cajas Ordóñez, Agastya Munnangi, Aldo Marzullo, Felipe Ocampo Osorio, Quang Bui, Mohammad Shahin, Armaan Grewal, Emmanuel Paul Kwesiga, Anqi Peter Li, Josephine Nanyonjo, Aaditya Panchal, Arshnoor Bhutani, Nikhil Jaiswal, Milit S. Patel, Maximin Lange, Leo Anthony Celi
cs.AI
要旨
臨床意思決定支援は、共有ワークスペース上で審議する言語モデルエージェントの委員会方式へと移行しつつある。我々は、そのような委員会が、ベンチマークが報いるが臨床医なら無視するような手がかり(ショートカット)によって攻略され得るかを問う。テキスト(MedQA-USMLE、MedMCQA、MIMIC-CXRレポート)、画像(NIH ChestX-ray14、MIMIC-CXR-JPG、CheXpert)、表形式ICU記録(SUPPORT2)にわたる6つの公開データセット上の7コホートにおいて、Gemini委員会はこれらの手がかりを単独では退ける(反転率5〜16%)が、社会的に妥当なショートカットは伝播する。すなわち、2人のピアが同じ誤答を主張すると、テスト対象のホールドアウトエージェントは38%のケースでそれを採用し、偽の「プレスクリーニング」システムフラグも、両方の性能レベルで同様の採用を引き起こす。3つの監視エージェントのうち、ゲートは同調による採用と正直な一致を区別できない(偽陽性率100%)。同一系統の判定エージェントは議事録のみを読んで、テキストでは採用を検出する(適合率100%、再現率93%)が、画像ではゲートと同様に機能しなくなる。ホールドアウトを非公開で再問い合わせするレフェリーは画像にも転移する(適合率77〜88%、偽陽性率13〜21%)。手がかりの視覚的顕著性を3倍にしても伝染度は変わらない一方、2人目のピアの発言はそれをさらに1.5倍に引き上げる。隠れたルーブリックを標的としたゲーミングはほぼ無言である:テキストでは10件中1件、画像では134件中1件のドリフターだけが、自分が近づいたルーブリックを名指しする。委員会を攻略するのは社会的妥当性であり、自己報告から独立したレフェリーだけがそれを捉える。コード:https://github.com/criticaldata/benchmaxxing
English
Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We ask whether such committees can be gamed by shortcuts, cues a benchmark rewards but a clinician would ignore. Across seven cohorts on six public datasets spanning text (MedQA-USMLE, MedMCQA, MIMIC-CXR reports), imaging (NIH ChestX-ray14, MIMIC-CXR-JPG, CheXpert) and tabular ICU records (SUPPORT2), Gemini committees resist these cues in isolation (flip 5-16%), yet a socially plausible shortcut spreads: when two peers assert the same wrong answer, the holdout under test adopts it in 38% of cases, as does a false "pre-screen" system flag, on both capability tiers. Of three oversight agents, a gate cannot separate adoption from honest agreement (false-positive rate 100%); a same-lineage judge reading only the transcript flags adoption on text (precision 100%, recall 93%) but collapses onto the gate in imaging; a referee that privately re-queries the holdout transfers to imaging (77-88% precision, 13-21% false-positive rate). Tripling a cue's visual salience does not move contagion, whereas a second peer voice raises it by half again. Gaming a hidden rubric is near-silent: only 1/10 text and 1/134 imaging drifters name the rubric they moved toward. What games a committee is social plausibility, and only a referee independent of self-report catches it. Code: https://github.com/criticaldata/benchmaxxing