ChatPaper.aiChatPaper

HealthAgentBench: Een uniforme benchmark suite van realistische agentische gezondheidszorgomgevingen voor uitdagende grensverleggende AI-agenten

HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents

June 30, 2026
Auteurs: Qianchu Liu, Sheng Zhang, Guanghui Qin, Jeya Maria Jose Valanarasu, Maximilian Rokuss, Mingyu Lu, Timothy Ossowski, Juan Manuel Zambrano Chaves, Cliff Wong, Peniel Argaw, Yashna Hasija, Mu Wei, Wen-wai Yim, Qin Liu, Zilin Jing, Jason Entenmann, Naoto Usuyama, Tristan Naumann, Hoifung Poon
cs.AI

Samenvatting

Naarmate AI-agenten steeds beter worden in complex redeneren over langere termijnen, is rigoureuze en holistische evaluatie essentieel om de vooruitgang richting realistische toepassingen in de gezondheidszorg te meten. We introduceren HealthAgentBench, een suite van 54 agentische zorgtaken verdeeld over 7 categorieën, elk met een unieke omgeving. De benchmark-suite bestrijkt uiteenlopende workflows gedurende het volledige patiënttraject en een breed scala aan modaliteiten. Elke taak is ontworpen om een end-to-end klinische workflow na te bootsen: op basis van minimale instructies moet een agent ruwe zorgdata verkennen, opereren binnen een complexe omgeving en meerstapsoplossingen uitvoeren die verder gaan dan naïef prompten. Er wordt een uiteindelijk taaksuccespercentage gerapporteerd als een enkele, interpreteerbare maatstaf voor de algehele prestaties van elke agent op HealthAgentBench. Bij het evalueren van geavanceerde agenten op HealthAgentBench stellen we vast dat het algehele taaksuccespercentage laag blijft, wat de moeilijkheidsgraad van de suite onderstreept. De sterkste en meest kosteneffectieve agent, Codex GPT-5.5, behaalt slechts een succespercentage van ongeveer 42%. Naast de geaggregeerde prestaties onthult HealthAgentBench genuanceerde sterktes en zwaktes over taakcategorieën heen. Geavanceerde agenten tonen veelbelovende resultaten in het automatisch ontwikkelen van onderzoeksmodelpijplijnen op basis van EPD-gegevens, maar medische beeldvorming blijft bijzonder uitdagend, met name voor Claude Code-modellen, terwijl Codex GPT-5.5 opkomende capaciteiten laat zien. Taken die grote zoekruimtes combineren met compositionele redeneervereisten blijven moeilijk voor alle huidige agenten. Samen wijzen deze resultaten erop dat HealthAgentBench een uitdagende en realistische benchmark biedt met aanzienlijke ruimte voor toekomstige vooruitgang. We publiceren onze benchmark op https://github.com/microsoft/HealthAgentBench.
English
As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare applications. We introduce HealthAgentBench, a suite of 54 agentic healthcare tasks across 7 categories each with its unique environment. The benchmark suite spans diverse workflows throughout the patient journey and a broad range of modalities. Each task is designed to replicate an end-to-end clinical workflow: given minimal instructions, an agent must explore raw healthcare data, operate within a complex environment, and execute multi-step solutions that go beyond naive prompting. A final task success rate is reported to provide a single, interpretable metric for HealthAgentBench overall performance for each agent. Evaluating frontier agents on HealthAgentBench, we find that overall task success rate remains low, underscoring the difficulty of the suite. The strongest and the most cost effective agent, Codex GPT-5.5, achieves only approximately 42% success rate. Beyond aggregate performance, HealthAgentBench reveals nuanced strengths and weaknesses across task categories. Frontier agents show promise in automatically developing research modeling pipelines over EHR data, but medical imaging remains especially challenging, particularly for Claude Code models, while Codex GPT-5.5 shows emerging capability. Tasks that combine large search spaces with compositional reasoning requirements remain difficult for all current agents. Together, these results suggest that HealthAgentBench provides a challenging and realistic benchmark with substantial room for future progress. We release our benchmark at https://github.com/microsoft/HealthAgentBench.