ChatPaper.aiChatPaper

자율적이고 감사 가능한 의료 영상 모델 개발을 향하여

Towards Autonomous and Auditable Medical Imaging Model Development

July 12, 2026
저자: Shengyuan Liu, Jia-Xuan Jiang, Boyun Zheng, Cheng Wang, Zipei Wang, Wentao Pan, Hongtao Wu, Houwen Peng, Yu Gu, Lichao Sun, Yixuan Yuan
cs.AI

초록

대규모 언어 모델(LLM) 에이전트는 계획 수립, 코드 실행, 디버깅, 실증적 피드백을 결합하여 머신러닝 엔지니어링(MLE)을 자동화하기 시작하고 있다. 이러한 역량을 의료 영상에 적용하는 것은 각 과제가 특정 양식에 특화된 실험과 검증 프로토콜 및 예측 결과물에 대한 엄격한 요구사항을 부과하기 때문에 여전히 어렵다. 본 논문에서는 의료 영상 모델 개발을 위한 자율적 다중 에이전트 프레임워크인 AMID를 소개한다. AMID는 먼저 데이터 조건 기반 방법 계획(Data-Conditioned Method Planning)을 제안하는데, 이는 작업 수준의 탐색 공간을 과제별 데이터 분석 및 실행 가능한 의료 영상 자원에 기반하여 실행 가능하고 병렬화가 가능한 방법 경로(method lanes)로 세분화한다. 이후 검증 기반 2단계 최적화(Verification-Guided Two-Stage Optimization)를 개발하여, 다양한 방법 경로에 대한 광범위한 초기 탐색에서 유망한 후보에 대한 선별적 활용으로 전환하며, 최적화 전반에 걸쳐 검증 프로토콜, 지표 계산, 예측 결과물에 대한 엄격한 검증을 강제한다. 다양한 양식과 예측 유형을 아우르는 20개의 의료 영상 챌린지 과제에 대해 AMID는 평가된 범용 MLE 시스템보다 뛰어난 성능을 보였으며, 여러 과제에서는 강력한 인간 설계 챌린지 솔루션에 근접하거나 일치하는 결과를 달성했다. 이러한 결과는 AMID가 과제 특화 의료 영상 모델 개발을 맞춤형 수동 엔지니어링에서 이질적인 과제 전반에 걸쳐 고성능의 감사 가능한 모델 결과물을 생성하는 에이전트 기반 워크플로우로 전환할 수 있음을 시사한다.
English
Large language model (LLM) agents are beginning to automate machine learning engineering (MLE) by coupling planning, code execution, debugging, and empirical feedback. Translating this capability to medical imaging remains difficult because each task imposes modality-specific experimentation and strict requirements for validation protocols and prediction artifacts. Here we introduce AMID, an autonomous multi-agent framework for medical imaging model development. AMID first proposes Data-Conditioned Method Planning, which refines coarse task-level search spaces into executable, parallelizable method lanes grounded in task-specific data analysis and runnable medical-imaging resources. It then develops Verification-Guided Two-Stage Optimization, moving from broad early exploration of diverse method lanes to selective exploitation of promising candidates while enforcing strict verification of validation protocols, metric computation, and prediction artifacts throughout the optimization. Across 20 medical imaging challenge tasks spanning diverse modalities and prediction types, AMID outperformed evaluated general-purpose MLE systems and, on several tasks, approached or matched strong human-designed challenge solutions. These results suggest that AMID can turn task-specific medical imaging model development from bespoke manual engineering into an agentic workflow for producing high-performing and auditable model artifacts across heterogeneous tasks.