ChatPaper.aiChatPaper

自律的かつ監査可能な医用画像モデル開発に向けて

Towards Autonomous and Auditable Medical Imaging Model Development

July 12, 2026
著者: Shengyuan Liu, Jia-Xuan Jiang, Boyun Zheng, Cheng Wang, Zipei Wang, Wentao Pan, Hongtao Wu, Houwen Peng, Yu Gu, Lichao Sun, Yixuan Yuan
cs.AI

要旨

大規模言語モデル(LLM)エージェントは、計画、コード実行、デバッグ、経験的フィードバックを連携させることで、機械学習エンジニアリング(MLE)の自動化を始めている。この能力を医用画像に適用することは依然として困難である。なぜなら、各タスクがモダリティ固有の実験、ならびに検証プロトコルと予測成果物に関する厳格な要件を課すからである。本稿では、医用画像モデル開発のための自律型マルチエージェントフレームワークAMIDを紹介する。AMIDはまず、データ条件付きメソッド計画(Data-Conditioned Method Planning)を提案する。これは、タスク固有のデータ分析と実行可能な医用画像リソースに基づいて、粗いタスクレベルの探索空間を実行可能かつ並列化可能なメソッドレーンに精緻化する。次に、検証誘導型二段階最適化(Verification-Guided Two-Stage Optimization)を開発する。これは、多様なメソッドレーンの広範な初期探索から、有望な候補の選択的活用へと移行し、最適化全体を通じて検証プロトコル、指標計算、予測成果物の厳格な検証を実施する。多様なモダリティと予測タイプにわたる20の医用画像チャレンジタスクにおいて、AMIDは評価された汎用MLEシステムを上回り、いくつかのタスクでは強力な人手によるチャレンジソリューションに近づくか同等の性能を示した。これらの結果は、AMIDがタスク固有の医用画像モデル開発を、特注の手動エンジニアリングから、異種タスクにわたって高性能かつ監査可能なモデル成果物を生成するエージェンティックワークフローへと転換できることを示唆している。
English
Large language model (LLM) agents are beginning to automate machine learning engineering (MLE) by coupling planning, code execution, debugging, and empirical feedback. Translating this capability to medical imaging remains difficult because each task imposes modality-specific experimentation and strict requirements for validation protocols and prediction artifacts. Here we introduce AMID, an autonomous multi-agent framework for medical imaging model development. AMID first proposes Data-Conditioned Method Planning, which refines coarse task-level search spaces into executable, parallelizable method lanes grounded in task-specific data analysis and runnable medical-imaging resources. It then develops Verification-Guided Two-Stage Optimization, moving from broad early exploration of diverse method lanes to selective exploitation of promising candidates while enforcing strict verification of validation protocols, metric computation, and prediction artifacts throughout the optimization. Across 20 medical imaging challenge tasks spanning diverse modalities and prediction types, AMID outperformed evaluated general-purpose MLE systems and, on several tasks, approached or matched strong human-designed challenge solutions. These results suggest that AMID can turn task-specific medical imaging model development from bespoke manual engineering into an agentic workflow for producing high-performing and auditable model artifacts across heterogeneous tasks.