ChatPaper.aiChatPaper

CTスキャンにおける信頼性が高く監査可能な空間関係検証のためのモジュール式エージェント

A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans

August 21, 2026
著者: Simon Vincent Abel, Heiko Hillenhagen, Michael Götz, Timo Ropinski, Ayhan Can Erdur, Daniel Santak Wolf
cs.AI

要旨

放射線レポート生成と構造化画像理解を支援することを目的とした将来の医用視覚言語システムにとって、信頼性の高い空間理解は重要な前提条件である。現代の視覚言語モデル(VLM)は多くの医用画像タスクで有望な性能を示しているものの、最近のエビデンスは、それらが制御された空間推論において依然として弱く、空間関係を画像エビデンスに確実に接地できないことが多いことを示唆している。放射線学的推論が解剖学的構造と所見の相対位置の理解に依存していることを考慮すると、この空間的理解の弱さは診断精度にリスクをもたらす。本稿では、軸位CTスライスにおける二値空間関係検証のためのモジュール型医用画像エージェントを提示する。本システムは、空間的答えをエンドツーエンドで直接予測するのではなく、タスクを言語解析、解剖学的位置特定、決定的幾何検証という明示的な段階に分解する。自然言語クエリは構造化関係タプルに変換され、クエリ対象の臓器はYOLOベースの検出器で位置特定され、最終的な空間判定はオブジェクト中心から決定的幾何ルールを用いて計算される。本アプローチをホールドアウトされたMIRP空間QAベンチマークで評価し、代表的なエンドツーエンドVLMベースラインと比較する。最良のハイブリッド構成は94.1%の精度と94.2%のF1を達成し、精度において直接的なQwen2-VLプロンプティングを42.5パーセントポイント上回るとともに、解釈可能な中間表現と監査可能な推論段階を保持する。これらの結果は、明示的なモジュール型空間検証が、将来のレポート生成指向医用画像エージェントの有望な構成要素となり得ることを示唆している。
English
Reliable spatial understanding is an important prerequisite for future medical vision-language systems that aim to support radiological report generation and structured image understanding. While modern vision-language models (VLMs) show promising performance on many medical imaging tasks, recent evidence suggests they remain weak in controlled spatial reasoning and often fail to reliably ground spatial relations in image evidence. Given that radiological reasoning hinges on understanding the relative positions of anatomical structures and findings, this spatial weakness poses risks to diagnostic accuracy. We present a modular medical imaging agent for binary spatial relation verification in axial CT slices. Instead of directly predicting spatial answers end-to-end, the system decomposes the task into explicit stages: language parsing, anatomical localization, and deterministic geometric verification. Natural-language queries are converted into structured relation tuples, queried organs are localized with a YOLO-based detector, and the final spatial decision is computed from object centers using deterministic geometric rules. We evaluate the approach on the held-out MIRP spatial QA benchmark and compare it against representative end-to-end VLM baselines. The best-performing hybrid configuration reaches 94.1% accuracy and 94.2% F1, outperforming direct Qwen2-VL prompting by 42.5 percentage points in accuracy, while preserving interpretable intermediate representations and auditable reasoning stages. The results suggest that explicit modular spatial verification can serve as a promising building block for future report-oriented medical imaging agents.