CT 스캔에서 신뢰할 수 있고 감사 가능한 공간 관계 검증을 위한 모듈식 에이전트
A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans
August 21, 2026
저자: Simon Vincent Abel, Heiko Hillenhagen, Michael Götz, Timo Ropinski, Ayhan Can Erdur, Daniel Santak Wolf
cs.AI
초록
신뢰할 수 있는 공간 이해는 영상의학 보고서 생성과 구조화된 영상 이해를 지원하는 미래 의료 비전-언어 시스템의 중요한 전제 조건이다. 최신 비전-언어 모델(VLM)은 많은 의료 영상 작업에서 유망한 성능을 보여주지만, 최근 증거에 따르면 이들은 통제된 공간 추론에 여전히 취약하며 이미지 증거에 공간 관계를 안정적으로 근거화하지 못한다. 영상의학적 추론이 해부학적 구조와 소견의 상대적 위치를 이해하는 데 달려 있다는 점을 고려할 때, 이러한 공간적 취약성은 진단 정확도에 위험을 초래한다. 우리는 축상 CT 슬라이스에서 이진 공간 관계 검증을 위한 모듈식 의료 영상 에이전트를 제시한다. 이 시스템은 공간 답변을 종단 간(end-to-end)으로 직접 예측하는 대신, 언어 파싱, 해부학적 위치 파악, 결정론적 기하 검증의 명시적 단계로 작업을 분해한다. 자연어 질의는 구조화된 관계 튜플로 변환되고, 질의된 장기는 YOLO 기반 탐지기로 위치를 파악하며, 최종 공간 판단은 결정론적 기하 규칙을 사용하여 객체 중심점으로부터 계산된다. 우리는 보류된 MIRP 공간 QA 벤치마크에서 이 접근법을 평가하고 대표적인 종단 간 VLM 기준 모델과 비교한다. 최고 성능의 하이브리드 구성은 정확도 94.1%, F1 점수 94.2%를 달성하여 직접적인 Qwen2-VL 프롬프팅보다 정확도에서 42.5퍼센트 포인트를 능가하며, 해석 가능한 중간 표현과 감사 가능한 추론 단계를 유지한다. 이러한 결과는 명시적 모듈식 공간 검증이 향후 보고서 지향 의료 영상 에이전트를 위한 유망한 구성 요소가 될 수 있음을 시사한다.
English
Reliable spatial understanding is an important prerequisite for future medical vision-language systems that aim to support radiological report generation and structured image understanding. While modern vision-language models (VLMs) show promising performance on many medical imaging tasks, recent evidence suggests they remain weak in controlled spatial reasoning and often fail to reliably ground spatial relations in image evidence. Given that radiological reasoning hinges on understanding the relative positions of anatomical structures and findings, this spatial weakness poses risks to diagnostic accuracy. We present a modular medical imaging agent for binary spatial relation verification in axial CT slices. Instead of directly predicting spatial answers end-to-end, the system decomposes the task into explicit stages: language parsing, anatomical localization, and deterministic geometric verification. Natural-language queries are converted into structured relation tuples, queried organs are localized with a YOLO-based detector, and the final spatial decision is computed from object centers using deterministic geometric rules. We evaluate the approach on the held-out MIRP spatial QA benchmark and compare it against representative end-to-end VLM baselines. The best-performing hybrid configuration reaches 94.1% accuracy and 94.2% F1, outperforming direct Qwen2-VL prompting by 42.5 percentage points in accuracy, while preserving interpretable intermediate representations and auditable reasoning stages. The results suggest that explicit modular spatial verification can serve as a promising building block for future report-oriented medical imaging agents.