유방암 분류를 위한 토큰 기반 이중 뷰 융합 및 대규모 비전 모델 적응
Token-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer Classification
July 7, 2026
저자: Aysan Ghayouri Pirsoltan, Shima Babakordi, Mohammad Reza Mohammadi
cs.AI
초록
유방촬영술을 통한 정확한 유방암 분류는 두개미측(CC) 및 내외사위(MLO) 영상의 상호보완적 정보를 효과적으로 통합해야 하며, 이를 통해 유방 이상에 대한 더 완전한 특성화가 가능해진다. 그러나 기존의 다중 뷰 학습 접근법은 일반적으로 특징 수준의 집합 또는 단일 단계 교차 주의(cross-attention)에 의존하는데, 이는 뷰 특이적 표현과 공유 표현을 혼동시키고 상호작용을 제한된 네트워크 깊이로 제한할 수 있다. 이러한 한계를 해결하기 위해, 우리는 고정된 비전 트랜스포머 백본 내에서 프롬프트 기반 적응과 뷰 간 융합을 통합하는 토큰 중심의 이중 뷰 학습 프레임워크를 제안한다. 이 프레임워크는 뷰 간 상호작용을 구조화된 토큰 수준 통신으로 재구성하며, 여기서 전용 융합 토큰이 교차 주의를 통해 CC 뷰와 MLO 뷰 간의 양방향 정보 교환을 명시적으로 인코딩하여 직접적인 특징 융합에 의존하지 않고 뷰 간 의존성의 중간 전달자 역할을 한다. 단일 계층에서 융합을 적용하는 기존 방법과 달리, 융합 모듈이 여러 트랜스포머 깊이에 삽입되어 인코더 계층 전반에 걸쳐 점진적이고 반복적인 상호작용을 가능하게 한다. 융합 토큰은 토큰 시퀀스에 재통합되고 후속 트랜스포머 계층에 의해 정제되어, 뷰 특이적 구조를 보존하면서 상보적 정보의 계층적 전파를 촉진한다. VinDr-Mammo 및 CMMD 데이터셋에 대한 실험은 선형 프로빙, 프롬프트 전용 적응 및 기존 융합 기준선에 비해 일관된 개선을 보여준다. VinDr-Mammo BI-RADS 분류 작업에서 프레임워크는 50.40%의 F1 점수와 0.8090의 AUC를 달성하며, 이진 설정에서 이중 뷰 융합 기준선 대비 AUC 0.10 향상을 포함한다. 절제 연구는 토큰 기반 융합 및 다중 깊이 상호작용 설계의 효과성을 추가로 검증한다.
English
Accurate breast cancer classification from mammography requires effective integration of complementary information from craniocaudal (CC) and mediolateral oblique (MLO) views, which provide a more complete characterization of breast abnormalities. However, existing multi-view learning approaches typically rely on feature-level aggregation or single-stage cross-attention, which can entangle view-specific and shared representations and restrict interaction to limited network depths. To address these limitations, we propose a token-centric dual-view learning framework that unifies prompt-based adaptation and cross-view fusion within a frozen vision transformer backbone. The framework reformulates inter-view interaction as structured token-level communication, where dedicated fusion tokens explicitly encode bidirectional information exchange between CC and MLO views via cross-attention, serving as intermediate carriers of cross-view dependencies rather than relying on direct feature fusion. Unlike conventional methods that apply fusion at a single layer, fusion modules are inserted at multiple transformer depths, enabling progressive and repeated interaction across the encoder hierarchy. Fusion tokens are reintegrated into the token sequence and refined by subsequent transformer layers, facilitating hierarchical propagation of complementary information while preserving view-specific structure. Experiments on VinDr-Mammo and CMMD datasets demonstrate consistent improvements over linear probing, prompt-only adaptation, and conventional fusion baselines. On the VinDr-Mammo BI-RADS classification task, the framework achieves 50.40% F1-score and 0.8090 AUC, including a 0.10 AUC improvement over a dual-view fusion baseline in the binary setting. Ablation studies further validate the effectiveness of token-based fusion and multi-depth interaction design.