基於令牌的雙視圖融合與大型視覺模型適應於乳癌分類
Token-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer Classification
July 7, 2026
作者: Aysan Ghayouri Pirsoltan, Shima Babakordi, Mohammad Reza Mohammadi
cs.AI
摘要
精確的乳房X光攝影乳癌分類需要有效整合頭尾向(CC)與內外斜位(MLO)視角所提供的互補資訊,從而更完整地刻劃乳房異常特徵。然而,現有多視角學習方法通常依賴於特徵層級的聚合或單階段交叉注意力,這可能導致視角特定表徵與共享表徵相互糾纏,並限制交互作用僅發生於有限的網路深度。為解決此問題,我們提出一個以令牌為中心的雙視角學習框架,該框架在凍結的視覺Transformer骨幹網路中統一了基於提示的適應性調整與跨視角融合。本框架將視角間交互重新建構為結構化的令牌層級通訊,其中專用融合令牌經由交叉注意力機制明確編碼CC與MLO視角之間的雙向資訊交換,作為跨視角依賴關係的中間載體,而非依賴直接的特徵融合。與傳統方法僅在單一層級進行融合不同,融合模組被插入於多個Transformer深度,從而實現編碼器階層中的漸進式與重複性交互。融合令牌被重新整合至令牌序列中,並經由後續Transformer層進行細化,在保留視角特定結構的同時促進互補資訊的階層式傳播。在VinDr-Mammo與CMMD資料集上的實驗結果顯示,本框架相較於線性探測、僅提示適應及傳統融合基線方法均有一致性提升。在VinDr-Mammo的BI-RADS分類任務中,本框架達成了50.40%的F1分數與0.8090的AUC,其中在二元設定下,相較於雙視角融合基線方法提升了0.10的AUC。消融研究進一步驗證了基於令牌的融合與多深度交互設計的有效性。
English
Accurate breast cancer classification from mammography requires effective integration of complementary information from craniocaudal (CC) and mediolateral oblique (MLO) views, which provide a more complete characterization of breast abnormalities. However, existing multi-view learning approaches typically rely on feature-level aggregation or single-stage cross-attention, which can entangle view-specific and shared representations and restrict interaction to limited network depths. To address these limitations, we propose a token-centric dual-view learning framework that unifies prompt-based adaptation and cross-view fusion within a frozen vision transformer backbone. The framework reformulates inter-view interaction as structured token-level communication, where dedicated fusion tokens explicitly encode bidirectional information exchange between CC and MLO views via cross-attention, serving as intermediate carriers of cross-view dependencies rather than relying on direct feature fusion. Unlike conventional methods that apply fusion at a single layer, fusion modules are inserted at multiple transformer depths, enabling progressive and repeated interaction across the encoder hierarchy. Fusion tokens are reintegrated into the token sequence and refined by subsequent transformer layers, facilitating hierarchical propagation of complementary information while preserving view-specific structure. Experiments on VinDr-Mammo and CMMD datasets demonstrate consistent improvements over linear probing, prompt-only adaptation, and conventional fusion baselines. On the VinDr-Mammo BI-RADS classification task, the framework achieves 50.40% F1-score and 0.8090 AUC, including a 0.10 AUC improvement over a dual-view fusion baseline in the binary setting. Ablation studies further validate the effectiveness of token-based fusion and multi-depth interaction design.