乳がん分類のためのトークンベースのデュアルビュー融合と大規模視覚モデルの適応
Token-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer Classification
July 7, 2026
著者: Aysan Ghayouri Pirsoltan, Shima Babakordi, Mohammad Reza Mohammadi
cs.AI
要旨
マンモグラフィーからの正確な乳がん分類には、頭尾方向(CC)像と内外斜位方向(MLO)像という相補的な情報の効果的な統合が必要であり、これにより乳房異常のより完全な特徴付けが可能となる。しかし、既存のマルチビュー学習アプローチは、通常、特徴レベルの集約や単一層のクロスアテンションに依存しており、ビュー固有の表現と共有表現が絡み合い、相互作用が限られたネットワーク深度に制限される可能性がある。これらの限界に対処するため、我々は、凍結されたビジョントランスフォーマーバックボーン内でプロンプトベースの適応とクロスビューフュージョンを統合する、トークン中心のデュアルビュー学習フレームワークを提案する。本フレームワークは、ビュー間の相互作用を構造化されたトークンレベルの通信として再構築し、専用の融合トークンがクロスアテンションを介してCC像とMLO像間の双方向情報交換を明示的に符号化し、直接的な特徴融合に依存するのではなく、クロスビュー依存関係の中間キャリアとして機能する。単一層で融合を適用する従来の方法とは異なり、融合モジュールは複数のトランスフォーマー深度に挿入され、エンコーダ階層全体にわたる漸進的かつ反復的な相互作用を可能にする。融合トークンはトークンシーケンスに再統合され、後続のトランスフォーマー層によって洗練され、ビュー固有の構造を保持しつつ、相補情報の階層的伝播を促進する。VinDr-MammoおよびCMMDデータセットでの実験により、線形プロービング、プロンプトのみの適応、および従来の融合ベースラインに対して一貫した改善が示された。VinDr-MammoのBI-RADS分類タスクにおいて、本フレームワークはF1スコア50.40%、AUC 0.8090を達成し、二値設定ではデュアルビュー融合ベースラインに対してAUCが0.10向上した。アブレーション研究により、トークンベースの融合と多深度相互作用設計の有効性がさらに検証された。
English
Accurate breast cancer classification from mammography requires effective integration of complementary information from craniocaudal (CC) and mediolateral oblique (MLO) views, which provide a more complete characterization of breast abnormalities. However, existing multi-view learning approaches typically rely on feature-level aggregation or single-stage cross-attention, which can entangle view-specific and shared representations and restrict interaction to limited network depths. To address these limitations, we propose a token-centric dual-view learning framework that unifies prompt-based adaptation and cross-view fusion within a frozen vision transformer backbone. The framework reformulates inter-view interaction as structured token-level communication, where dedicated fusion tokens explicitly encode bidirectional information exchange between CC and MLO views via cross-attention, serving as intermediate carriers of cross-view dependencies rather than relying on direct feature fusion. Unlike conventional methods that apply fusion at a single layer, fusion modules are inserted at multiple transformer depths, enabling progressive and repeated interaction across the encoder hierarchy. Fusion tokens are reintegrated into the token sequence and refined by subsequent transformer layers, facilitating hierarchical propagation of complementary information while preserving view-specific structure. Experiments on VinDr-Mammo and CMMD datasets demonstrate consistent improvements over linear probing, prompt-only adaptation, and conventional fusion baselines. On the VinDr-Mammo BI-RADS classification task, the framework achieves 50.40% F1-score and 0.8090 AUC, including a 0.10 AUC improvement over a dual-view fusion baseline in the binary setting. Ablation studies further validate the effectiveness of token-based fusion and multi-depth interaction design.