ChatPaper.aiChatPaper

iFAN: 추론을 고려한 평면 마스크 트랜스포머 학습

iFAN: Inference-Aware Learning for Plain Mask Transformers

August 7, 2026
저자: Fang Li, Yu He, Haoyang Tong, Lichen Ma, Jingling Fu, Wenxiao Fan, Tongxuan Liu, Luohang Liu, Ke Zhang, Junshi Huang
cs.AI

초록

쿼리 기반 마스크 트랜스포머는 최종 계층의 쿼리 예측 간 픽셀 단위 경쟁을 통해 분할 출력을 구성하지만, 이러한 추론 과정은 훈련 중에 명시적으로 최적화되지 않는다. 우리는 두 가지 주요 불일치를 식별한다: 가장 높은 확률-마스크 점수를 가진 쿼리가 반드시 가장 정확한 마스크를 생성하는 것은 아니며, 최종 계층 디코딩은 중간 계층의 우수한 예측을 버릴 수 있다. 이러한 문제를 해결하기 위해 우리는 일반 마스크 트랜스포머를 위한 일반적인 훈련 프레임워크인 추론 인지 학습(iFAN)을 제안한다. iFAN은 쿼리 경쟁을 예측된 마스크 품질과 정렬하고 높은 신뢰도지만 부정확한 경쟁자를 억제하는 조정된 확률-마스크 순위(APMR)를 도입한다. 또한 교차 계층 자기 증류(CLSD)를 통해 더 강력한 중간 예측을 최종 계층으로 전달한다. 순위 및 증류 목적은 훈련 전용이며, 추론 시에는 효율적인 최종 계층 디코딩을 유지한다. COCO, ADE20K 및 Cityscapes에서의 실험은 파놉틱, 인스턴스 및 시맨틱 분할과 다양한 아키텍처, 백본 규모, 입력 해상도에 걸쳐 일관된 성능 향상을 보여준다. 전반적으로 iFAN은 무시할 수 있는 수준의 추가 파라미터, FLOPs 및 추론 지연 시간만으로 평균 1.20 PQ, 1.30 AP, 0.63 mIoU의 성능 향상을 달성한다.
English
Query-based mask transformers assemble segmentation outputs through pixel-wise competition among query predictions of the final layer, yet this inference process is not explicitly optimized during training. We identify two key mismatches: the query with the highest probability-mask score does not necessarily produce the most accurate mask, and final-layer decoding may discard superior predictions from intermediate layers. To address these issues, we propose Inference-Aware Learning (iFAN), a general training framework for plain mask transformers. iFAN introduces Adjusted Probability-Mask Ranking (APMR), which aligns query competition with predicted mask quality and suppresses high-confidence but inaccurate competitors. We further employ Cross-Layer Self-Distillation (CLSD) to transfer stronger intermediate predictions to the final layer. The ranking and distillation objectives are training-only, while inference retains efficient final-layer decoding. Experiments on COCO, ADE20K, and Cityscapes demonstrate consistent improvements across panoptic, instance, and semantic segmentation, as well as across different architectures, backbone scales, and input resolutions. Overall, iFAN improves performance by an average of 1.20 PQ, 1.30 AP, and 0.63 mIoU, with negligible additional parameters, FLOPs and inference latency.