어텐션이 시각을 잃을 때: ALiBi 위치 인코딩의 수치적 실패
When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings
August 4, 2026
저자: Christopher Schröder, Lukas Gienapp, Ferdinand Schlatt, Martin Potthast, Gerhard Heyer
cs.AI
초록
우리는 ALiBi 위치 인코딩에서 이전에 간과되었던 실패 모드를 식별한다: 선형 바이어스 스케일링이 부동소수점 정밀도에서 언더플로되어 어텐션 가중치의 상당 부분을 0으로 만들고, 이로 인해 영향을 받은 어텐션 헤드가 부분적으로 실명 상태가 된다. 우리는 이 실패 모드를 분석하고 그 영향을 특성화하며 네 가지 완화 전략을 검토한다. 또한 ALiBi 기반의 최신 사전 학습 모델에서 이 실패 모드가 발생함을 입증한다. 1억 4800만 파라미터 디코더 모델을 사용한 종합적인 사전 학습 실험은 그 효과를 맥락 밖 성능 저하와 분리하는 데 도움을 준다. 우리는 ALiBi의 실패 모드가 표준 디코더 벤치마크에는 미미한 영향만 미치면서도 토큰 검색을 상당히 저하시킬 수 있음을 발견한다. 우리는 네 가지 학습 시점 완화 전략을 제안하고 이를 개별적으로 그리고 조합하여 평가하며, 로그 스케일 거리가 패스키 검색에서 가장 일관된 개선을 가져온다는 것을 발견한다. 이러한 문제에도 불구하고, 기본 ALiBi 기울기는 특히 건초 더미 속 바늘 찾기 검색에서 놀라울 정도로 강력한 기준선으로 남아 있다. 이러한 발견을 바탕으로 우리는 ALiBi로 모델을 학습하는 방법에 대한 구체적인 권고사항을 제공한다.
English
We identify a previously overlooked failure mode of ALiBi positional encoding: its linear bias scaling underflows floating-point precision, which zeroes out a large fraction of attention weights and renders the affected attention heads partially blind. We analyze this failure mode, characterize its impact, and examine four mitigation strategies. We further demonstrate its occurrence in state-of-the-art pretrained models based on ALiBi. Comprehensive pretraining experiments with 148M-parameter decoder models help us to disentangle its effects from out-of-context degradation. We find that ALiBi's failure mode can substantially impair token retrieval while having only a minor effect on standard decoder benchmarks. We propose four training-time mitigation strategies and evaluate them individually and in combinations, finding that log-scaled distances yield the most consistent improvements in passkey retrieval. Despite this problem, default ALiBi slopes remain a surprisingly strong baseline, particularly for needle-in-a-haystack retrieval. Based on these findings we provide concrete recommendations on how to train models with ALiBi.