심각한 음향 변이 하에서의 프로토타입 교정 기반 반복적 자기지도 매니폴드 잡음 제거
Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift
August 15, 2026
저자: Ashish Anand Shukla, Rini Smita Thakur, Aryan Das, Vinod K. Kurmi
cs.AI
초록
오디오-텍스트 기반 모델(ATM)은 심한 음향 잡음 하에서 치명적으로 실패하지만, 기존 적응 전략은 신호가 아닌 잡음을 강화하는 경사 기반 테스트-시간 적응(TTA)에 의존하거나, 추론 시점에는 이용할 수 없는 특권적 잡음 주석을 요구하는 프롬프트 튜닝에 의존한다. 우리는 이러한 실패를 PRISM(Prototype-Rectified Iterative Self-supervised Manifold Denoising)으로 해결한다. PRISM은 학습이 필요 없고 소스 데이터도 필요 없는 TTA 프레임워크로, 심한 음향 잡음이 다중 모달 잠재 공간에서 저랭크 아핀 이동을 유도하며 왜곡 에너지의 90% 이상이 상위 60개 주성분에 집중된다는 아핀 잡음 가설에 기반한다. PRISM은 레이블 없는 대상 배치로부터 이 왜곡을 추정하고, 고정된 텍스트 프로토타입을 기하학적 앵커로 사용하여 세 가지 폐쇄형 기하 보정을 통해 왜곡을 역전시킨다. 이 보정들은 아핀 편향 회귀에 의해 단일 정적 투영 행렬로 통합된다. 추론 시 적응은 0.0009ms의 행렬-벡터 곱셈 한 번으로 축소되므로, 경사 기반 TTA보다 훨씬 빠르며 추가 학습이 필요 없다. UrbanSound8K에서 PRISM은 제로샷 기준선보다 12.94퍼센트 포인트 향상되고, 오라클 지원 TTA 기준선이 사용하는 특권적 증강 잡음 프롬프트를 전혀 관찰하지 못했음에도 해당 기준선을 9.41퍼센트 포인트 능가한다. 또한 우리는 부분공간 축소가 광대역 클래스에 대해 갖는 원리적 실패 모드인 폴리포닉 트랩(Polyphonic Trap)을 식별하고, 신뢰 인식 회귀(CAR)를 통해 이를 해결하여 가장 큰 피해를 입은 클래스에서 최대 8.16퍼센트 포인트를 회복한다.
English
Audio-Text Foundation Models (ATMs) fail catastrophically under severe acoustic noise, yet existing adaptation strategies either rely on gradient-based Test-Time Adaptation (TTA), which reinforces noise rather than signal, or on prompt tuning that requires privileged noise annotations unavailable at inference. We address these failures with PRISM (Prototype-Rectified Iterative Self-supervised Manifold Denoising), a training-free, source-free TTA framework grounded in the Affine Noise Hypothesis: severe acoustic noise induces a low-rank affine shift in the multimodal latent space, with more than 90% of distortion energy confined to the leading 60 principal components. PRISM estimates and reverses this distortion from an unlabeled target batch using frozen text prototypes as geometric anchors via three closed-form geometric corrections compiled into a single static projection matrix by Affine Bias Regression. At inference, adaptation reduces to one matrix-vector multiplication in 0.0009 ms, making it substantially faster than gradient-based TTA while requiring no additional training. On UrbanSound8K, PRISM improves over the zero-shot baseline by 12.94 percentage points and surpasses an oracle-assisted TTA baseline by 9.41 percentage points, despite never observing its privileged augmented noise prompts. We further identify the Polyphonic Trap, a principled failure mode of subspace deflation for broadband classes, and resolve it via Confidence-Aware Regression (CAR), recovering up to 8.16 percentage points for the worst-affected class.