ChatPaper.aiChatPaper

嚴重聲學偏移下基於原型矯正之迭代自監督流形去噪

Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift

August 15, 2026
作者: Ashish Anand Shukla, Rini Smita Thakur, Aryan Das, Vinod K. Kurmi
cs.AI

摘要

音訊-文本基礎模型(ATMs)在嚴重的聲學噪聲下會災難性地失效。然而,現有的適應策略要麼依賴於基於梯度的測試時適應(TTA),此類方法強化的是噪聲而非訊號;要麼依賴於提示調整,但後者需要推論時無法取得的特權噪聲標註。我們透過PRISM(原型校正迭代自監督流形去噪)來應對這些失效。PRISM是一個免訓練、無需來源數據的TTA框架,其基礎是仿射噪聲假說:嚴重的聲學噪聲會在多模態潛在空間中引起低秩仿射偏移,且超過90%的失真能量集中於前60個主成分。PRISM以凍結的文字原型作為幾何錨點,透過仿射偏差回歸將三個閉式幾何校正整合為單一靜態投影矩陣,從而從未標記的目標批次中估計並逆轉此失真。在推論時,適應過程簡化為一次矩陣-向量乘法,耗時0.0009毫秒,因此遠快於基於梯度的TTA,且無需額外訓練。在UrbanSound8K上,PRISM比零樣本基線提升12.94個百分點,並超越具神諭輔助的TTA基線9.41個百分點,儘管它從未觀察過該基線的特權增強噪聲提示。我們進一步識別了複音陷阱,這是子空間縮減對寬頻類別的一種具原理性的失效模式,並透過置信度感知回歸(CAR)加以解決,為受影響最嚴重的類別恢復多達8.16個百分點。
English
Audio-Text Foundation Models (ATMs) fail catastrophically under severe acoustic noise, yet existing adaptation strategies either rely on gradient-based Test-Time Adaptation (TTA), which reinforces noise rather than signal, or on prompt tuning that requires privileged noise annotations unavailable at inference. We address these failures with PRISM (Prototype-Rectified Iterative Self-supervised Manifold Denoising), a training-free, source-free TTA framework grounded in the Affine Noise Hypothesis: severe acoustic noise induces a low-rank affine shift in the multimodal latent space, with more than 90% of distortion energy confined to the leading 60 principal components. PRISM estimates and reverses this distortion from an unlabeled target batch using frozen text prototypes as geometric anchors via three closed-form geometric corrections compiled into a single static projection matrix by Affine Bias Regression. At inference, adaptation reduces to one matrix-vector multiplication in 0.0009 ms, making it substantially faster than gradient-based TTA while requiring no additional training. On UrbanSound8K, PRISM improves over the zero-shot baseline by 12.94 percentage points and surpasses an oracle-assisted TTA baseline by 9.41 percentage points, despite never observing its privileged augmented noise prompts. We further identify the Polyphonic Trap, a principled failure mode of subspace deflation for broadband classes, and resolve it via Confidence-Aware Regression (CAR), recovering up to 8.16 percentage points for the worst-affected class.