ChatPaper.aiChatPaper

严重声学偏移下的原型校正迭代自监督流形去噪

Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift

August 15, 2026
作者: Ashish Anand Shukla, Rini Smita Thakur, Aryan Das, Vinod K. Kurmi
cs.AI

摘要

音频-文本基础模型(ATM)在严重声学噪声下会灾难性失效,然而现有自适应策略要么依赖基于梯度的测试时自适应(TTA),这往往会强化噪声而非信号;要么依赖提示调优,但后者需要推理时无法获取的特权噪声标注。针对这些问题,我们提出PRISM(原型校正迭代自监督流形去噪)——一种无需训练、无需源数据的TTA框架,其理论基础是仿射噪声假设:严重声学噪声在多模态潜在空间中诱发低秩仿射偏移,且超过90%的失真能量集中于前60个主成分。PRISM利用冻结的文本原型作为几何锚点,通过仿射偏差回归将三种闭式几何校正合并为单个静态投影矩阵,从而从未标注目标批次中估计并逆转该失真。推理时,自适应仅需一次矩阵-向量乘法(0.0009毫秒),因此远快于基于梯度的TTA,且无需任何额外训练。在UrbanSound8K上,PRISM相比零样本基线提升12.94个百分点,并且尽管从未观察过预言机辅助TTA基线所拥有的特权增强噪声提示,仍以9.41个百分点的优势超越该基线。我们还识别出“复音陷阱”——一种针对宽带类别的子空间紧缩原理性失效模式,并通过置信度感知回归(CAR)加以解决,为受影响最严重的类别恢复了高达8.16个百分点。
English
Audio-Text Foundation Models (ATMs) fail catastrophically under severe acoustic noise, yet existing adaptation strategies either rely on gradient-based Test-Time Adaptation (TTA), which reinforces noise rather than signal, or on prompt tuning that requires privileged noise annotations unavailable at inference. We address these failures with PRISM (Prototype-Rectified Iterative Self-supervised Manifold Denoising), a training-free, source-free TTA framework grounded in the Affine Noise Hypothesis: severe acoustic noise induces a low-rank affine shift in the multimodal latent space, with more than 90% of distortion energy confined to the leading 60 principal components. PRISM estimates and reverses this distortion from an unlabeled target batch using frozen text prototypes as geometric anchors via three closed-form geometric corrections compiled into a single static projection matrix by Affine Bias Regression. At inference, adaptation reduces to one matrix-vector multiplication in 0.0009 ms, making it substantially faster than gradient-based TTA while requiring no additional training. On UrbanSound8K, PRISM improves over the zero-shot baseline by 12.94 percentage points and surpasses an oracle-assisted TTA baseline by 9.41 percentage points, despite never observing its privileged augmented noise prompts. We further identify the Polyphonic Trap, a principled failure mode of subspace deflation for broadband classes, and resolve it via Confidence-Aware Regression (CAR), recovering up to 8.16 percentage points for the worst-affected class.