深刻な音響シフト下におけるプロトタイプ補正反復的自己教師あり多様体デノイジング
Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift
August 15, 2026
著者: Ashish Anand Shukla, Rini Smita Thakur, Aryan Das, Vinod K. Kurmi
cs.AI
要旨
音声テキスト基盤モデル(ATMs)は、深刻な音響ノイズの下で壊滅的に失敗する。しかし既存の適応戦略は、勾配ベースのテスト時適応(TTA)に依存するか(これは信号ではなくノイズを強化してしまう)、あるいは推論時には利用できない特権的なノイズ注釈を必要とするプロンプトチューニングに依存する。我々はこれらの失敗に対処するため、PRISM(Prototype-Rectified Iterative Self-supervised Manifold Denoising)を提案する。これは、アフィンノイズ仮説に基づく、訓練不要かつソースフリーのTTAフレームワークである。アフィンノイズ仮説とは、深刻な音響ノイズがマルチモーダル潜在空間に低ランクのアフィンシフトを引き起こし、歪みエネルギーの90%以上が先頭60主成分に集中するというものである。PRISMは、凍結されたテキストプロトタイプを幾何学的アンカーとして用い、ラベルなしのターゲットバッチからこの歪みを推定・反転させる。その際、3つの閉形式幾何補正をAffine Bias Regressionによって単一の静的射影行列に統合する。推論時には、適応は1回の行列ベクトル乗算(0.0009 ms)に帰着し、勾配ベースのTTAよりも大幅に高速であり、追加の訓練を一切必要としない。UrbanSound8Kにおいて、PRISMはゼロショットベースラインを12.94パーセンテージポイント上回り、特権的な拡張ノイズプロンプトを一切観測しないにもかかわらず、オラクル支援TTAベースラインを9.41パーセンテージポイント上回る。さらに我々は、広帯域クラスに対する部分空間デフレーションの原理的な失敗モードであるPolyphonic Trap(ポリフォニックトラップ)を特定し、Confidence-Aware Regression(CAR)によってこれを解決し、最も影響を受けるクラスで最大8.16パーセンテージポイントを回復する。
English
Audio-Text Foundation Models (ATMs) fail catastrophically under severe acoustic noise, yet existing adaptation strategies either rely on gradient-based Test-Time Adaptation (TTA), which reinforces noise rather than signal, or on prompt tuning that requires privileged noise annotations unavailable at inference. We address these failures with PRISM (Prototype-Rectified Iterative Self-supervised Manifold Denoising), a training-free, source-free TTA framework grounded in the Affine Noise Hypothesis: severe acoustic noise induces a low-rank affine shift in the multimodal latent space, with more than 90% of distortion energy confined to the leading 60 principal components. PRISM estimates and reverses this distortion from an unlabeled target batch using frozen text prototypes as geometric anchors via three closed-form geometric corrections compiled into a single static projection matrix by Affine Bias Regression. At inference, adaptation reduces to one matrix-vector multiplication in 0.0009 ms, making it substantially faster than gradient-based TTA while requiring no additional training. On UrbanSound8K, PRISM improves over the zero-shot baseline by 12.94 percentage points and surpasses an oracle-assisted TTA baseline by 9.41 percentage points, despite never observing its privileged augmented noise prompts. We further identify the Polyphonic Trap, a principled failure mode of subspace deflation for broadband classes, and resolve it via Confidence-Aware Regression (CAR), recovering up to 8.16 percentage points for the worst-affected class.