ChatPaper.aiChatPaper

GigaChat Audio:時間感知大型音訊語言模型

GigaChat Audio: Time-aware Large Audio Language Model

July 11, 2026
作者: Aleksandr Kutsakov, Mariia Sadovina, Georgii Gospodinov, Alexandr Maximenko, Oleg Kutuzov, Pavel Bogomolov, Fyodor Minkin
cs.AI

摘要

在長時間錄音中的時間定位對於音頻條件式大型語言模型仍是一項挑戰。我們提出了一種具時間感知能力的音頻LLM,可針對長達120分鐘的輸入,以明確時間戳回答問題。我們的方法透過級聯管線的大規模合成監督,將週期性時間標記與連續音頻區塊交錯處理。本模型在短時及長時基準測試中均展現出色的時間定位準確度,並支援以時間錨點為基礎的片段描述與摘要。廣泛的消融實驗探討了時間表示、標記頻率、分詞方式及時長混合設計對準確度與計算成本的影響。我們公開模型權重與資料集,以促進時間感知音頻理解的後續研究,網址:https://huggingface.co/ai-sage/GigaChat3.1-Audio-10B-A1.8B
English
Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input. Our approach interleaves periodic time markers with continuous audio tokens using large-scale synthetic supervision from a cascaded pipeline. Our model achieves strong temporal-grounding accuracy on short and long benchmarks and supports time-anchored fragment descriptions and summaries. Extensive ablations examine how time representation, marker frequency, tokenization, and duration-mixture design affect accuracy and computational cost. We release model weights and datasets to support further research on time-aware audio understanding, available at https://huggingface.co/ai-sage/GigaChat3.1-Audio-10B-A1.8B.