ChatPaper.aiChatPaper

GigaChat Audio: 時間認識型大規模音声言語モデル

GigaChat Audio: Time-aware Large Audio Language Model

July 11, 2026
著者: Aleksandr Kutsakov, Mariia Sadovina, Georgii Gospodinov, Alexandr Maximenko, Oleg Kutuzov, Pavel Bogomolov, Fyodor Minkin
cs.AI

要旨

長時間録音における時間的接地は、音声条件付き大規模言語モデル(LLM)にとって依然として課題です。我々は、最大120分の入力に対して明示的なタイムスタンプ付きで質問に答える時間認識型音声LLMを提案します。本手法では、カスケードパイプラインによる大規模な合成教師データを用いて、定期的な時間マーカーを連続音声トークンにインターリーブします。提案モデルは、短時間および長時間のベンチマークにおいて高い時間的接地精度を達成し、時間ベースの断片記述や要約をサポートします。広範なアブレーション研究により、時間表現、マーカー頻度、トークン化、および持続時間混合設計が精度と計算コストに与える影響を調査します。我々は、時間認識型音声理解のさらなる研究を支援するため、モデルの重みとデータセットを公開します。入手先:https://huggingface.co/ai-sage/GigaChat3.1-Audio-10B-A1.8B
English
Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input. Our approach interleaves periodic time markers with continuous audio tokens using large-scale synthetic supervision from a cascaded pipeline. Our model achieves strong temporal-grounding accuracy on short and long benchmarks and supports time-anchored fragment descriptions and summaries. Extensive ablations examine how time representation, marker frequency, tokenization, and duration-mixture design affect accuracy and computational cost. We release model weights and datasets to support further research on time-aware audio understanding, available at https://huggingface.co/ai-sage/GigaChat3.1-Audio-10B-A1.8B.