기가챗 오디오: 시간 인식 대규모 오디오 언어 모델
GigaChat Audio: Time-aware Large Audio Language Model
July 11, 2026
저자: Aleksandr Kutsakov, Mariia Sadovina, Georgii Gospodinov, Alexandr Maximenko, Oleg Kutuzov, Pavel Bogomolov, Fyodor Minkin
cs.AI
초록
긴 녹음에서의 시간적 접지는 오디오 조건 기반 LLM에게 여전히 어려운 과제입니다. 본 논문에서는 최대 120분 입력에 대해 명시적 타임스탬프가 포함된 질문에 답변하는 시간 인식 오디오 LLM을 제시합니다. 우리의 접근 방식은 캐스케이드 파이프라인의 대규모 합성 지도 학습을 활용하여 주기적 시간 마커를 연속 오디오 토큰과 인터리빙합니다. 우리 모델은 짧은 벤치마크와 긴 벤치마크 모두에서 강력한 시간적 접지 정확도를 달성하며, 시간 기준 조각 설명 및 요약을 지원합니다. 광범위한 절제 실험을 통해 시간 표현, 마커 빈도, 토큰화, 지속 시간 혼합 설계가 정확도와 계산 비용에 미치는 영향을 분석합니다. 시간 인식 오디오 이해에 대한 추가 연구를 지원하기 위해 모델 가중치와 데이터셋을 공개하며, https://huggingface.co/ai-sage/GigaChat3.1-Audio-10B-A1.8B에서 확인할 수 있습니다.
English
Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input. Our approach interleaves periodic time markers with continuous audio tokens using large-scale synthetic supervision from a cascaded pipeline. Our model achieves strong temporal-grounding accuracy on short and long benchmarks and supports time-anchored fragment descriptions and summaries. Extensive ablations examine how time representation, marker frequency, tokenization, and duration-mixture design affect accuracy and computational cost. We release model weights and datasets to support further research on time-aware audio understanding, available at https://huggingface.co/ai-sage/GigaChat3.1-Audio-10B-A1.8B.