TLive-Omni: 전자상거래 라이브 스트리밍을 위한 옴니모달 이해 모델
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
August 21, 2026
저자: Yibo Hu, Yu Qian, Mao Gu, Yingfan Tao, Yuhao Chen, Yongdong Luo, Zhuoqun Liu, Meiguang Jin, Junfeng Ma
cs.AI
초록
전자상거래 라이브 스트리밍은 잡음이 섞여 있고 시간적으로 길게 확장된 스트림에 대한 옴니모달 이해를 요구하며, 제품 정보는 음성, 비디오 프레임, 제품 이미지, 오버레이 텍스트, 사용자 질문에 분산되어 있다. 본 논문에서는 라이브커머스 시나리오에 특화된 옴니모달 이해 모델인 TLive-Omni를 제시한다. 이 모델은 이미지, 비디오, 오디오, 텍스트 입력을 통합 표현 공간으로 매핑한다. 장시간 라이브 스트리밍 분석을 위해, 명시적 경계 토큰 내에서 각 비디오 그리드를 시간적으로 대응하는 오디오와 함께 그룹화하는 타임스탬프 토큰 구성 방식인 Per-vGrid를 도입하여 시간적 정렬을 용이하게 한다. 우리는 옴니모달 지각부터 지시 수행 응답까지 라이브커머스 이해를 점진적으로 발전시키는 3단계 지도 학습 방식을 설계한다. 이어서 롤아웃 중 추론 스타일 탐색을 최적화하는 대신 작업 검증 가능한 피드백으로 최종 응답을 직접 평가하여, 실시간 요구를 충족하면서 답변의 충실성과 표현 품질을 추가로 향상시키는 강화 미세 조정 단계인 Faithful-RFT를 제안한다. 또한 TLive-Omni는 시나리오 중심 원자적 능력 분류체계와, 라이브커머스 오디오, 이미지, 비디오 스트림을 음성 인식, 화자 분석, 제품 시각적 근거 찾기, 텍스트 인식, 시간적 근거 찾기, 비디오 조밀 캡셔닝, 옴니모달 QA 등의 학습 신호로 변환하는 간결한 데이터 생성 엔진에 의해 지원된다. 확장 가능한 학습을 위해, 동기화된 길이 그룹 샘플러는 작업자 간 유사한 작업 부하를 유지하면서 패딩을 줄이며, 경량 동적 샘플링 전략은 GRPO를 위한 의미 있는 상대적 이점을 유지하도록 보상 분산이 거의 0에 가까운 롤아웃 그룹을 재생성한다. 전자상거래 라이브 스트리밍 벤치마크 실험은 라이브커머스 도메인 작업 전반에 걸쳐 강력한 성능을 입증하며, 일반 벤치마크에 대한 우수한 일반화 성능도 함께 보여준다.
English
E-commerce live streaming requires omni-modal understanding of noisy, temporally extended streams, where product facts are distributed across speech, video frames, product images, overlaid text, and user queries. We present TLive-Omni, an omni-modal understanding model tailored to live-commerce scenarios. It maps image, video, audio, and text inputs into a unified representation space. For long-form live streaming analysis, we introduce Per-vGrid, a timestamped token organization that groups each video grid with its temporally corresponding audio within explicit boundary tokens to facilitate temporal alignment. We design a three-stage supervised training recipe that progressively develops live-commerce understanding, from omni-modal perception to instruction-following responses. We then propose Faithful-RFT, a reinforcement fine-tuning stage that further improves answer faithfulness and expression quality while meeting real-time demands, scoring final responses directly with task-verifiable feedback rather than optimizing for reasoning-style exploration during rollout. Moreover, TLive-Omni is supported by a scenario-oriented atomic capability taxonomy and a compact data production engine that converts live-commerce audio, image, and video streams into training signals for speech recognition, speaker analysis, product visual grounding, text recognition, temporal grounding, video dense caption, and omni-modal QA, etc. For scalable training, a synchronized length-grouped sampler reduces padding while preserving comparable workloads across workers, while a lightweight dynamic sampling strategy regenerates rollout groups with near-zero reward variance to maintain meaningful relative advantages for GRPO. Experiments on e-commerce live streaming benchmarks demonstrate strong performance across live-commerce domain tasks, together with excellent generalization on general benchmarks.