Live-Omni:ECライブ配信のためのオムニモーダル理解モデル
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
August 21, 2026
著者: Yibo Hu, Yu Qian, Mao Gu, Yingfan Tao, Yuhao Chen, Yongdong Luo, Zhuoqun Liu, Meiguang Jin, Junfeng Ma
cs.AI
要旨
Eコマースライブストリーミングでは、ノイズが多く時間的に長いストリームのオムニモーダル理解が必要であり、商品の事実は音声、ビデオフレーム、商品画像、オーバーレイテキスト、ユーザークエリに分散している。我々は、ライブコマースシナリオに特化したオムニモーダル理解モデル「TLive-Omni」を提案する。本モデルは、画像、ビデオ、音声、テキスト入力を統一的表現空間にマッピングする。長時間のライブストリーミング解析のために、我々はPer-vGridを導入する。これは、明示的な境界トークンで囲んで各ビデオグリッドと時間的に対応するオーディオをグループ化し、時間的整合を容易にする、タイムスタンプ付きトークン編成である。我々は、オムニモーダル知覚から指示追従応答まで、ライブコマース理解を段階的に発展させる3段階の教師あり学習レシピを設計する。次に、我々はFaithful-RFTを提案する。これは、リアルタイム要件を満たしつつ回答の忠実性と表現品質をさらに向上させる強化学習ファインチューニング段階であり、ロールアウト中の推論スタイルの探索を最適化するのではなく、タスクで検証可能なフィードバックを用いて最終応答を直接スコアリングする。さらに、TLive-Omniは、シナリオ指向の原子的能力タクソノミーと、ライブコマースの音声・画像・ビデオストリームを、音声認識、話者分析、商品の視覚的グラウンディング、テキスト認識、時間的グラウンディング、ビデオのデンスキャプション、オムニモーダルQAなどの学習信号に変換するコンパクトなデータ生成エンジンによって支えられている。スケーラブルな学習のために、同期された長さグループ化サンプラーはワーカー間で同等のワークロードを維持しつつパディングを削減し、一方で軽量な動的サンプリング戦略は、GRPOの有意義な相対的アドバンテージを維持するために、ほぼゼロの報酬分散でロールアウトグループを再生成する。Eコマースライブストリーミングベンチマークに関する実験は、ライブコマース領域のタスク全般で高い性能を示し、一般的なベンチマークでも優れた汎化性能を示す。
English
E-commerce live streaming requires omni-modal understanding of noisy, temporally extended streams, where product facts are distributed across speech, video frames, product images, overlaid text, and user queries. We present TLive-Omni, an omni-modal understanding model tailored to live-commerce scenarios. It maps image, video, audio, and text inputs into a unified representation space. For long-form live streaming analysis, we introduce Per-vGrid, a timestamped token organization that groups each video grid with its temporally corresponding audio within explicit boundary tokens to facilitate temporal alignment. We design a three-stage supervised training recipe that progressively develops live-commerce understanding, from omni-modal perception to instruction-following responses. We then propose Faithful-RFT, a reinforcement fine-tuning stage that further improves answer faithfulness and expression quality while meeting real-time demands, scoring final responses directly with task-verifiable feedback rather than optimizing for reasoning-style exploration during rollout. Moreover, TLive-Omni is supported by a scenario-oriented atomic capability taxonomy and a compact data production engine that converts live-commerce audio, image, and video streams into training signals for speech recognition, speaker analysis, product visual grounding, text recognition, temporal grounding, video dense caption, and omni-modal QA, etc. For scalable training, a synchronized length-grouped sampler reduces padding while preserving comparable workloads across workers, while a lightweight dynamic sampling strategy regenerates rollout groups with near-zero reward variance to maintain meaningful relative advantages for GRPO. Experiments on e-commerce live streaming benchmarks demonstrate strong performance across live-commerce domain tasks, together with excellent generalization on general benchmarks.