ChatPaper.aiChatPaper

TLive-Omni:面向電子商務直播的全模態理解模型

TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming

August 21, 2026
作者: Yibo Hu, Yu Qian, Mao Gu, Yingfan Tao, Yuhao Chen, Yongdong Luo, Zhuoqun Liu, Meiguang Jin, Junfeng Ma
cs.AI

摘要

電商直播需要對嘈雜且時間上延伸的串流進行全模態理解,其中產品事實分佈於語音、影片幀、產品圖片、疊加文字和使用者查詢中。我們提出 TLive-Omni,一個專為直播電商場景設計的全模態理解模型,能將影像、影片、音訊和文字輸入映射至統一的表示空間。針對長時直播串流分析,我們引入 Per-vGrid,一種帶有時間戳記的 token 組織方式,將每個影片網格與其時間上對應的音訊分組於明確的邊界 token 之內,以促進時間對齊。我們設計了一套三階段監督式訓練流程,從全模態感知到指令跟隨回應,逐步發展直播電商理解能力。接著,我們提出 Faithful-RFT 強化微調階段,在滿足即時性需求的同時,進一步提升答案忠實度與表達品質,直接以任務可驗證的回饋評分最終回應,而非在 rollout 期間針對推理式探索進行最佳化。此外,TLive-Omni 由一個面向場景的原子能力分類法及一個緊湊的資料生產引擎所支援;該引擎可將直播電商的音訊、影像和影片串流轉換為訓練訊號,用於語音辨識、說話者分析、產品視覺定位、文字辨識、時間定位、影片密集描述及全模態問答等任務。為了可擴展的訓練,一個同步的長度分組採樣器在維持各工作節點工作量相當的同時減少填充;而一個輕量級動態採樣策略會重新生成 rollout 群組,其獎勵變異數接近零,以維持 GRPO 的有意義相對優勢。在電商直播基準上的實驗顯示,TLive-Omni 在直播電商領域任務上表現強勁,同時在一般基準上亦具備優異的泛化能力。
English
E-commerce live streaming requires omni-modal understanding of noisy, temporally extended streams, where product facts are distributed across speech, video frames, product images, overlaid text, and user queries. We present TLive-Omni, an omni-modal understanding model tailored to live-commerce scenarios. It maps image, video, audio, and text inputs into a unified representation space. For long-form live streaming analysis, we introduce Per-vGrid, a timestamped token organization that groups each video grid with its temporally corresponding audio within explicit boundary tokens to facilitate temporal alignment. We design a three-stage supervised training recipe that progressively develops live-commerce understanding, from omni-modal perception to instruction-following responses. We then propose Faithful-RFT, a reinforcement fine-tuning stage that further improves answer faithfulness and expression quality while meeting real-time demands, scoring final responses directly with task-verifiable feedback rather than optimizing for reasoning-style exploration during rollout. Moreover, TLive-Omni is supported by a scenario-oriented atomic capability taxonomy and a compact data production engine that converts live-commerce audio, image, and video streams into training signals for speech recognition, speaker analysis, product visual grounding, text recognition, temporal grounding, video dense caption, and omni-modal QA, etc. For scalable training, a synchronized length-grouped sampler reduces padding while preserving comparable workloads across workers, while a lightweight dynamic sampling strategy regenerates rollout groups with near-zero reward variance to maintain meaningful relative advantages for GRPO. Experiments on e-commerce live streaming benchmarks demonstrate strong performance across live-commerce domain tasks, together with excellent generalization on general benchmarks.