ChatPaper.aiChatPaper

OpenLongTail:長尾駕駛數據的生成式擴展

OpenLongTail: Generative Scaling of Long-Tail Driving Data

July 10, 2026
作者: Lulin Liu, Nuo Chen, Yan Wang, Bangya Liu, Wenyan Cong, Hezhen Hu, Boris Ivanovic, Hao Wang, Ziyao Zeng, Xinyu Gong, Yang Zhou, Zixiang Xiong, Dilin Wang, Zhangyang Wang, Weisong Shi, Ruohan Zhang, Marco Pavone, Zhiwen Fan
cs.AI

摘要

強化穩健駕駛策略的發展,其根本瓶頸在於策展資料集中邊緣案例的稀缺性。雖然真實世界持續捕捉這些關鍵事件,但此類長尾事件在從異質來源收集時仍未被充分利用。具體而言,多樣但具價值的野外長尾影片缺乏訓練策略模型所需的完整視角覆蓋,經常遺失多視角姿態,或僅來自單目行車記錄器。這種模態差距阻礙了將這些普遍存在的觀測資料轉化為可擴展的訓練資料,以因應長尾事件的一般化。我們提出 OpenLongTail,一個開源生成式資料引擎,用於擴展自動駕駛策略在長尾事件下的應用。為將異質資料來源轉換為視角對齊且時序連貫的多視角資產(對策略學習有用),我們開發了一套基於姿態推斷的外推式視角合成流程,以生成缺失的視角。我們進一步透過將 Plücker 射線幾何注入可擴展的生成引擎,增強了新生成視角的跨視角一致性與時序對齊。透過合成異質長尾資料,我們觀察到在處理長尾事件時,閉環駕駛的穩健性有顯著提升。藉由評估外推式視角合成與姿態指標,我們驗證了 OpenLongTail 在視覺保真度、跨視角一致性及自我軌跡恢復方面的有效性。
English
Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. While the real world continuously captures these critical events, such long-tail events remain underutilized when collected from heterogeneous sources. Specifically, diverse but valuable in-the-wild long-tail videos lack the full view coverage required for training policy models, often missing multi-view poses or originating solely from monocular dash cameras. This modality gap prevents these ubiquitous observations from being converted into scalable training data for long-tail generalization. We introduce OpenLongTail, an open-source generative data engine for scaling autonomous driving policies under long-tail events. To transform heterogeneous data sources into view-aligned and temporally coherent multi-view assets that are useful for policy learning, we develop a pose-informed extrapolative view synthesis pipeline that generates the missing views. We further enhance cross-view consistency and the temporal alignment for the newly generated views by injecting Plücker ray geometry into the scalable generation engine. By synthesizing heterogeneous long-tail data, we observe a significant improvement in closed-loop driving robustness in handling long-tail events. By measuring the extrapolative view synthesis and pose metrics, we validate the effectiveness of OpenLongTail in visual fidelity, cross-view consistency, and ego-trajectory recovery.