ChatPaper.aiChatPaper

OpenLongTail:长尾驾驶数据的生成式扩展

OpenLongTail: Generative Scaling of Long-Tail Driving Data

July 10, 2026
作者: Lulin Liu, Nuo Chen, Yan Wang, Bangya Liu, Wenyan Cong, Hezhen Hu, Boris Ivanovic, Hao Wang, Ziyao Zeng, Xinyu Gong, Yang Zhou, Zixiang Xiong, Dilin Wang, Zhangyang Wang, Weisong Shi, Ruohan Zhang, Marco Pavone, Zhiwen Fan
cs.AI

摘要

扩展鲁棒驾驶策略的根本瓶颈在于精心整理的数据集中边缘案例的稀缺性。尽管现实世界持续捕获这些关键事件,但当从异构来源收集时,此类长尾事件仍未得到充分利用。具体而言,多样但宝贵的野外长尾视频缺乏训练策略模型所需的完整视角覆盖,常常缺失多视图位姿,或仅源自单目行车记录仪摄像头。这种模态差距阻碍了将这些普遍存在的观测数据转化为可扩展的长尾泛化训练数据。我们提出OpenLongTail——一个用于在长尾事件下扩展自动驾驶策略的开源生成式数据引擎。为将异构数据源转化为对策略学习有用的视角对齐且时间连贯的多视图资产,我们开发了一种基于位姿推断的外推式视图合成管道,以生成缺失的视图。我们进一步通过将普吕克射线几何注入可扩展生成引擎,增强了新生成视图的跨视角一致性和时间对齐。通过合成异构长尾数据,我们观察到在处理长尾事件时闭环驾驶鲁棒性的显著提升。通过衡量外推式视图合成和位姿指标,我们验证了OpenLongTail在视觉保真度、跨视角一致性和自我轨迹恢复方面的有效性。
English
Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. While the real world continuously captures these critical events, such long-tail events remain underutilized when collected from heterogeneous sources. Specifically, diverse but valuable in-the-wild long-tail videos lack the full view coverage required for training policy models, often missing multi-view poses or originating solely from monocular dash cameras. This modality gap prevents these ubiquitous observations from being converted into scalable training data for long-tail generalization. We introduce OpenLongTail, an open-source generative data engine for scaling autonomous driving policies under long-tail events. To transform heterogeneous data sources into view-aligned and temporally coherent multi-view assets that are useful for policy learning, we develop a pose-informed extrapolative view synthesis pipeline that generates the missing views. We further enhance cross-view consistency and the temporal alignment for the newly generated views by injecting Plücker ray geometry into the scalable generation engine. By synthesizing heterogeneous long-tail data, we observe a significant improvement in closed-loop driving robustness in handling long-tail events. By measuring the extrapolative view synthesis and pose metrics, we validate the effectiveness of OpenLongTail in visual fidelity, cross-view consistency, and ego-trajectory recovery.