ChatPaper.aiChatPaper

OpenLongTail: 롱테일 주행 데이터의 생성적 확장

OpenLongTail: Generative Scaling of Long-Tail Driving Data

July 10, 2026
저자: Lulin Liu, Nuo Chen, Yan Wang, Bangya Liu, Wenyan Cong, Hezhen Hu, Boris Ivanovic, Hao Wang, Ziyao Zeng, Xinyu Gong, Yang Zhou, Zixiang Xiong, Dilin Wang, Zhangyang Wang, Weisong Shi, Ruohan Zhang, Marco Pavone, Zhiwen Fan
cs.AI

초록

강건한 주행 정책을 확장하는 데 있어 근본적인 병목 현상은 선별된 데이터셋 내 에지 사례의 부족에서 비롯됩니다. 실제 세계는 이러한 중요한 이벤트를 지속적으로 포착하지만, 롱테일 사건들은 이질적 소스로부터 수집될 때 충분히 활용되지 못합니다. 구체적으로, 다양하지만 가치 있는 현장 롱테일 영상은 정책 모델 훈련에 필요한 전체 시점 범위가 부족하여, 다중 시점 포즈가 누락되거나 단일 시점 대시카메라에서만 획득되는 경우가 많습니다. 이러한 모달리티 차이로 인해 보편적으로 존재하는 관측 데이터가 롱테일 일반화를 위한 확장 가능한 훈련 데이터로 전환되지 못합니다. 본 논문에서는 롱테일 이벤트에서 자율주행 정책을 확장하기 위한 오픈소스 생성형 데이터 엔진인 OpenLongTail을 소개합니다. 이질적 데이터 소스를 정책 학습에 유용한 시점 정렬 및 시간적 일관성을 갖춘 다중 시점 자산으로 변환하기 위해, 누락된 시점을 생성하는 포즈 기반 외삽적 시점 합성 파이프라인을 개발합니다. 또한, 확장 가능한 생성 엔진에 플뤼커 광선 기하학을 주입하여 새로 생성된 시점의 교차 시점 일관성 및 시간적 정렬을 강화합니다. 이질적 롱테일 데이터를 합성함으로써 롱테일 이벤트 처리에 있어 폐루프 주행 강건성이 크게 향상됨을 관찰했습니다. 외삽적 시점 합성 및 포즈 메트릭을 측정하여 시각적 충실도, 교차 시점 일관성, 자차 궤적 복원 측면에서 OpenLongTail의 효과를 검증합니다.
English
Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. While the real world continuously captures these critical events, such long-tail events remain underutilized when collected from heterogeneous sources. Specifically, diverse but valuable in-the-wild long-tail videos lack the full view coverage required for training policy models, often missing multi-view poses or originating solely from monocular dash cameras. This modality gap prevents these ubiquitous observations from being converted into scalable training data for long-tail generalization. We introduce OpenLongTail, an open-source generative data engine for scaling autonomous driving policies under long-tail events. To transform heterogeneous data sources into view-aligned and temporally coherent multi-view assets that are useful for policy learning, we develop a pose-informed extrapolative view synthesis pipeline that generates the missing views. We further enhance cross-view consistency and the temporal alignment for the newly generated views by injecting Plücker ray geometry into the scalable generation engine. By synthesizing heterogeneous long-tail data, we observe a significant improvement in closed-loop driving robustness in handling long-tail events. By measuring the extrapolative view synthesis and pose metrics, we validate the effectiveness of OpenLongTail in visual fidelity, cross-view consistency, and ego-trajectory recovery.