ChatPaper.aiChatPaper

OpenLongTail: 長尾運転データの生成的スケーリング

OpenLongTail: Generative Scaling of Long-Tail Driving Data

July 10, 2026
著者: Lulin Liu, Nuo Chen, Yan Wang, Bangya Liu, Wenyan Cong, Hezhen Hu, Boris Ivanovic, Hao Wang, Ziyao Zeng, Xinyu Gong, Yang Zhou, Zixiang Xiong, Dilin Wang, Zhangyang Wang, Weisong Shi, Ruohan Zhang, Marco Pavone, Zhiwen Fan
cs.AI

要旨

スケーラブルなロバスト運転ポリシーは、キュレーションデータセットにおけるエッジケースの不足に根本的にボトルネックを抱えている。実世界ではこれらの重要なイベントが継続的に捉えられているものの、異種ソースから収集された場合、こうしたロングテールイベントは十分に活用されていない。特に、多様で貴重な実環境のロングテール映像は、ポリシーモデルの訓練に必要な全視点カバレッジを欠いており、マルチビューのポーズ情報が不足していたり、単眼ダッシュカメラのみから取得されていることが多い。このモダリティギャップにより、これらの遍在する観測データがロングテール汎化のためのスケーラブルな訓練データに変換されることが妨げられている。本稿では、ロングテールイベント下での自動運転ポリシーをスケーリングするためのオープンソース生成型データエンジンであるOpenLongTailを紹介する。異種データソースを、ポリシー学習に有用な視点整合性・時間的一貫性を備えたマルチビューアセットに変換するため、我々はポーズ情報に基づく外挿的視点合成パイプラインを開発し、欠損した視点を生成する。さらに、スケーラブルな生成エンジンにプリュッカー線幾何を注入することで、新たに生成された視点間のクロスビュー一貫性と時間的アライメントを強化する。異種ロングテールデータを合成することにより、ロングテールイベント処理におけるクローズドループ運転のロバスト性が大幅に向上することを確認した。外挿的視点合成とポーズ指標の評価を通じて、視覚的忠実度、クロスビュー一貫性、および自我軌跡復元におけるOpenLongTailの有効性を検証した。
English
Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. While the real world continuously captures these critical events, such long-tail events remain underutilized when collected from heterogeneous sources. Specifically, diverse but valuable in-the-wild long-tail videos lack the full view coverage required for training policy models, often missing multi-view poses or originating solely from monocular dash cameras. This modality gap prevents these ubiquitous observations from being converted into scalable training data for long-tail generalization. We introduce OpenLongTail, an open-source generative data engine for scaling autonomous driving policies under long-tail events. To transform heterogeneous data sources into view-aligned and temporally coherent multi-view assets that are useful for policy learning, we develop a pose-informed extrapolative view synthesis pipeline that generates the missing views. We further enhance cross-view consistency and the temporal alignment for the newly generated views by injecting Plücker ray geometry into the scalable generation engine. By synthesizing heterogeneous long-tail data, we observe a significant improvement in closed-loop driving robustness in handling long-tail events. By measuring the extrapolative view synthesis and pose metrics, we validate the effectiveness of OpenLongTail in visual fidelity, cross-view consistency, and ego-trajectory recovery.