DriveDNA: 運転スタイル識別のための大規模マルチモーダル自然運転データセットとベンチマーク
DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification
July 26, 2026
著者: Yuhang Wang, Lingyao Li, Hao Zhou
cs.AI
要旨
運転スタイルは、車両の運転方法における安定した運転者固有のパターンを捉える。しかし、自然主義データにおいては、運転者が異なる車両、異なる道路、異なる条件下で観察されるため、この信号を分離することは困難であり、モデルが車両固有または状況固有の規則性を運転者固有のスタイルと誤認する可能性がある。我々は、個人化された運転スタイルモデリングのための大規模自然主義データセットおよびベンチマークであるDriveDNAを提案する。これは、115車種にわたる465人の運転者による4,121回の走行から構成され、合計975時間の人間制御による運転を10Hzで前方ビデオとともに収録したものであり、日常使用におけるコミュニティドライバーから収集された。DriveDNAは、運転スタイルを、類似条件下での車両の動き方における一貫した運転者固有の行動パターンと定義する。本ベンチマークは、少数ショット運転者再同定、個人化行動予測、条件一致比較の3つの中核タスクを通じてこの信号を評価し、行動アノテーションに加え、大規模な人間による監査を経た6クラスにわたる276,248件のルール生成操縦イベントを提供する。我々は、古典的記述子、教師ありおよび自己教師あり時系列エンコーダ、マルチモーダル融合、確率的予測、ゼロショット基盤モデルに及ぶベースラインを、固定マルチシードプロトコルの下で評価する。学習された表現は、未見の運転者において古典的記述子を大幅に上回り(AUROC .935 vs .707)、一致した運転条件下でも運転者固有の情報を保持する一方、記述子の性能は偶然のレベルに近づく。ビデオのみのモデルは同等の再同定精度を達成するが、深刻な経路リークを示し、強い認識が運転行動ではなく文脈上の近道から生じる可能性があることを示している。これらの知見は、信頼性の高い運転スタイル評価には、学習された表現の行動的価値と、車両、走行、条件の交絡に対するロバスト性の両方を評価する必要があることを示している。
English
Driving style captures stable, driver-specific patterns in how a vehicle is driven. In naturalistic data, however, this signal is hard to isolate because drivers are observed in different vehicles, on different roads, and under different conditions, so models may mistake vehicle- or situation-specific regularities for driver-specific style. We introduce DriveDNA, a large-scale naturalistic dataset and benchmark for personalized driving-style modeling, comprising 4,121 drives from 465 drivers across 115 vehicle models and totaling 975 hours of human-controlled driving at 10 Hz with forward video, collected from community drivers in everyday use. DriveDNA defines driving style as a consistent, driver-specific behavioral pattern in how a vehicle moves under similar conditions. The benchmark evaluates this signal through three core tasks: few-shot driver re-identification, personalized behavior prediction, and condition-matched comparison, and provides behavioral annotations plus 276,248 rule-generated maneuver events across six classes with large-scale human auditing. We evaluate baselines spanning classical descriptors, supervised and self-supervised time-series encoders, multimodal fusion, probabilistic prediction, and zero-shot foundation models under a fixed multi-seed protocol. Learned representations substantially outperform classical descriptors on unseen drivers (AUROC .935 vs. .707) and retain driver-specific information under matched driving conditions, while descriptor performance approaches chance. Video-only models achieve comparable re-identification accuracy but exhibit severe route leakage, showing that strong recognition may arise from contextual shortcuts rather than driving behavior. These findings show that reliable driving-style evaluation must assess both the behavioral value of learned representations and their robustness to vehicle, drive, and condition confounds.