DriveDNA:一种用于驾驶风格识别的大规模多模态自然驾驶数据集与基准
DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification
July 26, 2026
作者: Yuhang Wang, Lingyao Li, Hao Zhou
cs.AI
摘要
驾驶风格捕捉了车辆驾驶中稳定且具有驾驶员特异性的模式。然而,在自然驾驶数据中,这种信号难以分离,因为驾驶员是在不同车辆、不同道路及不同条件下被观测的,因此模型可能将车辆或情境的特定规律误判为驾驶员的特异性风格。我们提出了DriveDNA,这是一个用于个性化驾驶风格建模的大规模自然数据集和基准,包含来自465名驾驶员的4121次行程,覆盖115种车型,总计975小时由人类驾驶的10Hz频率数据(含前视视频),数据采集自社区驾驶员的日常使用场景。DriveDNA将驾驶风格定义为:在相似条件下,车辆运动过程中表现出的稳定且具有驾驶员特异性的行为模式。该基准通过三个核心任务评估这一信号:少样本驾驶员重识别、个性化行为预测以及条件匹配比较,并提供行为标注及经大规模人工审核的276,248个规则生成事件(涵盖六类机动事件)。我们在固定多种子协议下评估了涵盖经典描述符、监督与自监督时间序列编码器、多模态融合、概率预测及零样本基础模型的基线方法。学习得到的表示在未见驾驶员上显著优于经典描述符(AUROC 0.935 vs. 0.707),并在匹配驾驶条件下保留了驾驶员特异性信息,而描述符性能接近随机。仅使用视频的模型在重识别精度上表现相当,但存在严重的路线泄露问题,表明强识别能力可能源于情境捷径而非驾驶行为。这些发现表明,可靠的驾驶风格评估必须同时评估学习表示的行为价值及其对车辆、行程和条件混杂因素的鲁棒性。
English
Driving style captures stable, driver-specific patterns in how a vehicle is driven. In naturalistic data, however, this signal is hard to isolate because drivers are observed in different vehicles, on different roads, and under different conditions, so models may mistake vehicle- or situation-specific regularities for driver-specific style. We introduce DriveDNA, a large-scale naturalistic dataset and benchmark for personalized driving-style modeling, comprising 4,121 drives from 465 drivers across 115 vehicle models and totaling 975 hours of human-controlled driving at 10 Hz with forward video, collected from community drivers in everyday use. DriveDNA defines driving style as a consistent, driver-specific behavioral pattern in how a vehicle moves under similar conditions. The benchmark evaluates this signal through three core tasks: few-shot driver re-identification, personalized behavior prediction, and condition-matched comparison, and provides behavioral annotations plus 276,248 rule-generated maneuver events across six classes with large-scale human auditing. We evaluate baselines spanning classical descriptors, supervised and self-supervised time-series encoders, multimodal fusion, probabilistic prediction, and zero-shot foundation models under a fixed multi-seed protocol. Learned representations substantially outperform classical descriptors on unseen drivers (AUROC .935 vs. .707) and retain driver-specific information under matched driving conditions, while descriptor performance approaches chance. Video-only models achieve comparable re-identification accuracy but exhibit severe route leakage, showing that strong recognition may arise from contextual shortcuts rather than driving behavior. These findings show that reliable driving-style evaluation must assess both the behavioral value of learned representations and their robustness to vehicle, drive, and condition confounds.