ChatPaper.aiChatPaper

DriveDNA:一個大規模多模態自然駕駛資料集與駕駛風格識別基準

DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification

July 26, 2026
作者: Yuhang Wang, Lingyao Li, Hao Zhou
cs.AI

摘要

駕駛風格捕捉了車輛駕駛過程中穩定且具駕駛員特定性的模式。然而,在自然駕駛資料中,此訊號難以分離,因為駕駛員在不同車輛、不同道路及不同條件下被觀測,因此模型可能將車輛或情境特定規律誤認為駕駛員特定風格。我們提出 DriveDNA,一個大規模自然駕駛資料集與基準測試,用於個人化駕駛風格建模,包含來自 465 位駕駛員、涵蓋 115 種車輛型號的 4,121 趟行程,總計 975 小時以 10 Hz 取樣的人類操控駕駛,並附有前方影像,資料源自社群駕駛員的日常使用。DriveDNA 將駕駛風格定義為在相似條件下車輛移動時具一致性且駕駛員特定的行為模式。該基準測試透過三項核心任務評估此訊號:少樣本駕駛員重新識別、個人化行為預測及條件匹配比較,並提供行為標註,以及經大規模人工審核的 276,248 筆涵蓋六類規則生成的操作事件。我們在固定多重種子協議下評估了多種基準方法,包括經典描述符、有監督與自監督時間序列編碼器、多模態融合、概率預測及零樣本基礎模型。學習表徵在未見過的駕駛員上顯著優於經典描述符(AUROC .935 對比 .707),且在匹配駕駛條件下保留駕駛員特定資訊,而描述符表現則接近隨機。僅使用影像的模型達到可比的重新識別準確度,但表現出嚴重的路線洩漏,顯示強識別能力可能源於情境捷徑而非駕駛行為。這些發現表明,可靠的駕駛風格評估必須同時評估學習表徵的行為價值,及其對車輛、行程與條件混雜因素的穩健性。
English
Driving style captures stable, driver-specific patterns in how a vehicle is driven. In naturalistic data, however, this signal is hard to isolate because drivers are observed in different vehicles, on different roads, and under different conditions, so models may mistake vehicle- or situation-specific regularities for driver-specific style. We introduce DriveDNA, a large-scale naturalistic dataset and benchmark for personalized driving-style modeling, comprising 4,121 drives from 465 drivers across 115 vehicle models and totaling 975 hours of human-controlled driving at 10 Hz with forward video, collected from community drivers in everyday use. DriveDNA defines driving style as a consistent, driver-specific behavioral pattern in how a vehicle moves under similar conditions. The benchmark evaluates this signal through three core tasks: few-shot driver re-identification, personalized behavior prediction, and condition-matched comparison, and provides behavioral annotations plus 276,248 rule-generated maneuver events across six classes with large-scale human auditing. We evaluate baselines spanning classical descriptors, supervised and self-supervised time-series encoders, multimodal fusion, probabilistic prediction, and zero-shot foundation models under a fixed multi-seed protocol. Learned representations substantially outperform classical descriptors on unseen drivers (AUROC .935 vs. .707) and retain driver-specific information under matched driving conditions, while descriptor performance approaches chance. Video-only models achieve comparable re-identification accuracy but exhibit severe route leakage, showing that strong recognition may arise from contextual shortcuts rather than driving behavior. These findings show that reliable driving-style evaluation must assess both the behavioral value of learned representations and their robustness to vehicle, drive, and condition confounds.