DriveDNA: 운전 스타일 식별을 위한 대규모 다중 모달 자연적 주행 데이터셋 및 벤치마크
DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification
July 26, 2026
저자: Yuhang Wang, Lingyao Li, Hao Zhou
cs.AI
초록
운전 스타일은 차량이 어떻게 운전되는지에 대한 안정적이고 운전자 특유의 패턴을 포착한다. 그러나 자연주의 데이터에서는 운전자가 다른 차량, 다른 도로, 다른 조건에서 관찰되기 때문에 이 신호를 분리하기 어렵다. 따라서 모델이 차량 또는 상황 특이적 규칙성을 운전자 특유의 스타일로 오인할 수 있다. 본 논문에서는 대규모 자연주의 데이터셋이자 벤치마크인 DriveDNA를 소개한다. DriveDNA는 115개 차량 모델의 465명 운전자로부터 수집된 4,121회 주행 데이터를 포함하며, 총 975시간의 인간 제어 주행 데이터를 10Hz로 수집하였고 전방 비디오도 포함한다. 데이터는 일상적인 사용 환경에서 커뮤니티 운전자들로부터 수집되었다. DriveDNA는 운전 스타일을 유사한 조건에서 차량이 움직이는 방식에 대한 일관된 운전자 특유의 행동 패턴으로 정의한다. 벤치마크는 세 가지 핵심 과제, 즉 few-shot 운전자 재식별, 개인화된 행동 예측, 조건 일치 비교를 통해 이 신호를 평가하며, 행동 주석과 함께 276,248개의 규칙 기반 기동 이벤트를 6개 클래스로 제공하고 대규모 인간 검증을 거쳤다. 고정된 다중 시드 프로토콜 하에서 고전적 기술자, 지도 및 자기지도 시계열 인코더, 다중 모달 융합, 확률론적 예측, 제로샷 기반 모델을 포함한 기준선들을 평가한다. 학습된 표현은 보지 못한 운전자에 대해 고전적 기술자보다 월등히 뛰어난 성능을 보였으며(AUROC .935 대 .707), 일치된 운전 조건에서도 운전자 특유의 정보를 유지한 반면, 기술자의 성능은 우연 수준에 근접했다. 비디오 전용 모델은 비슷한 재식별 정확도를 달성했지만 심각한 경로 누출(route leakage)을 보였으며, 이는 강한 인식 성능이 운전 행동보다 맥락적 지름길에서 비롯될 수 있음을 보여준다. 이러한 발견은 신뢰할 수 있는 운전 스타일 평가가 학습된 표현의 행동적 가치와 차량, 주행 및 조건 혼란 변수에 대한 강건성을 모두 평가해야 함을 시사한다.
English
Driving style captures stable, driver-specific patterns in how a vehicle is driven. In naturalistic data, however, this signal is hard to isolate because drivers are observed in different vehicles, on different roads, and under different conditions, so models may mistake vehicle- or situation-specific regularities for driver-specific style. We introduce DriveDNA, a large-scale naturalistic dataset and benchmark for personalized driving-style modeling, comprising 4,121 drives from 465 drivers across 115 vehicle models and totaling 975 hours of human-controlled driving at 10 Hz with forward video, collected from community drivers in everyday use. DriveDNA defines driving style as a consistent, driver-specific behavioral pattern in how a vehicle moves under similar conditions. The benchmark evaluates this signal through three core tasks: few-shot driver re-identification, personalized behavior prediction, and condition-matched comparison, and provides behavioral annotations plus 276,248 rule-generated maneuver events across six classes with large-scale human auditing. We evaluate baselines spanning classical descriptors, supervised and self-supervised time-series encoders, multimodal fusion, probabilistic prediction, and zero-shot foundation models under a fixed multi-seed protocol. Learned representations substantially outperform classical descriptors on unseen drivers (AUROC .935 vs. .707) and retain driver-specific information under matched driving conditions, while descriptor performance approaches chance. Video-only models achieve comparable re-identification accuracy but exhibit severe route leakage, showing that strong recognition may arise from contextual shortcuts rather than driving behavior. These findings show that reliable driving-style evaluation must assess both the behavioral value of learned representations and their robustness to vehicle, drive, and condition confounds.