자율주행 VLM에서 검증 가능한 추론을 위한 미래 궤적의 지연 노출
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs
August 3, 2026
저자: Zixuan Huang, Yang Zhou, Kaixuan Wang, Guli Zhang, Hongyan Xie, Yakun Zhu, Hao Geng, Yikun Ban, Deqing Wang
cs.AI
초록
최근 자율주행(AD)을 위한 비전-언어-행동(VLA) 모델들은 비전-언어 모델(VLM) 구성 요소의 추론 능력을 향상시키기 위해 사고 사슬(CoT) 감독을 점점 더 활용하고 있다. 그러나 기존의 주석 파이프라인은 일반적으로 교사 모델을 기록된 정답(GT) 미래 궤적에 노출시킨다. 우리는 이것이 궤적 앵커링 편향을 유발함을 실증적으로 보여준다: 교사 모델은 장면 증거로부터 결정을 추론하기보다 드러난 결과를 합리화하여, 인과적으로 덜 충실한 CoT를 생성하고 특히 인과적 추론이 까다로운 장면에서 훨씬 더 심각한 환각을 일으킨다. GT 궤적을 제거하면 이러한 지름길이 사라지지만, 개방형 궤적 생성은 고수준 의사결정을 정밀한 기하학적 합성 및 저수준 역학과 얽히게 만든다. 개방형 궤적 합성 없이 궤적 수준의 주행 결정을 검증 가능하게 만들기 위해, 우리는 명시적 궤적 후보들 중에서의 선택으로 계획을 정식화하는 자율주행 객관식 질문(AD-MCQ)을 도입한다. 여기서 한 걸음 더 나아가, 우리는 미래 궤적을 결정 전 앵커에서 결정 후 검증 대상으로 전환하는 DEFT-RLVR(Deferred Exposure of Future Trajectories for RLVR)을 제안한다. 실험 결과는 DEFT-RLVR이 일반적인 시각 능력을 유지하거나 오히려 향상시키면서 AD 추론을 개선함을 보여준다. VLM 단독 추론과 후보 구성을 통한 제어 가능한 난이도를 갖춘 AD-MCQ는 검증 가능한 AD 추론에 관한 향후 연구를 위한 유연하고, 확장 가능하며, 기능 확장이 용이한 기반을 제공한다.
English
Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reasoning capabilities of their Vision-Language Model (VLM) components, yet existing annotation pipelines commonly expose the teacher model to the logged ground-truth (GT) future trajectory. We empirically show that this induces trajectory anchoring bias: teacher models rationalize the revealed outcome rather than infer a decision from scene evidence, producing less causally faithful CoTs and substantially more severe hallucinations, especially in causally challenging scenes. Removing the GT trajectory eliminates this shortcut, but open-ended trajectory generation entangles high-level decision-making with precise geometric synthesis and low-level dynamics. To make trajectory-level driving decisions verifiable without requiring open-ended trajectory synthesis, we introduce Autonomous-Driving Multiple-Choice Question (AD-MCQ), which casts planning as selection among explicit trajectory candidates. Taking this a step further, we propose Deferred Exposure of Future Trajectories for RLVR (DEFT-RLVR) to transform future trajectories from pre-decision anchors into post-decision verification targets. Experimental results show that DEFT-RLVR improves AD reasoning while preserving or even enhancing general visual capabilities. With VLM-only inference and controllable difficulty through candidate construction, AD-MCQ provides a flexible, scalable, and extensible foundation for future research on verifiable AD reasoning.