ChatPaper.aiChatPaper

延遲揭露未來軌跡以實現自動駕駛視覺語言模型中的可驗證推理

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs

August 3, 2026
作者: Zixuan Huang, Yang Zhou, Kaixuan Wang, Guli Zhang, Hongyan Xie, Yakun Zhu, Hao Geng, Yikun Ban, Deqing Wang
cs.AI

摘要

近期用於自動駕駛(AD)的視覺-語言-行動(VLA)模型日益利用思維鏈(CoT)監督來增強其視覺語言模型(VLM)組件的推理能力,然而現有的標註流程通常會將教師模型暴露於記錄在案的真實(GT)未來軌跡。我們通過實驗證明,這會誘發軌跡錨定偏差:教師模型會合理化所揭示的結果,而非從場景證據中推斷決策,從而產生因果忠實度較低的思維鏈,並導致明顯更嚴重的幻覺,尤其是在因果關係具有挑戰性的場景中。移除真實軌跡能消除這一捷徑,但開放式的軌跡生成會將高層次決策制定與精確的幾何合成及低層次動力學糾纏在一起。為了使軌跡層面的駕駛決策可驗證,又無需開放式軌跡合成,我們引入了自動駕駛選擇題(AD-MCQ),將規劃轉化為在明確軌跡候選項之間進行選擇。更進一步,我們提出了用於可驗證獎勵強化學習的未來軌跡延遲暴露方法(DEFT-RLVR),將未來軌跡從決策前的錨點轉變為決策後的驗證目標。實驗結果表明,DEFT-RLVR在提升AD推理能力的同時,能保持甚至增強通用視覺能力。憑藉僅使用VLM的推理以及通過候選項構建實現的可控難度,AD-MCQ為未來可驗證AD推理研究提供了靈活、可擴展且可延伸的基礎。
English
Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reasoning capabilities of their Vision-Language Model (VLM) components, yet existing annotation pipelines commonly expose the teacher model to the logged ground-truth (GT) future trajectory. We empirically show that this induces trajectory anchoring bias: teacher models rationalize the revealed outcome rather than infer a decision from scene evidence, producing less causally faithful CoTs and substantially more severe hallucinations, especially in causally challenging scenes. Removing the GT trajectory eliminates this shortcut, but open-ended trajectory generation entangles high-level decision-making with precise geometric synthesis and low-level dynamics. To make trajectory-level driving decisions verifiable without requiring open-ended trajectory synthesis, we introduce Autonomous-Driving Multiple-Choice Question (AD-MCQ), which casts planning as selection among explicit trajectory candidates. Taking this a step further, we propose Deferred Exposure of Future Trajectories for RLVR (DEFT-RLVR) to transform future trajectories from pre-decision anchors into post-decision verification targets. Experimental results show that DEFT-RLVR improves AD reasoning while preserving or even enhancing general visual capabilities. With VLM-only inference and controllable difficulty through candidate construction, AD-MCQ provides a flexible, scalable, and extensible foundation for future research on verifiable AD reasoning.