面向自动驾驶VLM可验证推理的未来轨迹延迟暴露
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs
August 3, 2026
作者: Zixuan Huang, Yang Zhou, Kaixuan Wang, Guli Zhang, Hongyan Xie, Yakun Zhu, Hao Geng, Yikun Ban, Deqing Wang
cs.AI
摘要
近期用于自动驾驶(AD)的视觉-语言-动作(VLA)模型越来越多地利用思维链(CoT)监督来增强其视觉-语言模型(VLM)组件的推理能力,然而现有标注流程通常会让教师模型接触记录中的真值(GT)未来轨迹。我们通过实验证明,这会诱发轨迹锚定偏差:教师模型会对已揭示的结果进行合理化解释,而非根据场景证据推断决策,从而产生因果忠实度较低的思维链,并在因果挑战性场景中产生明显更严重的幻觉。移除GT轨迹可消除这一捷径,但开放式轨迹生成会将高层决策与精确的几何合成及低层动力学纠缠在一起。为使轨迹级驾驶决策无需开放式轨迹合成即可验证,我们引入了自动驾驶多项选择题(AD-MCQ),将规划问题转化为对显式轨迹候选项的选择。在此基础上,我们进一步提出DEFT-RLVR(Deferred Exposure of Future Trajectories for RLVR),将未来轨迹从决策前锚点转变为决策后验证目标。实验结果表明,DEFT-RLVR在提升AD推理能力的同时,保持甚至增强了通用视觉能力。凭借仅依赖VLM的推理方式以及通过候选轨迹构建实现的可控难度,AD-MCQ为未来可验证AD推理研究提供了灵活、可扩展且具备良好延展性的基础。
English
Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reasoning capabilities of their Vision-Language Model (VLM) components, yet existing annotation pipelines commonly expose the teacher model to the logged ground-truth (GT) future trajectory. We empirically show that this induces trajectory anchoring bias: teacher models rationalize the revealed outcome rather than infer a decision from scene evidence, producing less causally faithful CoTs and substantially more severe hallucinations, especially in causally challenging scenes. Removing the GT trajectory eliminates this shortcut, but open-ended trajectory generation entangles high-level decision-making with precise geometric synthesis and low-level dynamics. To make trajectory-level driving decisions verifiable without requiring open-ended trajectory synthesis, we introduce Autonomous-Driving Multiple-Choice Question (AD-MCQ), which casts planning as selection among explicit trajectory candidates. Taking this a step further, we propose Deferred Exposure of Future Trajectories for RLVR (DEFT-RLVR) to transform future trajectories from pre-decision anchors into post-decision verification targets. Experimental results show that DEFT-RLVR improves AD reasoning while preserving or even enhancing general visual capabilities. With VLM-only inference and controllable difficulty through candidate construction, AD-MCQ provides a flexible, scalable, and extensible foundation for future research on verifiable AD reasoning.