ChatPaper.aiChatPaper

自動運転VLMにおける検証可能な推論のための将来軌跡の遅延露出

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs

August 3, 2026
著者: Zixuan Huang, Yang Zhou, Kaixuan Wang, Guli Zhang, Hongyan Xie, Yakun Zhu, Hao Geng, Yikun Ban, Deqing Wang
cs.AI

要旨

近年、自動運転(AD)向けのVision-Language-Action(VLA)モデルでは、視覚言語モデル(VLM)コンポーネントの推論能力を高めるために、チェーン・オブ・ソート(CoT)教師信号を利用する手法が増えている。しかし、既存のアノテーションパイプラインでは、教師モデルに記録済みのグラウンドトゥルース(GT)未来軌跡を提示することが一般的である。本稿では、これが軌跡アンカリングバイアスを誘発することを実証的に示す。すなわち、教師モデルはシーンの証拠から意思決定を推論するのではなく、開示された結果を事後的に正当化してしまうため、因果的に忠実でないCoTを生成し、特に因果的推論が困難なシーンでは、ハルシネーションが著しく悪化する。GT軌跡を除去すればこのショートカットを排除できるが、オープンエンドな軌跡生成は、高次レベルの意思決定と精密な幾何学的合成、そして低次レベルのダイナミクスとを複雑に絡み合わせてしまう。軌跡レベルの運転意思決定を、オープンエンドな軌跡合成を必要とせずに検証可能にするため、本稿ではAutonomous-Driving Multiple-Choice Question(AD-MCQ)を導入する。これは、プランニングを明示的な軌跡候補の中からの選択として定式化するものである。さらに一歩進めて、Deferred Exposure of Future Trajectories for RLVR(DEFT-RLVR)を提案し、未来軌跡を意思決定前のアンカーから意思決定後の検証対象へと転換する。実験結果は、DEFT-RLVRがADの推論能力を向上させるとともに、一般的な視覚能力を維持ないし強化することを示している。VLMのみの推論が可能であり、候補構築により難易度を制御できるAD-MCQは、検証可能なAD推論に関する今後の研究のための、柔軟でスケーラブルかつ拡張可能な基盤を提供する。
English
Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reasoning capabilities of their Vision-Language Model (VLM) components, yet existing annotation pipelines commonly expose the teacher model to the logged ground-truth (GT) future trajectory. We empirically show that this induces trajectory anchoring bias: teacher models rationalize the revealed outcome rather than infer a decision from scene evidence, producing less causally faithful CoTs and substantially more severe hallucinations, especially in causally challenging scenes. Removing the GT trajectory eliminates this shortcut, but open-ended trajectory generation entangles high-level decision-making with precise geometric synthesis and low-level dynamics. To make trajectory-level driving decisions verifiable without requiring open-ended trajectory synthesis, we introduce Autonomous-Driving Multiple-Choice Question (AD-MCQ), which casts planning as selection among explicit trajectory candidates. Taking this a step further, we propose Deferred Exposure of Future Trajectories for RLVR (DEFT-RLVR) to transform future trajectories from pre-decision anchors into post-decision verification targets. Experimental results show that DEFT-RLVR improves AD reasoning while preserving or even enhancing general visual capabilities. With VLM-only inference and controllable difficulty through candidate construction, AD-MCQ provides a flexible, scalable, and extensible foundation for future research on verifiable AD reasoning.