ChatPaper.aiChatPaper

意図を持って行動する:視覚言語行動モデルのための行動意図の蒸留

Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models

August 24, 2026
著者: Sangoh Lee, Sangwoo Mo, Wook-Shin Han
cs.AI

要旨

視覚-言語-行動(VLA)モデルはマルチモーダルなコンテキストをロボットの行動に変換できるが、その行動デコーダは依然として主に行動模倣(ビヘイビアクローニング)によって訓練されている。これは、どの運動指令が実証されたかを監視する一方で、指示の下でその行動が果たす局所的な目的を暗黙のままにしている。未来ベースの監視は、フレーム、潜在観測、軌跡、または動作表現を用いて行動学習を豊かにするが、これらの信号は、今後起こり得ることの特定の実現形態を捉えるものであり、今後の行動の共有された意味的目標を捉えるものではない。我々は、行動レベルの意図を行動デコーダに蒸留する意図蒸留(INDI)を提案する。訓練中、凍結された教師VLMが、現在の観測、指示、粗い行動要約、および対応する実行ビデオから実証されたセグメントを解釈する。展開されたVLAは、その標準的な入力から、中間デコーダ層で得られたマルチモーダル意図表現を回復し、それを使用して、行動がどのように展開し、何を達成するかの表現とともに行動予測を組織化する。SimplerEnv-Bridgeでは、INDIはGR00T-N1.7を64.3%から84.7%に改善し、RoboCasa Kitchenでは、対照GR00T-N1.7ベースラインを64.1%から70.3%に改善し、両ベンチマークでπ_{0.5}に一貫した改善が見られる。実世界のタスクでは、INDIは平均成功率を62.0%から68.7%に改善し、より長いホライズンのタスクでは最大12.0ポイントの改善が見られる。さらなる分析により、回復された潜在表現がデコーダによって使用され、行動の目的と実行進捗を捉え、下流の予測を目的依存的に組織化することが示される。これらの結果は、行動デコーダが生成する行動の意味的目標を明示的にモデル化することの利点を示している。
English
Vision-Language-Action (VLA) models can turn multimodal context into robot actions, but their action decoders are still trained largely by behavior cloning. This supervises which motor command was demonstrated while leaving implicit the local objective served by the behavior under the instruction. Future-based supervision enriches action learning with frames, latent observations, trajectories, or motion representations, but these signals capture particular realizations of what may happen rather than the shared semantic objective of the forthcoming behavior. We propose Intention Distillation (INDI), which distills behavior-level intent into the action decoder. During training, a frozen teacher VLM interprets a demonstrated segment from the current observation, instruction, coarse action summary, and corresponding execution video. From its standard inputs, the deployed VLA recovers the resulting multimodal intent representation at an intermediate decoder layer and uses it to organize action prediction together with representations of how the behavior unfolds and what it achieves. On SimplerEnv-Bridge, INDI improves GR00T-N1.7 from 64.3% to 84.7%, and on RoboCasa Kitchen it improves the controlled GR00T-N1.7 baseline from 64.1% to 70.3%, with consistent gains on π_{0.5} across both benchmarks. In real-world tasks, INDI improves average success from 62.0% to 68.7%, with gains of up to 12.0 pp on longer-horizon tasks. Further analyses show that the recovered latent is used by the decoder, captures behavior objective and execution progress, and organizes downstream predictions in an objective-dependent manner. These results show that action decoders benefit from explicitly modeling the semantic objective of the behavior they generate.