Leren bewegen voordat je leert doen: Taak-agnostische pre-training voor VLA's
Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs
July 2, 2026
Auteurs: Junhao Shi, Siyin Wang, Xiaopeng Yu, Li Ji, Jingjing Gong, Xipeng Qiu
cs.AI
Samenvatting
Visie-Taal-Actie (VTA) modellen worden fundamenteel beperkt door de schaarste aan expert demonstraties — tripletten van observaties, instructies en acties die op grote schaal kostbaar zijn om te verzamelen. Wij stellen dat dit knelpunt voortkomt uit het door elkaar halen van twee afzonderlijke leerdoelen: het verwerven van fysieke competentie (hoe te bewegen) en het verwerven van semantische afstemming (wat te doen). Cruciaal is dat alleen het laatste taal-supervisie vereist. Voortbouwend op deze Decompositiehypothese introduceren we Taak-agnostische Pretraining (TAP), een tweefasenkader dat eerst overdraagbare motorische priorissen leert uit goedkope, ongelabelde interactiegegevens — waaronder weggegooide off-task trajecten en autonoom robotspel — via een zelfgecontroleerde Inverse Dynamica-doelstelling. Een lichte tweede fase verankert deze priorissen vervolgens in taal met minimale expertgegevens. Op de SIMPLER benchmark bereikt TAP de prestaties van modellen getraind op meer dan 1M expert trajecten, terwijl het ordes van grootte minder gelabelde data gebruikt, wat een absolute winst van 10% oplevert ten opzichte van standaard gedragsclonen. Op een echte WidowX-platform behoudt TAP 25% succes onder camerastoringen waar op internet-schaal gebaseerde basislijnen instorten tot 0%, wat aantoont dat taak-agnostische pretraining robuuste, overdraagbare fysieke representaties produceert en een schaalbare weg vooruit biedt voor Embodied AI.
English
Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations -- triplets of observations, instructions, and actions that are costly to collect at scale. We argue that this bottleneck stems from conflating two distinct learning objectives: acquiring physical competence (how to move) and acquiring semantic alignment (what to do). Crucially, only the latter requires language supervision. Building on this Decomposition Hypothesis, we propose Task-Agnostic Pretraining (TAP), a two-stage framework that first learns transferable motor priors from cheap, unlabeled interaction data -- including discarded off-task trajectories and autonomous robot play -- via a self-supervised Inverse Dynamics objective. A lightweight second stage then grounds these priors in language using minimal expert data. On the SIMPLER benchmark, TAP matches models trained on over 1M expert trajectories while using orders of magnitude less labeled data, yielding a 10% absolute gain over standard behavior cloning. On a real-world WidowX platform, TAP retains 25% success under camera perturbations where internet-scale baselines collapse to 0%, demonstrating that task-agnostic pretraining produces robust, transferable physical representations and offers a scalable path forward for Embodied AI.