ChatPaper.aiChatPaper

Van basis tot toepassing: VLA-modellen verbeteren in de praktijk

From Foundation to Application: Improving VLA Models in Practice

July 7, 2026
Auteurs: Wei Wu, Fangjing Wang, Fan Lu, He Sun, Shi Liu, Yunnan Wang, Yibin Yan, Yong Wang, Shuailei Ma, Xinyang Wang, Yibin Liu, Shuai Yang, Tianxiang Zhou, Kejia Zhang, Lei Zhou, Cheng Su, Nan Xue, Bin Tan, Han Zhang, Youchao Zhang, Fei Liao, Xing Zhu, Yujun Shen, Kecheng Zheng
cs.AI

Samenvatting

Ondanks de recente vooruitgang van VLA-fundamentmodellen blijft de kloof tussen laboratoriumomstandigheden en toepassingen in de praktijk een belemmering voor hun daadwerkelijke implementatie. Om deze kloof te overbruggen presenteren we LingBot-VLA 2.0, een verbetering van LingBot-VLA op drie functionele domeinen. (1) Generalisatie over taken en belichamingen. Vergeleken met de vorige versie hebben we de dataproverwerkingspijplijn herzien en ongeveer 60.000 uur aan data samengesteld voor pretraining, waaronder 50.000 uur aan robottrajecten verspreid over 20 robotconfiguraties en 10.000 uur aan egocentrische menselijke video's. (2) Uitgebreide actieruimte naast hardwareplatformen met twee armen. In het bijzonder ondersteunt ons systeem vrijheidsgraden voor hoofden, tailles, mobiele bases en behendige handen, waardoor robots worden bekrachtigd om complexere taken aan te pakken in praktische scenario's. (3) Voorspellende dynamica-modellering voor verbeterd temporeel redeneren. Specifiek formuleren we toekomstvoorspelling als een proxitaak, ondersteund door een videorepresentatiemodel voor semantische voorkennis en een diepteschattingsmodel voor geometrische aanwijzingen. Evaluaties op de GM-100 benchmark, uitgevoerd in een generalistische setting, bevestigen de gunstige impact van deze voorgestelde aanpassingen. Bovendien, profiterend van de uitgebreide pretrainingdata die vrijheidsgraden van het hele lichaam omvat, toont LingBot-VLA-2.0 een sterk cross-embodiment vermogen tot mobiele manipulatie over lange tijdsperioden op de twee robotplatforms.
English
Despite recent progress of VLA foundation models, the disparity between laboratory conditions and real-world applications continues to impede their practical implementation. To bridge this gap, we present LingBot-VLA 2.0, which advances LingBot-VLA through improvements in three functional domains. (1) Generalization across tasks and embodiments. Compared to the previous version, we revamp the data processing pipeline and curate around 60,000 hours of data for pretraining, including 50,000 hours of robot trajectories spanning 20 robot configurations and 10,000 hours of egocentric human videos. (2) Expanded action space in addition to dual-arm hardware platforms. In particular, our system accommodates degrees of freedom for the heads, waists, mobile bases, and dexterous hands, thereby empowering the robots to tackle more complex tasks in practical scenarios. (3) Predictive dynamics modeling for improved temporal reasoning. Specifically, we formulate future prediction as a proxy task, facilitated by a video representation model for semantic priors and a depth estimation model for geometric cues. Evaluations on the GM-100 benchmark, conducted in a generalist setting, validate the beneficial impact of these proposed modifications. Furthermore, benefiting from the expanded pretraining data that covers whole-body degrees of freedom, LingBot-VLA-2.0 demonstrates strong cross-embodiment long-horizon mobile manipulation capability across the two robotic platforms.