從基礎到應用:在實踐中改進VLA模型
From Foundation to Application: Improving VLA Models in Practice
July 7, 2026
作者: Wei Wu, Fangjing Wang, Fan Lu, He Sun, Shi Liu, Yunnan Wang, Yibin Yan, Yong Wang, Shuailei Ma, Xinyang Wang, Yibin Liu, Shuai Yang, Tianxiang Zhou, Kejia Zhang, Lei Zhou, Cheng Su, Nan Xue, Bin Tan, Han Zhang, Youchao Zhang, Fei Liao, Xing Zhu, Yujun Shen, Kecheng Zheng
cs.AI
摘要
儘管近期視覺語言動作(VLA)基礎模型取得了進展,但實驗室條件與真實世界應用之間的差距仍持續阻礙其實際部署。為彌合此差距,我們提出 LingBot-VLA 2.0,透過三個功能領域的改進來推進 LingBot-VLA。(1)任務與具身形態的通用化。相較於前一版本,我們徹底改造資料處理流程,並整理約六萬小時的預訓練資料,其中包括五萬小時涵蓋二十種機器人配置的機器人軌跡,以及一萬小時的自我中心人類影片。(2)除了雙臂硬體平台之外,擴展動作空間。具體來說,我們的系統支援頭部、腰部、移動底座及靈巧手的自由度,從而使機器人能夠在實際場景中處理更複雜的任務。(3)預測動力學建模以改善時間推理。具體而言,我們將未來預測制定為代理任務,並透過影片表徵模型提供語意先驗知識,以及深度估計模型提供幾何線索。在通用設定下於 GM-100 基準進行的評估,驗證了這些提出的修改所帶來的正面影響。此外,得益於涵蓋全身自由度的擴展預訓練資料,LingBot-VLA-2.0 在兩個機器人平台上展現了強大的跨具身形態長程移動操作能力。
English
Despite recent progress of VLA foundation models, the disparity between laboratory conditions and real-world applications continues to impede their practical implementation. To bridge this gap, we present LingBot-VLA 2.0, which advances LingBot-VLA through improvements in three functional domains. (1) Generalization across tasks and embodiments. Compared to the previous version, we revamp the data processing pipeline and curate around 60,000 hours of data for pretraining, including 50,000 hours of robot trajectories spanning 20 robot configurations and 10,000 hours of egocentric human videos. (2) Expanded action space in addition to dual-arm hardware platforms. In particular, our system accommodates degrees of freedom for the heads, waists, mobile bases, and dexterous hands, thereby empowering the robots to tackle more complex tasks in practical scenarios. (3) Predictive dynamics modeling for improved temporal reasoning. Specifically, we formulate future prediction as a proxy task, facilitated by a video representation model for semantic priors and a depth estimation model for geometric cues. Evaluations on the GM-100 benchmark, conducted in a generalist setting, validate the beneficial impact of these proposed modifications. Furthermore, benefiting from the expanded pretraining data that covers whole-body degrees of freedom, LingBot-VLA-2.0 demonstrates strong cross-embodiment long-horizon mobile manipulation capability across the two robotic platforms.