기초에서 응용까지: VLA 모델의 실무 개선하기
From Foundation to Application: Improving VLA Models in Practice
July 7, 2026
저자: Wei Wu, Fangjing Wang, Fan Lu, He Sun, Shi Liu, Yunnan Wang, Yibin Yan, Yong Wang, Shuailei Ma, Xinyang Wang, Yibin Liu, Shuai Yang, Tianxiang Zhou, Kejia Zhang, Lei Zhou, Cheng Su, Nan Xue, Bin Tan, Han Zhang, Youchao Zhang, Fei Liao, Xing Zhu, Yujun Shen, Kecheng Zheng
cs.AI
초록
최근 VLA 기반 모델의 발전에도 불구하고, 실험실 환경과 실제 응용 간의 격차는 여전히 실용적 구현을 저해하고 있다. 이러한 격차를 해소하기 위해, 우리는 LingBot-VLA 2.0을 제시한다. 이는 세 가지 기능적 영역의 개선을 통해 LingBot-VLA를 발전시킨 것이다. (1) 과제 및 구현체 간 일반화. 이전 버전과 비교하여, 데이터 처리 파이프라인을 재구성하고 사전 학습을 위해 약 60,000시간의 데이터를 수집하였으며, 이 중 50,000시간은 20가지 로봇 구성을 포괄하는 로봇 궤적 데이터이고 10,000시간은 에고센트릭 인간 비디오 데이터이다. (2) 이중 팔 하드웨어 플랫폼에 추가된 확장된 행동 공간. 특히, 우리 시스템은 헤드, 허리, 이동 베이스, 다자유도 손의 자유도를 수용함으로써 로봇이 실제 시나리오에서 더 복잡한 과제를 수행할 수 있도록 지원한다. (3) 향상된 시간적 추론을 위한 예측 역학 모델링. 구체적으로, 미래 예측을 대리 과제로 설정하고, 의미적 사전 정보를 위한 비디오 표현 모델과 기하학적 단서를 위한 깊이 추정 모델을 활용하여 이를 지원한다. 일반주의 환경에서 수행된 GM-100 벤치마크 평가는 이러한 제안된 수정 사항들의 긍정적 영향을 검증한다. 또한, 전신 자유도를 포괄하는 확장된 사전 학습 데이터의 이점을 바탕으로, LingBot-VLA-2.0은 두 로봇 플랫폼에 걸쳐 강력한 교차 구현체 장기 수평 이동 조작 능력을 입증한다.
English
Despite recent progress of VLA foundation models, the disparity between laboratory conditions and real-world applications continues to impede their practical implementation. To bridge this gap, we present LingBot-VLA 2.0, which advances LingBot-VLA through improvements in three functional domains. (1) Generalization across tasks and embodiments. Compared to the previous version, we revamp the data processing pipeline and curate around 60,000 hours of data for pretraining, including 50,000 hours of robot trajectories spanning 20 robot configurations and 10,000 hours of egocentric human videos. (2) Expanded action space in addition to dual-arm hardware platforms. In particular, our system accommodates degrees of freedom for the heads, waists, mobile bases, and dexterous hands, thereby empowering the robots to tackle more complex tasks in practical scenarios. (3) Predictive dynamics modeling for improved temporal reasoning. Specifically, we formulate future prediction as a proxy task, facilitated by a video representation model for semantic priors and a depth estimation model for geometric cues. Evaluations on the GM-100 benchmark, conducted in a generalist setting, validate the beneficial impact of these proposed modifications. Furthermore, benefiting from the expanded pretraining data that covers whole-body degrees of freedom, LingBot-VLA-2.0 demonstrates strong cross-embodiment long-horizon mobile manipulation capability across the two robotic platforms.