SynthGait-19K: 보행 파라미터 추정을 위한 물리 기반 합성 비디오 데이터셋
SynthGait-19K: A Physically Grounded Synthetic Video Dataset for Gait Parameter Estimation
September 8, 2026
저자: Soroush Mehraban, Xin Lei Lin, Vida Adeli, Majid Mirmehdi, Amirhossein Dadashzadeh, Clint Hansen, Andrea Iaboni, Babak Taati
cs.AI
초록
단안 비디오로부터 임상적으로 의미 있는 보행 매개변수를 정확하게 추정하는 것은 확장 가능한 이동성 평가에 중요하지만, 기존 데이터셋의 작은 규모, 제한된 시점, 제한된 시각적 다양성으로 인해 진전이 제한되어 왔다. 우리는 437명의 피험자에 걸친 6,427개 MoCap 시퀀스에서 파생된 19,272개 보행 비디오를 포함하며, SMPL 모션과 여섯 가지 보행 매개변수에 대한 주석이 함께 제공되는 물리적 근거에 기반한 합성 비디오 데이터셋 SynthGait-19k를 소개한다. 데이터셋을 구축하기 위해 우리는 Gait2Vid를 개발하는데, 이는 SMPL을 통해 이질적인 MoCap 기록을 통합하고 제어 가능한 시점과 장면 외형에서 다양한 RGB 보행 비디오를 합성한다. 우리는 생성된 비디오가 조건화된 보행 운동학과 일치하는지 평가하고, 추출된 보행 이벤트를 지면반력판 측정값과 대조하여 검증한다. SynthGait-19K를 사용하여 우리는 직접 RGB, 포즈 기반, 생체역학적, 인체 메시 복원 접근법을 벤치마킹하고 시점, 학습 데이터 규모, 합성-실제 도메인 이동을 분석한다. 우리는 또한 보행 매개변수 추정을 위한 직접 RGB 기준 모델로서 GaitXFormer를 소개한다. 합성 데이터 기반 지도는 GaitXFormer와 포즈 기반 아키텍처 모두에 걸쳐 실제 비디오로 효과적으로 전이되며, 서로 다른 표현 전반에서 유용성을 입증한다. 우리는 추가로 공간적 보행 매개변수가 시각적 도메인 이동에 더 민감하며, 개선된 HMR 복원만으로는 개선된 다운스트림 보행 추정으로 반드시 이어지지 않는다는 것을 발견한다.
English
Accurate estimation of clinically meaningful gait parameters from monocular video is important for scalable mobility assessment, yet progress is limited by the small scale, restricted viewpoints, and limited visual diversity of existing datasets. We introduce SynthGait-19k, a physically grounded synthetic video dataset containing 19,272 walking videos derived from 6,427 MoCap sequences across 437 subjects, with paired SMPL motion and annotations for six gait parameters. To construct the dataset, we develop Gait2Vid, which unifies heterogeneous MoCap recordings through SMPL and synthesizes diverse RGB walking videos under controllable viewpoints and scene appearances. We assess the generated videos for consistency with their conditioning gait kinematics and validate extracted gait events against force-platform measurements. Using SynthGait-19K, we benchmark direct RGB, pose-based, biomechanical, and human-mesh-recovery approaches and analyze viewpoint, training-data scale, and synthetic-to-real domain shift. We also introduce GaitXFormer as a direct RGB reference model for estimating gait parameters. Synthetic supervision transfers effectively to real videos across both GaitXFormer and a pose-based architecture, demonstrating utility across different representations. We further find that spatial gait parameters are more sensitive to visual domain shift and that improved HMR reconstruction alone does not necessarily translate to improved downstream gait estimation.