SynthGait-19K:一個具物理基礎的合成視訊資料集,用於步態參數估測
SynthGait-19K: A Physically Grounded Synthetic Video Dataset for Gait Parameter Estimation
September 8, 2026
作者: Soroush Mehraban, Xin Lei Lin, Vida Adeli, Majid Mirmehdi, Amirhossein Dadashzadeh, Clint Hansen, Andrea Iaboni, Babak Taati
cs.AI
摘要
準確地從單目影片估計具臨床意義的步態參數,對於可擴展的行動能力評估至關重要,然而現有資料集的規模小、視角受限及視覺多樣性有限,限制了進展。我們提出 SynthGait-19k,一個具物理基礎的合成影片資料集,包含 19,272 段步行影片,衍生自涵蓋 437 位受試者的 6,427 段動作捕捉序列,並配對 SMPL 動作與六項步態參數標註。為建構此資料集,我們開發 Gait2Vid,其透過 SMPL 統一異質動作捕捉記錄,並在可控視角與場景外觀下合成多樣化的 RGB 步行影片。我們評估生成影片與其條件步態運動學的一致性,並以測力板量測驗證所擷取的步態事件。使用 SynthGait-19K,我們對直接 RGB、基於姿態、生物力學及人體網格重建方法進行基準測試,並分析視角、訓練資料規模及合成至真實的領域偏移。我們亦提出 GaitXFormer,作為估計步態參數的直接 RGB 參考模型。無論在 GaitXFormer 或基於姿態的架構上,合成監督皆能有效遷移至真實影片,展現其在不同表徵上的實用性。我們進一步發現,空間步態參數對視覺領域偏移更為敏感,且單獨改善 HMR 重建不一定能轉化為更佳的下游步態估計。
English
Accurate estimation of clinically meaningful gait parameters from monocular video is important for scalable mobility assessment, yet progress is limited by the small scale, restricted viewpoints, and limited visual diversity of existing datasets. We introduce SynthGait-19k, a physically grounded synthetic video dataset containing 19,272 walking videos derived from 6,427 MoCap sequences across 437 subjects, with paired SMPL motion and annotations for six gait parameters. To construct the dataset, we develop Gait2Vid, which unifies heterogeneous MoCap recordings through SMPL and synthesizes diverse RGB walking videos under controllable viewpoints and scene appearances. We assess the generated videos for consistency with their conditioning gait kinematics and validate extracted gait events against force-platform measurements. Using SynthGait-19K, we benchmark direct RGB, pose-based, biomechanical, and human-mesh-recovery approaches and analyze viewpoint, training-data scale, and synthetic-to-real domain shift. We also introduce GaitXFormer as a direct RGB reference model for estimating gait parameters. Synthetic supervision transfers effectively to real videos across both GaitXFormer and a pose-based architecture, demonstrating utility across different representations. We further find that spatial gait parameters are more sensitive to visual domain shift and that improved HMR reconstruction alone does not necessarily translate to improved downstream gait estimation.