SynthGait-19K:歩行パラメータ推定のための物理法則に基づく合成動画データセット
SynthGait-19K: A Physically Grounded Synthetic Video Dataset for Gait Parameter Estimation
September 8, 2026
著者: Soroush Mehraban, Xin Lei Lin, Vida Adeli, Majid Mirmehdi, Amirhossein Dadashzadeh, Clint Hansen, Andrea Iaboni, Babak Taati
cs.AI
要旨
単眼動画から臨床的に意味のある歩行パラメータを正確に推定することは、スケーラブルな移動機能評価にとって重要であるが、その進展は既存データセットの規模の小ささ、制限された視点、限られた視覚的多様性によって妨げられている。我々はSynthGait-19kを導入する。これは、437名の被験者にわたる6,427件のMoCapシーケンスから生成された19,272本の歩行動画を含み、SMPLモーションと6種類の歩行パラメータのアノテーションを対にして備えた、物理に基づく合成動画データセットである。データセットを構築するために、我々はGait2Vidを開発した。これは、異種のMoCap記録をSMPLを通じて統合し、制御可能な視点およびシーン外観の下で多様なRGB歩行動画を合成する。生成された動画について、条件とした歩行運動学との整合性を評価し、抽出された歩行イベントをフォースプレート測定に対して検証する。SynthGait-19Kを用いて、直接RGB、姿勢ベース、生体力学、人体メッシュ復元のアプローチをベンチマークし、視点、学習データ規模、合成から実データへのドメインシフトを分析する。また、歩行パラメータを推定するための直接RGB参照モデルとしてGaitXFormerを導入する。合成データによる教師あり学習は、GaitXFormerと姿勢ベースアーキテクチャの両方において実動画へ効果的に転移し、異なる表現にわたる有用性を示す。さらに、空間的歩行パラメータは視覚的ドメインシフトに対してより敏感であり、HMR再構築の改善だけでは下流の歩行推定の改善に必ずしもつながらないことを見出した。
English
Accurate estimation of clinically meaningful gait parameters from monocular video is important for scalable mobility assessment, yet progress is limited by the small scale, restricted viewpoints, and limited visual diversity of existing datasets. We introduce SynthGait-19k, a physically grounded synthetic video dataset containing 19,272 walking videos derived from 6,427 MoCap sequences across 437 subjects, with paired SMPL motion and annotations for six gait parameters. To construct the dataset, we develop Gait2Vid, which unifies heterogeneous MoCap recordings through SMPL and synthesizes diverse RGB walking videos under controllable viewpoints and scene appearances. We assess the generated videos for consistency with their conditioning gait kinematics and validate extracted gait events against force-platform measurements. Using SynthGait-19K, we benchmark direct RGB, pose-based, biomechanical, and human-mesh-recovery approaches and analyze viewpoint, training-data scale, and synthetic-to-real domain shift. We also introduce GaitXFormer as a direct RGB reference model for estimating gait parameters. Synthetic supervision transfers effectively to real videos across both GaitXFormer and a pose-based architecture, demonstrating utility across different representations. We further find that spatial gait parameters are more sensitive to visual domain shift and that improved HMR reconstruction alone does not necessarily translate to improved downstream gait estimation.