HumanTracker:包括的かつ人間整合的なモーショントラッキングベンチマークを目指して
HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark
August 13, 2026
著者: Dairu Liu, Zekun Qi, Jiayu Zeng, Ruixi Yu, Yu Guan, Yintianrun Zhang, Xuchuan Chen, Sikai Liang, Zekai Li, Chenghuai Lin, Xinqiang Yu, Wenyao Zhang, He Wang, Li Yi
cs.AI
要旨
ヒューマノイドの動作追跡は遠隔操作や全身模倣の中核をなす技術であるが、その評価は映像に対する人間の知覚と乖離することが多い。運動学的誤差はフレーム単位の姿勢差分を平均化するものの、最も重要となる物理的アーティファクト、特に不安定な支持や誤った接触(足の滑りやタイミングを誤った接地など)を見逃してしまう。一方、広く用いられているテストスイートは規模が小さく、接触を多く含む長期的な行動を評価するために必要な多様性を欠いている。本稿では、ヒューマノイド追跡評価を知覚的に整合させつつ、スケーラブルにするHumanTrackerを提案する。HumanTrackerベンチマークには、複数のプロのパフォーマーから収集した約153時間分の光学式モーション軌跡が含まれており、きめ細かな診断のためのテキストラベルを備えた4つの動作ファミリーに整理されている。さらに、2万4千の動作ペアを含む1万2千の動作ペアで学習された選好整合メトリクスであるHumanScoreを提案する。代表的かつ最先端のトラッカー群に対する評価において、HumanScoreは人間の選好をより高精度に予測し、運動学的メトリクスではしばしば見逃される接触の失敗や安定性の欠如を明らかにする。
English
Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos. Kinematic errors average per-frame pose differences but miss the physical artifacts that matter most, particularly unstable support and incorrect contacts such as foot skating and mistimed touch-downs. Meanwhile, widely used test suites are small and lack the diversity needed to stress contact-rich, long-horizon behaviors. We introduce HumanTracker to make humanoid tracking evaluation both perceptually aligned and scalable. The HumanTracker benchmark contains approximately 153 hours of optical motion trajectories from multiple professional performers, organized into four motion families with text labels for fine-grained diagnosis. We further propose HumanScore, a preference-aligned metric trained on 12K motion pairs containing 24K motions. Across representative state-of-the-art trackers, HumanScore better predicts human preferences and reveals contact and stability failures that kinematic metrics often miss.