ChatPaper.aiChatPaper

HumanTracker:迈向全面且与人类对齐的运动跟踪基准

HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

August 13, 2026
作者: Dairu Liu, Zekun Qi, Jiayu Zeng, Ruixi Yu, Yu Guan, Yintianrun Zhang, Xuchuan Chen, Sikai Liang, Zekai Li, Chenghuai Lin, Xinqiang Yu, Wenyao Zhang, He Wang, Li Yi
cs.AI

摘要

人形运动跟踪是遥操作和全身模仿的核心任务,然而其评估结果往往与人们在视频中感知到的质量不一致。运动学误差逐帧计算姿态差异的平均值,却忽略了最关键的身体物理伪影,特别是支撑不稳定以及接触错误,如脚部打滑和触地时机不当。与此同时,广泛使用的测试集规模较小,缺乏评估接触密集、长时域行为所需的多样性。我们提出HumanTracker,使人形跟踪评估兼具感知对齐性和可扩展性。HumanTracker基准测试包含来自多位专业表演者约153小时的光学运动轨迹数据,按四类运动族进行组织,并配有文本标签以支持细粒度诊断。我们进一步提出HumanScore,一种基于偏好对齐的度量指标,在包含24K段运动、由12K个运动对组成的数据集上训练而成。在具有代表性的最先进跟踪器上,HumanScore能够更准确地预测人类偏好,并揭示运动学度量指标常常遗漏的接触与稳定性失效问题。
English
Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos. Kinematic errors average per-frame pose differences but miss the physical artifacts that matter most, particularly unstable support and incorrect contacts such as foot skating and mistimed touch-downs. Meanwhile, widely used test suites are small and lack the diversity needed to stress contact-rich, long-horizon behaviors. We introduce HumanTracker to make humanoid tracking evaluation both perceptually aligned and scalable. The HumanTracker benchmark contains approximately 153 hours of optical motion trajectories from multiple professional performers, organized into four motion families with text labels for fine-grained diagnosis. We further propose HumanScore, a preference-aligned metric trained on 12K motion pairs containing 24K motions. Across representative state-of-the-art trackers, HumanScore better predicts human preferences and reveals contact and stability failures that kinematic metrics often miss.