arXiv: 2607.13216

日常活动的分类需要姿态,其重建需要运动。

Classifying daily activities needs posture, reconstructing them needs motion

July 14, 2026
作者: Arefeh Farahmandi, Gunnar Blohm
q-bio.NCq-bio.NCcs.AIcs.IRcs.LG

摘要

人类能够毫不费力地识别动作,即使视觉输入嘈杂且复杂。但刺激中的何种信息使人类能够快速分类动作?尚无框架系统比较不同的运动分析策略来回答这一问题。本研究使用MoVi数据集中16种日常活动的视频,比较了三种策略:时间运动基元(TMPs),将动作分解为时间平滑基函数的加权求和;勒让德多项式系数,将关节坐标轨迹投影到正交多项式基上;以及自编码器潜变量嵌入。勒让德系数和TMPs实现了最高的分类器准确率,其次是自编码器。我们发现了两个用于动作分类的判别性特征:最具信息量的是身体的一般姿态,即区分不同活动的平均空间构型;此外,还识别出9个对动作分类最具预测性的关键关节。有趣的是,高分类准确率并未自动带来好的动作生成:当我们重建每种活动的动作时,TMPs保留了时间动态并产生了感知上自然的运动,而勒让德系数的重建仅保留了平均姿态,显得像冻结一般。这些结果揭示了运动信息组织方式的分离:身体的静态构型足以分类所执行的活动,但重建动作如何展开则需要运动的时间动态。这一区分阐明了视觉系统可能依赖哪些特征进行快速动作识别,并提示在临床应用中,姿态特征可实现高效的运动筛查,而动态信息在需要动作生成的场景中仍不可或缺。
English
Humans recognize movements effortlessly, even from noisy and complex visual input. But what information in the stimulus allows humans to rapidly classify movements? No framework has systematically compared different strategies of movement analysis to address this question. Here, we used videos of 16 daily activities from the MoVi dataset and compared three strategies: Temporal Movement Primitives (TMPs), which decompose movements into weighted sums of temporally smooth basis functions; Legendre polynomial coefficients, which project joint-coordinate trajectories onto an orthogonal polynomial basis; and Autoencoder latent embeddings. Legendre coefficients and TMPs achieved the highest classifier accuracy, followed by autoencoders. We found two discriminative features for movement classification. The most informative is the general posture of the body, the average spatial configuration that distinguishes one activity from another. Additionally, we identified 9 critical joints that are most predictive for movement classification. Interestingly, good classification accuracy did not automatically lead to good movement generation: when we reconstructed movements for each activity, TMPs preserved the temporal dynamics and produced perceptually natural motion, whereas reconstructions from Legendre coefficients retained only the average posture and appeared frozen. These results reveal a dissociation in how movement information is organized: the static configuration of the body suffices to classify what activity is performed, but the temporal dynamics of movement are required to reconstruct how it unfolds. This distinction clarifies which features the visual system may rely upon for rapid action recognition, and suggests that postural features could enable efficient movement screening in clinical applications, while dynamic information remain essential wherever movement generation is the goal.
PDFJuly 19, 2026