arXiv: 2607.13216
分類日常活動需要姿態,重建它們則需要運動
Classifying daily activities needs posture, reconstructing them needs motion
July 14, 2026
作者: Arefeh Farahmandi, Gunnar Blohm
q-bio.NCq-bio.NCcs.AIcs.IRcs.LG
摘要
人類能夠毫不費力地辨識動作,即使面對充滿雜訊且複雜的視覺輸入。然而,究竟是刺激中的哪些訊息讓人類能夠快速分類動作?目前尚無系統性比較不同動作分析策略的框架來探討此問題。本研究採用MoVi資料集中16種日常活動的影片,比較三種策略:時間動作基元(TMPs),可將動作分解為時間平滑基底函數的加權總和;勒讓德多項式係數,可將關節座標軌跡投影至正交多項式基底;以及自編碼器潛在嵌入。勒讓德係數與TMPs達到最高的分類器準確率,其次為自編碼器。我們發現兩個有助於動作分類的區別性特徵。最具資訊性的特徵是身體的整體姿勢,即區分不同活動的平均空間配置。此外,我們識別出9個對動作分類最具預測力的關鍵關節。有趣的是,良好的分類準確率並未自動導向良好的動作生成:當我們針對每項活動重建動作時,TMPs保留了時間動態,產生知覺上自然的運動;而勒讓德係數的重建僅保留平均姿勢,呈現如凍結般的狀態。這些結果揭示了動作訊息組織方式的分離:身體的靜態配置足以分類正在執行的活動,但動作的時間動態卻是重建其如何展開所必需的。此區別有助於釐清視覺系統可能依賴哪些特徵進行快速動作辨識,並暗示姿勢特徵可在臨床應用中實現高效的動作篩檢,而動態訊息在動作生成為目標的情境中仍不可或缺。
English
Humans recognize movements effortlessly, even from noisy and complex visual input. But what information in the stimulus allows humans to rapidly classify movements? No framework has systematically compared different strategies of movement analysis to address this question. Here, we used videos of 16 daily activities from the MoVi dataset and compared three strategies: Temporal Movement Primitives (TMPs), which decompose movements into weighted sums of temporally smooth basis functions; Legendre polynomial coefficients, which project joint-coordinate trajectories onto an orthogonal polynomial basis; and Autoencoder latent embeddings. Legendre coefficients and TMPs achieved the highest classifier accuracy, followed by autoencoders. We found two discriminative features for movement classification. The most informative is the general posture of the body, the average spatial configuration that distinguishes one activity from another. Additionally, we identified 9 critical joints that are most predictive for movement classification. Interestingly, good classification accuracy did not automatically lead to good movement generation: when we reconstructed movements for each activity, TMPs preserved the temporal dynamics and produced perceptually natural motion, whereas reconstructions from Legendre coefficients retained only the average posture and appeared frozen. These results reveal a dissociation in how movement information is organized: the static configuration of the body suffices to classify what activity is performed, but the temporal dynamics of movement are required to reconstruct how it unfolds. This distinction clarifies which features the visual system may rely upon for rapid action recognition, and suggests that postural features could enable efficient movement screening in clinical applications, while dynamic information remain essential wherever movement generation is the goal.