arXiv: 2607.13216
日常生活活動を分類するには姿勢が必要であり、それらを再構成するには動作が必要である。
Classifying daily activities needs posture, reconstructing them needs motion
July 14, 2026
著者: Arefeh Farahmandi, Gunnar Blohm
q-bio.NCq-bio.NCcs.AIcs.IRcs.LG
要旨
人間は、ノイズが多く複雑な視覚入力からでも、動きを容易に認識することができる。しかし、刺激の中のどのような情報が、人間の迅速な動作分類を可能にしているのだろうか。この疑問に取り組むために、動作分析の異なる戦略を体系的に比較した枠組みはこれまで存在しなかった。本研究では、MoViデータセットから16種類の日常生活活動の動画を用い、三つの戦略を比較した。すなわち、動きを時間的に滑らかな基底関数の重み付き和に分解する時間的動作プリミティブ(TMP)、関節座標の軌跡を直交多項式基底に投影するルジャンドル多項式係数、そしてオートエンコーダの潜在埋め込みである。分類器の精度が最も高かったのはルジャンドル係数とTMPであり、次いでオートエンコーダであった。動作分類において、二つの識別的特徴が明らかになった。最も情報量の多い特徴は体の全体的な姿勢、すなわち活動を互いに区別する平均的な空間配置である。さらに、動作分類に最も予測力を持つ9つの重要な関節を特定した。興味深いことに、高い分類精度は必ずしも良好な動作生成には結びつかなかった。各活動の動作を再構成したところ、TMPは時間的ダイナミクスを保持し、知覚的に自然な動きを生成したのに対し、ルジャンドル係数からの再構成は平均姿勢のみを保持し、静止したように見えた。これらの結果は、動作情報の組織化における乖離を明らかにしている。すなわち、身体の静的な配置は行われている活動の分類には十分であるが、動作がどのように展開するかを再構成するには、動作の時間的ダイナミクスが必要である。この区別は、視覚系が素早い動作認識のためにどのような特徴に依存している可能性があるかを明確にし、臨床応用においては姿勢特徴が効率的な動作スクリーニングを可能にする一方、動作生成が目的となる場面では動的情報が不可欠であることを示唆している。
English
Humans recognize movements effortlessly, even from noisy and complex visual input. But what information in the stimulus allows humans to rapidly classify movements? No framework has systematically compared different strategies of movement analysis to address this question. Here, we used videos of 16 daily activities from the MoVi dataset and compared three strategies: Temporal Movement Primitives (TMPs), which decompose movements into weighted sums of temporally smooth basis functions; Legendre polynomial coefficients, which project joint-coordinate trajectories onto an orthogonal polynomial basis; and Autoencoder latent embeddings. Legendre coefficients and TMPs achieved the highest classifier accuracy, followed by autoencoders. We found two discriminative features for movement classification. The most informative is the general posture of the body, the average spatial configuration that distinguishes one activity from another. Additionally, we identified 9 critical joints that are most predictive for movement classification. Interestingly, good classification accuracy did not automatically lead to good movement generation: when we reconstructed movements for each activity, TMPs preserved the temporal dynamics and produced perceptually natural motion, whereas reconstructions from Legendre coefficients retained only the average posture and appeared frozen. These results reveal a dissociation in how movement information is organized: the static configuration of the body suffices to classify what activity is performed, but the temporal dynamics of movement are required to reconstruct how it unfolds. This distinction clarifies which features the visual system may rely upon for rapid action recognition, and suggests that postural features could enable efficient movement screening in clinical applications, while dynamic information remain essential wherever movement generation is the goal.