arXiv: 2607.13216
일상 활동 분류에는 자세가 필요하고, 재구성에는 움직임이 필요하다.
Classifying daily activities needs posture, reconstructing them needs motion
July 14, 2026
저자: Arefeh Farahmandi, Gunnar Blohm
q-bio.NCq-bio.NCcs.AIcs.IRcs.LG
초록
인간은 잡음이 많고 복잡한 시각적 입력 속에서도 움직임을 쉽게 인식한다. 그러나 자극 내 어떤 정보가 인간이 움직임을 빠르게 분류할 수 있게 하는가? 이 질문을 다루기 위해 움직임 분석의 다양한 전략을 체계적으로 비교한 프레임워크는 지금까지 없었다. 본 연구에서는 MoVi 데이터셋의 16가지 일상 활동 동영상을 활용하여 세 가지 전략을 비교하였다: 시간적 움직임 기본 요소(TMPs)는 움직임을 시간적으로 매끄러운 기저 함수의 가중합으로 분해하고, 르장드르 다항식 계수는 관절 좌표 궤적을 직교 다항식 기저에 투영하며, 오토인코더 잠재 임베딩을 사용한다. 르장드르 계수와 TMPs는 가장 높은 분류 정확도를 보였으며, 오토인코더가 그 뒤를 이었다. 움직임 분류에 대해 두 가지 변별적 특징을 발견하였다. 가장 정보량이 많은 것은 신체의 일반적인 자세, 즉 한 활동을 다른 활동과 구별해주는 평균 공간 구성이다. 또한, 움직임 분류에 가장 예측력이 높은 9개의 주요 관절을 확인하였다. 흥미롭게도, 좋은 분류 정확도가 자동적으로 좋은 움직임 생성으로 이어지지는 않았다: 각 활동에 대해 움직임을 재구성했을 때, TMPs는 시간적 동역학을 보존하여 지각적으로 자연스러운 움직임을 생성한 반면, 르장드르 계수의 재구성은 평균 자세만 유지하고 마치 멈춰 있는 것처럼 보였다. 이러한 결과는 움직임 정보가 조직되는 방식에 있어 해리가 존재함을 보여준다: 신체의 정적 구성은 어떤 활동이 수행되는지 분류하기에 충분하지만, 움직임이 어떻게 전개되는지 재구성하기 위해서는 움직임의 시간적 동역학이 필요하다. 이러한 구분은 시각 시스템이 빠른 행동 인식을 위해 어떤 특징에 의존할 수 있는지를 명확히 하며, 임상 응용에서 효율적인 움직임 선별을 위해 자세적 특징이 활용될 수 있는 반면, 움직임 생성이 목표인 경우에는 동적 정보가 필수적임을 시사한다.
English
Humans recognize movements effortlessly, even from noisy and complex visual input. But what information in the stimulus allows humans to rapidly classify movements? No framework has systematically compared different strategies of movement analysis to address this question. Here, we used videos of 16 daily activities from the MoVi dataset and compared three strategies: Temporal Movement Primitives (TMPs), which decompose movements into weighted sums of temporally smooth basis functions; Legendre polynomial coefficients, which project joint-coordinate trajectories onto an orthogonal polynomial basis; and Autoencoder latent embeddings. Legendre coefficients and TMPs achieved the highest classifier accuracy, followed by autoencoders. We found two discriminative features for movement classification. The most informative is the general posture of the body, the average spatial configuration that distinguishes one activity from another. Additionally, we identified 9 critical joints that are most predictive for movement classification. Interestingly, good classification accuracy did not automatically lead to good movement generation: when we reconstructed movements for each activity, TMPs preserved the temporal dynamics and produced perceptually natural motion, whereas reconstructions from Legendre coefficients retained only the average posture and appeared frozen. These results reveal a dissociation in how movement information is organized: the static configuration of the body suffices to classify what activity is performed, but the temporal dynamics of movement are required to reconstruct how it unfolds. This distinction clarifies which features the visual system may rely upon for rapid action recognition, and suggests that postural features could enable efficient movement screening in clinical applications, while dynamic information remain essential wherever movement generation is the goal.