arXiv: 2607.12909
低消費電力エッジプラットフォーム向けのリアルタイム視覚に基づく転倒検知
Real-time fall detection based on vision for low-power edge platforms
July 14, 2026
著者: Wenjun Xia, Zhicheng Peng, Haopeng Li, Zhengdi Zhang
q-bio.NCq-bio.NCcs.AIcs.CV
要旨
転倒検出は高齢者介護およびインテリジェント監視において極めて重要であるが、主流の視覚ベース手法はこれを主に静的姿勢分類または離散的時間パターンマッチングとして捉えており、人間の支持系の不安定性ダイナミクスを根本的に見落としている。本論文では、転倒を連成動的システムにおける安定性喪失事象として再定義する、物理情報に基づく転倒検出フレームワークを提案する。我々は、重心(CoM)サブシステムと支持基底(BoS)サブシステムからなる新規なデュアルLTCアーキテクチャを導入し、両サブシステムをLiquid Time-Constant(LTC)ニューラルネットワークとして実装することで、適応的時間定数を通じて慣性軌道の時間発展と接地調整を連続的にモデル化し、転倒動作の物理的解釈可能性を実現する。学習可能な結合モジュールは両サブシステム間の物理的相互作用を模倣し、一方、安定性多様体分類器は結合潜在空間において動作し、リアプノフに着想を得た安定性指標を用いて境界越えを検出する。補完的な反実仮想軌道投影と衝突余裕時間(TTC)推定により、不可逆性評価と早期警告がさらに可能となる。本アーキテクチャは3状態予測パラダイム(通常、転倒中、転倒後)をサポートするように設計されているが、本予備研究では2クラスデータセット(通常 vs. 転倒中)で核心的な安定性判別能力を検証し、完全な3状態時間遷移は将来の課題とする。従来のCNN-RNNパイプラインとは異なり、提案手法は連続時間の機械的慣性を符号化し、5万パラメータ未満のネットワークでリソース制約のあるエッジデバイス上でのリアルタイム推論を可能にする。広範な実験により、優れた物理的解釈可能性を伴う競争力のある精度が実証され、低計算コストの視覚的転倒検出における有効性が確認された。
English
Falling detection is vital for elderly care and intelligent surveillance; however, prevailing vision-based approaches predominantly frame it as static pose classification or discrete temporal pattern matching, fundamentally overlooking the instability dynamics of the human support system. This paper proposes a physics-informed falling detection framework that recasts falling as a stability-loss event in a coupled dynamical system. We introduce a novel dual-LTC architecture comprising a Center-of-Mass (CoM) subsystem and a Base-of-Support (BoS) subsystem, both instantiated as Liquid Time-Constant (LTC) neural networks to continuously model inertial trajectory evolution and ground-contact adjustment through adaptive time constants, Physical interpretability of falling motion. A learnable coupling module emulates physical interaction between the two subsystems, while a Stability Manifold classifier operates in the joint latent space to detect boundary crossing via Lyapunov-inspired stability metrics. Complementary counterfactual trajectory projection and Time-to-Collision (TTC) estimation further enable irreversibility assessment and early warning. The architecture is designed to support a three-state prediction paradigm (Normal, Falling, Fallen); in this preliminary study, we validate the core stability discrimination capability on a two-class dataset (Normal vs. Falling), leaving the full three-state temporal transition to future work. Unlike conventional CNN--RNN pipelines, the proposed formulation encodes continuous-time mechanical inertia, yielding a sub-50K-parameter network capable of real-time inference on resource-constrained edge devices. Extensive experiments demonstrate competitive accuracy with superior physical interpretability, validating its efficacy for low-compute visual fall detection.