arXiv: 2607.12909
適用於低功耗邊緣平台的基於視覺的即時跌倒偵測
Real-time fall detection based on vision for low-power edge platforms
July 14, 2026
作者: Wenjun Xia, Zhicheng Peng, Haopeng Li, Zhengdi Zhang
q-bio.NCq-bio.NCcs.AIcs.CV
摘要
跌倒偵測對於年長者照護與智慧監控至關重要,然而當前主流的基於視覺的方法大多將其框架化為靜態姿勢分類或離散時序模式匹配,根本上忽略了人體支撐系統的不穩定性動態。本文提出一個基於物理啟發的跌倒偵測框架,將跌倒重新定義為耦合動態系統中的穩定度喪失事件。我們引入一個新穎的雙LTC架構,包含質心子系統與支撐基底子系統,兩者均採用液態時間常數神經網路進行實例化,以自適應時間常數連續建模慣性軌跡演化與地面接觸調整,賦予跌倒運動物理可解釋性。一個可學習的耦合模組模擬兩子系統間的物理交互作用,同時一個穩定性流形分類器在聯合潛在空間中,透過李亞普諾夫啟發的穩定度指標偵測邊界跨越。輔以反事實軌跡投影與碰撞時間估計,進一步實現不可逆性評估與早期預警。該架構設計支援三狀態(正常、跌倒中、已跌倒)預測範式;在本初步研究中,我們在二類資料集(正常 vs. 跌倒中)上驗證核心穩定度判別能力,完整的二至三狀態時序轉換則留待未來工作。與傳統CNN-RNN流程不同,本方法編碼連續時間機械慣性,形成一個參數少於五萬的網路,可於資源受限的邊緣裝置進行即時推論。大量實驗證明其具有競爭性的準確度與優越的物理可解釋性,驗證了其在低運算視覺跌倒偵測中的有效性。
English
Falling detection is vital for elderly care and intelligent surveillance; however, prevailing vision-based approaches predominantly frame it as static pose classification or discrete temporal pattern matching, fundamentally overlooking the instability dynamics of the human support system. This paper proposes a physics-informed falling detection framework that recasts falling as a stability-loss event in a coupled dynamical system. We introduce a novel dual-LTC architecture comprising a Center-of-Mass (CoM) subsystem and a Base-of-Support (BoS) subsystem, both instantiated as Liquid Time-Constant (LTC) neural networks to continuously model inertial trajectory evolution and ground-contact adjustment through adaptive time constants, Physical interpretability of falling motion. A learnable coupling module emulates physical interaction between the two subsystems, while a Stability Manifold classifier operates in the joint latent space to detect boundary crossing via Lyapunov-inspired stability metrics. Complementary counterfactual trajectory projection and Time-to-Collision (TTC) estimation further enable irreversibility assessment and early warning. The architecture is designed to support a three-state prediction paradigm (Normal, Falling, Fallen); in this preliminary study, we validate the core stability discrimination capability on a two-class dataset (Normal vs. Falling), leaving the full three-state temporal transition to future work. Unlike conventional CNN--RNN pipelines, the proposed formulation encodes continuous-time mechanical inertia, yielding a sub-50K-parameter network capable of real-time inference on resource-constrained edge devices. Extensive experiments demonstrate competitive accuracy with superior physical interpretability, validating its efficacy for low-compute visual fall detection.