WearableQA:針對真實世界穿戴式資料進行健康推理的基準測試
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
September 4, 2026
作者: Ji Soo Lee, Xilun Chen, Pierce Chuang, Ashish Shenoy, Jason Wei, Dohwan Ko, Hyunwoo J. Kim, Benoit Corda
cs.AI
摘要
穿戴式感測的最新進展使得生理與行為訊號的連續監測成為可能,然而現有基準測試鮮少評估 AI 系統是否能對真實使用者的縱貫性穿戴式紀錄進行推理。我們提出 WearableQA,這是一個包含 4,084 道 10 選項選擇題的基準測試,題目由 200 位真實使用者的穿戴式時間序列、血液生物標記與人口統計資料建構而成,每位使用者最多具有 500 天的每日測量資料。WearableQA 保留真實的穿戴式資料分布,其中包括裝置雜訊與個體間變異性。為評估不同的推理能力,我們引入 16 種題型,並沿著兩個互補軸向組織:資料推理與健康推理,用於區分對縱貫性測量進行運算與生理詮釋;以及單一訊號推理與跨訊號推理,用於區分對個別訊號的推理與多個訊號的整合。為大規模建構可靠題目,我們採用雙重依據框架,結合以文獻為依據的生理發現與經統計驗證、以族群為依據的生理模式。這使得該基準測試能夠捕捉在真實世界穿戴式資料中觀察到的有意義關係。對 14 個專有與開源 LLM 的評估顯示,WearableQA 能有效區分模型能力,其效能範圍從 19.6% 到 72.9%,相對於 10% 的隨機猜測基準。此外,WearableQA 仍遠未被解決:多數模型的準確率低於 60%。總體而言,WearableQA 為評估 LLM 對真實世界穿戴式資料的推理提供了一個真實且具診斷性的基準測試。
English
Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing benchmarks rarely evaluate whether AI systems can reason over a real user's longitudinal wearable record. We introduce WearableQA, a benchmark comprising 4,084 10-option multiple-choice questions constructed from the wearable time series, blood biomarkers, and demographics of 200 real users, each with up to 500 days of daily measurements. WearableQA preserves authentic wearable distributions that include device noise and inter-individual variability. To evaluate distinct reasoning capabilities, we introduce 16 question types organized along two complementary axes: data versus health reasoning, which distinguishes computation over longitudinal measurements from physiological interpretation; and single- versus cross-signal reasoning, which separates reasoning about individual signals from the integration of multiple signals. To construct reliable questions at scale, we adopt a dual-grounding framework that combines literature-grounded physiological findings with statistically validated population-grounded physiological patterns. This enables the capture of meaningful relationships observed in real-world wearable data. Evaluation of 14 proprietary and open-source LLMs demonstrates that WearableQA effectively differentiates model capabilities, with performance ranging from 19.6% to 72.9% against a 10% chance baseline. Moreover, WearableQA remains far from solved: most models achieve accuracies below 60%. Overall, WearableQA provides a realistic and diagnostic benchmark for evaluating LLM reasoning over real-world wearable data.