WearableQA:実世界のウェアラブルデータに対する健康推論のためのベンチマーク
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
September 4, 2026
著者: Ji Soo Lee, Xilun Chen, Pierce Chuang, Ashish Shenoy, Jason Wei, Dohwan Ko, Hyunwoo J. Kim, Benoit Corda
cs.AI
要旨
ウェアラブルセンシングの最近の進歩により、生理学的および行動的シグナルの継続的モニタリングが可能になったが、既存のベンチマークは、AIシステムが実ユーザーの縦断的ウェアラブル記録に対して推論できるかどうかをほとんど評価していない。我々はWearableQAを導入する。これは、200名の実ユーザーのウェアラブル時系列、血液バイオマーカー、人口統計情報から構築された4,084問の10択多肢選択問題からなるベンチマークであり、各ユーザーは最大500日分の日次測定値を有する。WearableQAは、デバイスノイズや個人間変動を含む、実際のウェアラブルデータの分布を保持している。異なる推論能力を評価するため、我々は2つの相補的な軸に沿って整理された16種類の問題タイプを導入する。すなわち、データ推論対健康推論であり、これは縦断的測定値に対する計算と生理学的解釈を区別する。また、単一信号推論対クロスシグナル推論であり、これは個々の信号に関する推論と複数信号の統合を分離する。大規模に信頼できる問題を構築するために、我々は文献に基づく生理学的知見と、統計的に検証された母集団に基づく生理学的パターンを組み合わせるデュアルグラウンディング・フレームワークを採用する。これにより、実世界のウェアラブルデータで観察される意味のある関係を捉えることが可能になる。14個のプロプライエタリおよびオープンソースLLMの評価は、WearableQAがモデル能力を効果的に識別し、その性能が10%のチャンスレベルベースラインに対して19.6%から72.9%の範囲に及ぶことを示す。さらに、WearableQAは依然として解決にはほど遠い。ほとんどのモデルは60%未満の正解率を達成している。全体として、WearableQAは、実世界のウェアラブルデータに対するLLM推論を評価するための現実的で診断的なベンチマークを提供する。
English
Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing benchmarks rarely evaluate whether AI systems can reason over a real user's longitudinal wearable record. We introduce WearableQA, a benchmark comprising 4,084 10-option multiple-choice questions constructed from the wearable time series, blood biomarkers, and demographics of 200 real users, each with up to 500 days of daily measurements. WearableQA preserves authentic wearable distributions that include device noise and inter-individual variability. To evaluate distinct reasoning capabilities, we introduce 16 question types organized along two complementary axes: data versus health reasoning, which distinguishes computation over longitudinal measurements from physiological interpretation; and single- versus cross-signal reasoning, which separates reasoning about individual signals from the integration of multiple signals. To construct reliable questions at scale, we adopt a dual-grounding framework that combines literature-grounded physiological findings with statistically validated population-grounded physiological patterns. This enables the capture of meaningful relationships observed in real-world wearable data. Evaluation of 14 proprietary and open-source LLMs demonstrates that WearableQA effectively differentiates model capabilities, with performance ranging from 19.6% to 72.9% against a 10% chance baseline. Moreover, WearableQA remains far from solved: most models achieve accuracies below 60%. Overall, WearableQA provides a realistic and diagnostic benchmark for evaluating LLM reasoning over real-world wearable data.