WearableQA: 실제 환경 웨어러블 데이터에 대한 건강 추론 벤치마크
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
September 4, 2026
저자: Ji Soo Lee, Xilun Chen, Pierce Chuang, Ashish Shenoy, Jason Wei, Dohwan Ko, Hyunwoo J. Kim, Benoit Corda
cs.AI
초록
웨어러블 센싱의 최근 발전은 생리적 및 행동적 신호의 지속적 모니터링을 가능하게 하지만, 기존 벤치마크는 AI 시스템이 실제 사용자의 종단적 웨어러블 기록에 대해 추론할 수 있는지를 거의 평가하지 않는다. 우리는 실제 사용자 200명의 웨어러블 시계열, 혈액 바이오마커, 인구통계학적 정보로부터 구축된 4,084개의 10지선다형 객관식 문제로 구성된 벤치마크인 WearableQA를 제안한다. 각 사용자는 최대 500일간의 일일 측정값을 가진다. WearableQA는 기기 노이즈와 개인 간 변동성을 포함하는 실제 웨어러블 분포를 보존한다. 서로 다른 추론 능력을 평가하기 위해 우리는 두 개의 상호 보완적 축을 따라 구성된 16가지 문제 유형을 제안한다. 데이터 추론 대 건강 추론은 종단적 측정값에 대한 계산과 생리학적 해석을 구분하며, 단일 신호 추론 대 교차 신호 추론은 개별 신호에 대한 추론과 여러 신호의 통합을 분리한다. 대규모로 신뢰할 수 있는 문제를 구축하기 위해 우리는 문헌에 근거한 생리학적 발견과 통계적으로 검증된 인구집단 기반 생리학적 패턴을 결합하는 이중 그라운딩 프레임워크를 채택한다. 이는 실제 웨어러블 데이터에서 관찰되는 의미 있는 관계를 포착할 수 있게 한다. 14개의 독점 및 오픈소스 LLM을 평가한 결과, WearableQA는 모델 능력을 효과적으로 구별하며, 성능은 10% 우연 기준선 대비 19.6%에서 72.9% 범위에 이른다. 더욱이 WearableQA는 아직 해결과 거리가 멀다. 대부분의 모델이 60% 미만의 정확도를 달성한다. 종합적으로 WearableQA는 실제 웨어러블 데이터에 대한 LLM 추론을 평가하기 위한 현실적이고 진단적인 벤치마크를 제공한다.
English
Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing benchmarks rarely evaluate whether AI systems can reason over a real user's longitudinal wearable record. We introduce WearableQA, a benchmark comprising 4,084 10-option multiple-choice questions constructed from the wearable time series, blood biomarkers, and demographics of 200 real users, each with up to 500 days of daily measurements. WearableQA preserves authentic wearable distributions that include device noise and inter-individual variability. To evaluate distinct reasoning capabilities, we introduce 16 question types organized along two complementary axes: data versus health reasoning, which distinguishes computation over longitudinal measurements from physiological interpretation; and single- versus cross-signal reasoning, which separates reasoning about individual signals from the integration of multiple signals. To construct reliable questions at scale, we adopt a dual-grounding framework that combines literature-grounded physiological findings with statistically validated population-grounded physiological patterns. This enables the capture of meaningful relationships observed in real-world wearable data. Evaluation of 14 proprietary and open-source LLMs demonstrates that WearableQA effectively differentiates model capabilities, with performance ranging from 19.6% to 72.9% against a 10% chance baseline. Moreover, WearableQA remains far from solved: most models achieve accuracies below 60%. Overall, WearableQA provides a realistic and diagnostic benchmark for evaluating LLM reasoning over real-world wearable data.