WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
Ji Soo Lee, Xilun Chen, Pierce Chuang, Ashish Shenoy, Jason Wei, Dohwan Ko, et al.
arXiv:2609.05405v1Today’s pick is WearableQA, a new benchmark for asking whether AI can actually reason over real wearable data, not just summarize it. The challenge is that wearable streams are messy: they stretch across hundreds of days, include device noise, and differ a lot from person to person. WearableQA builds 4,084 multiple-choice questions from longitudinal records for 200 real users, combining time-series signals, blood biomarkers, and demographics. The questions are designed to test two kinds of reasoning: computing over the data itself, and interpreting what those measurements mean for health. They also probe single-signal versus cross-signal reasoning, so models have to connect patterns across multiple sources when needed. The authors use both physiological literature and population-level statistical validation to make the questions reliable. When they test 14 language models, performance ranges widely, but most still fall below 60 percent. That makes WearableQA both a realistic benchmark and a clear warning: health reasoning over everyday wearable data is still far from solved.
Also spotted that day
Previous daily papers