WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
Ji Soo Lee, Xilun Chen, Pierce Chuang, Ashish Shenoy, Jason Wei, Dohwan Ko, et al.
arXiv:2609.05405v1Wearable devices collect a rich, messy stream of health data, but most AI benchmarks don’t really test whether a model can reason over that kind of real-world record. This paper introduces WearableQA, a benchmark built from up to 500 days of data for 200 real users, combining wearable time series, blood biomarkers, and demographics. The questions are multiple choice, but they’re designed to probe different skills: simple computation over measurements, physiological interpretation, and reasoning across multiple signals at once. What makes it especially useful is that the questions are grounded in both medical literature and statistically validated patterns found in the data, so they reflect relationships that actually occur in practice. When the authors tested 14 proprietary and open-source language models, performance ranged widely, and most models still scored below 60 percent. That matters because wearables are becoming a major source of continuous health information, and we need benchmarks that reveal whether AI can interpret them responsibly, not just summarize them.
Also spotted that day
Previous daily papers