Daily AI Paper

Daily AI Paper

2026-09-07

Archived
2026-09-07

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

Ji Soo Lee, Xilun Chen, Pierce Chuang, Ashish Shenoy, Jason Wei, Dohwan Ko, et al.

arXiv:2609.05405v1

Today’s pick is WearableQA, a new benchmark for asking whether AI can actually reason over real wearable data, not just summarize it. The challenge is that wearable streams are messy: they stretch across hundreds of days, include device noise, and differ a lot from person to person. WearableQA builds 4,084 multiple-choice questions from longitudinal records for 200 real users, combining time-series signals, blood biomarkers, and demographics. The questions are designed to test two kinds of reasoning: computing over the data itself, and interpreting what those measurements mean for health. They also probe single-signal versus cross-signal reasoning, so models have to connect patterns across multiple sources when needed. The authors use both physiological literature and population-level statistical validation to make the questions reliable. When they test 14 language models, performance ranges widely, but most still fall below 60 percent. That makes WearableQA both a realistic benchmark and a clear warning: health reasoning over everyday wearable data is still far from solved.

Previous daily papers

2026-09-07UniMate: One Unified Model to Animate Diverse Skeletons2026-09-06Compile by Training: Turning Natural-Language Specifications into Local Neural Functions2026-09-06Rethinking On-Policy Distillation of Large Language Models II: One Training Example2026-09-06SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents2026-09-05From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research2026-09-05Last Translation Benchmark2026-09-03A Common Measure of Communication for Speech Brain-Computer Interfaces2026-09-03Discriminative World Models for Web Agents2026-09-03Post-Training Language Models for Gold-Medal Performance in Coding Competitions2026-09-02Mechanism Design for Alignment and Control2026-09-02Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation2026-09-02The Rise of Verbal Reinforcement Learning2026-09-01Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions2026-09-01Aspire: Can Models Self-Evolve from Vague Goals?2026-09-01PaperGym: Rubric-Centered Evolution for Research-Plan Generation2026-08-31DARTS: Decoder-Aware Representation Tuning via Surgery for Model Merging2026-08-31Video Generative Models as Geometry Learner2026-08-31When Robots Mishear Us: Mapping the Safety Risks of Voice-Controlled Embodied AI2026-08-30CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes2026-08-30Boosting LLM Exploration via Weak-Model Guidance in RLVR2026-08-30CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators2026-08-30TTPO: Test-Time Policy Optimization2026-08-29TTPO: Test-Time Policy Optimization2026-08-28TTPO: Test-Time Policy Optimization2026-08-27Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings2026-08-26BrowserForge: Scaling Web Episode via Parallel Browser Sandboxes2026-08-25ConvergeFlow: Language Flow with Provable Convergence to Token Embeddings2026-08-24VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences2026-08-23Inducing Task Models from Computer-Use Traces2026-08-22Inducing Task Models from Computer-Use Traces2026-08-21Inducing Task Models from Computer-Use Traces2026-08-20SPADE: Self-Play in Adaptive Synthetic Executable Environments2026-08-19From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation2026-08-18Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text2026-08-17Universal Thermodynamic Interatomic Potentials for Crystalline Materials2026-08-16OmniScientist: An Omni-Modal Omni-Discipline AI Scientist