OpenMHC: Accelerating the Science of Wearable Foundation Models
This paper introduces OpenMHC, a comprehensive open-access dataset containing over 60 million hours of wearable health data from nearly 12,000 participants, alongside open-source implementations of foundation models and a unified benchmark, to democratize and accelerate research in wearable health AI.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine your smartwatch is a tiny, tireless detective that never sleeps. Every minute of the day, it whispers secrets to your phone: how many steps you took, how fast your heart raced, how deep you slept, and even how many flights of stairs you conquered. For years, scientists have wanted to teach computers to listen to these whispers and understand our health better. They call these super-smart computers "foundation models." Think of them like a master chef who has tasted every dish in the world; once they've learned the flavors, they can invent new recipes or spot a bad ingredient instantly. But here's the problem: to become a master chef, you need a massive kitchen full of ingredients. In the world of health data, the biggest kitchens have been locked behind closed doors, owned by big tech companies, leaving researchers with only tiny, empty cupboards. Without enough data, these AI chefs can't learn the full story of how our bodies move and feel.
This is where a team of researchers from Stanford, Imperial College London, and others steps in with a massive key. They have unlocked the doors to the "My Heart Counts" study, a digital health experiment that has been running for over a decade. They are releasing a treasure trove called OpenMHC. This isn't just a small snack; it's a feast of over 67 million hours of wearable data from nearly 12,000 people. It includes everything from step counts and heart rates to sleep patterns and workout logs, collected from iPhones and Apple Watches. Along with this giant dataset, they are also sharing the "recipes" (code) and the "trained chefs" (AI models) they built to test it. Their goal is to let anyone, anywhere, try to cook up new ways to predict health issues, fill in missing data, or forecast how our bodies might behave tomorrow.
Important Note on the Ingredients: Before you start cooking, it's important to know that this kitchen has some specific quirks. The people who contributed data are mostly from the US, and the group is skewed toward white men in their late thirties. This means the "flavors" the AI learns might not taste exactly right for everyone else in the world. Also, because this was a digital study, all the health information (like whether someone has diabetes or how happy they feel) was self-reported by the users. This introduces some "noise" or uncertainty into the labels, which can make it harder for the AI to learn perfect patterns. Finally, the AI is designed to detect health traits based on the data collected at the time of the survey, not to predict future medical events like a heart attack happening next year.
The Big Reveal: What They Found
The researchers didn't just dump the data; they built a giant playground to see which AI models could actually make sense of it. They set up three main challenges:
- Prediction: Can the AI guess your health traits (like if you have diabetes or how happy you are) just by looking at your watch data?
- Imputation (Filling in the Blanks): Since people take their watches off to shower or charge them, the data has holes. Can the AI guess what happened during those missing minutes?
- Forecasting: Can the AI predict what your activity will look like in the next 24 hours?
Here is what they discovered:
1. The "Specialized Chef" Wins, but the "Simple Cook" is Still Tough
When they tested different AI models, a model called LSM-2 (a type of foundation model trained specifically on wearable data) turned out to be the best overall chef. It was the most consistent at predicting health outcomes and filling in missing data. However, the researchers found something surprising: a much simpler, older-style model called XGBoost (which is like a very organized, rule-based calculator) came in a very close second. In fact, for some tasks, the simple model was just as good as the fancy, complex AI. This suggests that while big AI is powerful, we shouldn't assume that "more complex" always means "better." Sometimes, a well-tuned, simple tool works wonders.
2. Big Language Models (Like Chatbots) Struggled
The team also tried using massive language models (the kind that write essays and chat with you) to solve these health puzzles. They fed the models the data as text or pictures of graphs. The result? These giant chatbots generally performed worse than the specialized health models and even the simple calculators. It seems that for reading the specific "language" of heartbeats and steps, a specialized tool is much better than a general-purpose one.
3. Long-Term Memory Helps
When it came to filling in the missing data (imputation), the models that could look at a user's history over several days did much better than those that only looked at a single day. It's like trying to guess what someone is doing right now: if you only see them for one minute, you might be wrong. But if you know they usually go for a run at 6 PM, you can make a much smarter guess. The paper suggests that using a user's long-term history is a key to making these models smarter.
4. Fairness Matters
The researchers also checked if the models treated everyone equally. They found that while the fancy AI models were accurate, they sometimes made bigger mistakes for certain groups of people (like different age groups or genders) compared to the simpler models. This is a crucial reminder that as we build these health tools, we must ensure they work fairly for everyone, not just the majority.
Why This Changes the Game
Before this paper, if you wanted to build an AI to understand wearable health data, you were stuck. You couldn't get the data, and you couldn't see how the big companies trained their models. It was like trying to learn to fly a plane without ever seeing a cockpit.
OpenMHC changes that by giving everyone the cockpit, the flight manual, and the training simulator. By releasing 11,894 participants' data spanning 13 years, along with the code to run the models, the authors have democratized the field. They aren't claiming to have solved all health mysteries yet; in fact, they admit that for some conditions, the AI still struggles to beat a simple baseline due to the nature of the self-reported data. But they have provided the foundation for the entire scientific community to start cooking, experimenting, and hopefully, discovering new ways to keep us healthy. The kitchen is now open, and the ingredients are ready for anyone to use.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.