← Latest papers
📄 health informatics

Sense and Sensibility: Sensible Quality Control and Data Decision-Making in Passive Sensing Data for Adolescent Substance Use Risk Prediction

This paper provides principled recommendations for preprocessing passive sensing data from smartphones and wearables to improve adolescent substance use risk prediction, using the ABCD Study to illustrate how rigorous data quality control and transparent feature engineering can address challenges like outlier handling, platform differences, and protocol deviations.

Original authors: Hong, A., Diaz, J. L., Kennedy, T. M., Wang, F. L.

Published 2026-10-08
📖 6 min read🧠 Deep dive

Original authors: Hong, A., Diaz, J. L., Kennedy, T. M., Wang, F. L.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine trying to understand the daily habits of teenagers by watching them through a window. You might see them sleeping, walking, or staring at their phones, but you cannot see the whole picture unless you are inside the room with them. For decades, scientists have relied on teenagers to fill out surveys about their lives, asking them to remember how many hours they slept or how much they moved. But memory is fallible, and people often forget or misreport their habits. A newer approach involves giving young people smartphones and wearable fitness trackers that quietly record their movements, sleep patterns, and phone usage in the real world. This method, known as passive sensing, offers a window into behavior that is continuous and objective. However, just because a device records data does not mean that data is always accurate or easy to use. The devices can malfunction, the software can glitch, and the way teenagers interact with technology varies wildly. Before scientists can use this flood of information to predict health risks, such as the likelihood of a teenager starting to use drugs or alcohol, they must first learn how to clean the data, deciding which numbers are real and which are mistakes.

A team of researchers set out to solve these messy problems using data from a massive, long-term study of American youth called the Adolescent Brain Cognitive Development Study. They focused on a group of nearly 5,000 teenagers who were around 13 or 14 years old and had agreed to wear a fitness tracker and use a special app on their phones for three weeks. The goal was to see if the patterns in this digital data could help identify which young people were at higher risk for substance use. But as they began to look at the numbers, the researchers found that the data was far from perfect. Some teenagers had recorded only a few days of activity, while others had weeks of data. Some had strange readings, like sleeping for 24 hours straight or not moving at all. The team had to figure out how to handle these oddities without throwing away the very behaviors that might signal a problem.

The researchers discovered that the most extreme and unlikely numbers usually came from the teenagers who had worn their devices the least. If a person only wore their tracker for one day and happened to be very inactive that day, the computer might calculate that they were inactive all the time. This suggested that the strange numbers were not necessarily signs of a dangerous lifestyle, but rather signs that the data was incomplete. Instead of deleting every single data point that looked weird, the team decided to look at the person first. They set a rule that a teenager had to have at least five weekdays and two weekend days of data to be included in the study. This approach allowed them to keep teenagers who had irregular sleep or activity patterns, as long as they had enough data to show that those patterns were real and not just a fluke of a single bad day. By doing this, they preserved the messy, irregular reality of teenage life, which is often where the most important clues about health risks are hidden.

The team also found that the type of phone a teenager used changed the data in surprising ways. The study used a system to record how much time teens spent typing on their phones. However, the system worked differently on iPhones compared to Android phones. On iPhones, the system could not record typing unless the user switched to a special keyboard provided by the study. The researchers suspected that many iPhone users found this switch annoying and went back to their normal keyboard, meaning the study missed a lot of their typing activity. When they compared the recorded data with what the teenagers said they did, they found that iPhone users seemed to type much less than they actually did, especially when their usage was low. This meant that the data was not just different; it was systematically biased against one group of people. The researchers concluded that they could not simply mix the data from both phone types together without accounting for this difference, and they had to treat the phone operating system as a major factor in their analysis.

Another challenge was deciding which days to count. The study had a specific three-week window where the researchers were actively monitoring the participants. However, the devices kept recording data before the study started and after it ended. The researchers compared the behavior during the official study weeks with the behavior during the extra days. They found that teenagers acted differently when they knew they were being watched versus when they were not. During the official weeks, their sleep and activity patterns were consistent with each other. But on the days outside the study window, the patterns became erratic and unpredictable. This led the team to decide that only the data collected during the official three-week period was reliable enough to use. They discarded the extra days, realizing that more data was not always better if that data came from a different context.

Finally, the researchers had to build the tools to turn these raw numbers into something a computer could understand. They had to create features that described not just how much a teen slept, but how consistent their sleep was. They had to figure out how to measure the difference between a bedtime of 11:00 PM and 1:00 AM without getting confused by the way clocks wrap around midnight. They also had to clean up errors in the data that the study organizers had missed, such as missing values that were accidentally marked as zero, which would make a teenager look like they never moved when they actually had. After all this careful cleaning and organizing, the team was left with a smaller group of 428 teenagers, but they were confident that the data from this group was trustworthy.

The paper does not claim to have found a perfect way to predict substance use, nor does it say that the problems with digital data are solved. Instead, it offers a set of practical rules for other scientists who want to use this kind of data. The main lesson is that researchers need to be careful about how they handle outliers and missing information. They should not just delete strange numbers, because those numbers might be the most important ones. They should also check if the data changes depending on the device used or the time it was collected. By following these steps, scientists can ensure that their findings reflect the true lives of the teenagers they are studying, rather than the mistakes of the machines they are using. This work provides a roadmap for turning a chaotic stream of digital signals into a clear picture of adolescent behavior, paving the way for better tools to help young people stay healthy.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →