Micro-randomized Trials with Categorical Treatments and Binary Proximal Outcome: Causal Effect Estimation and Sample Size Calculation
This paper addresses the challenges of micro-randomized trials with categorical treatments and binary proximal outcomes by defining the causal excursion effect, proposing the EMEE-catA estimator, and deriving a robust sample size formula to ensure adequate power for comparing treatment levels, as demonstrated through simulations and a real-world application.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are the captain of a spaceship trying to get your crew to eat their vegetables. You have a fleet of smart devices on their wrists that can send little nudges: a funny cartoon, a stern reminder, or a friendly high-five. But you don't know which nudge works best, or if the crew is even looking at their wrists when you send them. This is the world of mobile health (mHealth): using phones and wearables to help people change their habits, like drinking less alcohol or exercising more.
To figure out what works, scientists use a special kind of experiment called a Micro-Randomized Trial (MRT). Think of it like a massive, high-speed game of "choose your own adventure." Instead of assigning a person to one strategy for the whole month, the computer flips a coin hundreds of times a day. At every single moment a decision is made, it randomly picks a different nudge to send. This happens so fast and so often that researchers can see, almost instantly, if a specific nudge made someone open their app or take a step. The big question they are trying to answer is: "Did this specific message, sent right now, cause the person to do the good thing?"
The tricky part is that sometimes the messages aren't just "yes" or "no." Sometimes you have a whole menu of options: a short message, a long story, a video, or no message at all. And the result you are looking for is often a simple "yes" or "no" (did they open the app? yes/no). Figuring out how to design these experiments so you have enough people to get a clear answer, without wasting time and money, is like trying to bake a cake where you don't know exactly how many eggs you need until you've already started mixing.
This paper is about baking that cake correctly. The authors, Jeremy Lin and Tianchen Qian, are statisticians who built a new recipe for these mobile health experiments. They created a tool to help researchers figure out exactly how many people they need to recruit to make sure their experiment actually works.
Here is the problem they solved: In the past, these tools only worked well when there were two choices (like "send a message" vs. "don't send a message"). But in the real world, scientists often want to test many different types of messages at once. Imagine a menu with a "funny joke," a "serious warning," and a "motivational quote." The old tools couldn't handle this menu properly, especially when the result is a simple "yes/no" (like "did they open the app?").
The authors invented a new way to measure the effect of these multiple choices. They call their new method EMEE-catA. It's like a super-powered calculator that can look at a messy history of thousands of random nudges and tell you, "Hey, the funny joke actually made people open the app 3.6 times more often than doing nothing!" Even better, this calculator is tough. It doesn't break if the researchers guess the wrong details about how people usually behave. It's designed to be robust, meaning it gives the right answer even when the real world is messy and unpredictable.
But the real magic is the sample size formula. This is a mathematical rule that tells a researcher: "If you want to be 80% sure you can tell the difference between your 'funny joke' and your 'serious warning,' you need exactly 1,197 people." Before this paper, researchers were often guessing this number, which could lead to experiments that were too small to find anything or too big and wasteful.
The authors tested their new recipe using data from a real study called Drink Less, which helped people cut down on alcohol. In that study, 349 people were nudged every day for 30 days. When the authors ran their new formula on this data, it confirmed that the study design was solid. They also ran thousands of computer simulations—virtual experiments where they pretended to be the researchers—to see if their formula would fail if the real world didn't follow the rules perfectly.
The results were promising. The formula worked great when the rules were followed. But even when the researchers broke the rules in their simulations (like when people's availability changed in weird ways or when the messages had delayed effects), the formula was surprisingly sturdy. It didn't always give the perfect number, but it usually gave a number that was close enough to ensure the experiment would succeed.
The paper also gives a set of "best practices" for anyone using this tool. It suggests that if you aren't sure about the details, it's better to be a little conservative. For example, if you think your message will work 50% of the time, but you aren't sure, the formula suggests planning for a scenario where it works 40% of the time. This ensures you don't run out of participants halfway through.
In short, this paper provides a new, reliable map for navigating the complex world of mobile health experiments. It allows scientists to design studies with multiple options and simple yes/no results without getting lost in the math. By using this new tool, researchers can be more confident that their experiments will actually tell them which digital nudges are the most effective at helping people change their lives.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.