← Latest papers
📊 statistics

How can the use of different modes of survey data collection introduce bias? A simple introduction to mode effects using directed acyclic graphs (DAGs)

This paper utilizes directed acyclic graphs (DAGs) to explain how mixed-mode survey designs introduce bias through "mode effects" and "mode selection," demonstrating that naive statistical adjustments like conditioning can inadvertently create collider bias while advocating for quantitative bias analysis as a more robust solution.

Original authors: Georgia D Tomova, Richard J Silverwood, Peter WG Tennant, Liam Wright

Published 2026-01-26
📖 5 min read🧠 Deep dive

Original authors: Georgia D Tomova, Richard J Silverwood, Peter WG Tennant, Liam Wright

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to figure out how much people really like a specific type of food. To do this, you ask them to fill out a survey. But here's the twist: you let them choose how they answer. Some fill out a paper form, some answer on a website, and some talk to a person over the phone.

This paper argues that giving people choices on how to answer can actually mess up your results in two sneaky ways. The authors use a tool called "Directed Acyclic Graphs" (DAGs)—which are basically simple flowcharts that show cause-and-effect relationships—to explain why this happens and how to avoid getting fooled.

Here is the breakdown of the problem and the solutions, using everyday analogies.

The Two Big Problems

1. The "Mode Effect" (The Mask Problem)
Imagine you are asking people, "Do you eat vegetables?"

  • Scenario A: You ask them face-to-face with a friendly interviewer. People might say "Yes!" because they don't want to look bad in front of a human (they are being socially polite).
  • Scenario B: You ask them on a private computer screen. They might type "No" because they feel safe and honest.

The truth about their eating habits hasn't changed, but the answer they give has changed just because of the "mode" (the method) they used. This is called a Mode Effect. It's like wearing a mask that changes your face depending on who is looking at you.

2. The "Mode Selection" (The Self-Selection Problem)
Now, imagine who chooses which method.

  • Tech-savvy young people might prefer the website.
  • Older people or those without internet might prefer the phone call.

This means the group of people answering on the website is different from the group answering on the phone. This is Mode Selection. It's like if you only asked people who own a car about their driving habits; your sample is already biased because non-car owners aren't there.

The Trap: Why "Fixing" It Can Make It Worse

The paper's main warning is about what happens when researchers try to fix the "Mask Problem" without thinking about the "Self-Selection Problem."

The Naive Fix:
If you notice that people on the website say "No" to vegetables and people on the phone say "Yes," a researcher might think, "Okay, I'll just separate the data and compare them." Or, "I'll just add a 'mode' box to my math equation to cancel out the difference."

The Trap (Collider Bias):
The authors say this is dangerous if the type of person (e.g., their age or education) influenced both which method they chose AND what they answered.

Think of it like a bouncer at a club:

  • The "Bouncer" is the Survey Mode (Website vs. Phone).
  • The "VIPs" are the people who chose the website (maybe they are younger).
  • The "Regulars" are the people who chose the phone (maybe they are older).

If you try to "fix" the data by forcing the VIPs and Regulars to be compared directly, you might accidentally create a fake link between things that aren't related.

The Paper's Warning:
If you try to "control for" the survey mode (by separating the data or adjusting the math) when the people themselves chose the mode based on their characteristics, you might accidentally create a fake connection between your variables. It's like looking at a group of people who all own red cars and all like pizza, and concluding that "Red cars cause a love for pizza," when really, the people who like both things just happened to choose the red car club.

The paper uses flowcharts (DAGs) to show that:

  • If the survey mode is just a "mask" (affecting the answer but not chosen by the person's traits), you can fix it by separating the modes.
  • If the survey mode is a "bouncer" (chosen by the person's traits), trying to fix it by separating the modes actually creates new errors (called Collider Bias).

What Should Researchers Do?

The paper suggests three ways to handle this, depending on the situation:

  1. Don't touch it (Conditioning): If you are sure the people didn't choose the mode based on their traits, you can simply separate the data by mode. But if they did choose, don't do this!
  2. Guess and Fill (Imputation): You could try to guess what the "website people" would have said if they used the "phone method" and fill in the blanks. The paper warns this is risky because it assumes the missing data is random, which it usually isn't. It's like trying to guess a friend's answer to a question they didn't ask, assuming they would have answered the same as everyone else.
  3. The "What If" Game (Quantitative Bias Analysis): This is the method the authors recommend most. Instead of trying to magically fix the data, you admit the data is messy. You say, "Okay, let's assume the mode effect was this big. How would that change our results?" Then you try a different size. If your main conclusion stays the same no matter how big you assume the error is, you can be confident in your findings.

The Bottom Line

When you mix different ways of collecting survey data (like phone, web, and paper), you have to be careful.

  • The Method matters: It changes how people answer (The Mask).
  • The People matter: Different people pick different methods (The Bouncer).

If you try to fix the "Mask" problem without realizing the "Bouncer" is there, you might accidentally invent a fake story. The best approach is to map out your assumptions (using those flowcharts), realize you can't always fix the data perfectly, and instead test how much your results would change if the errors were bigger or smaller.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →