Discovering the critical number of respondents to validate an item in a questionnaire: The Binomial Cut-level Content Validity proposal
This paper proposes a refined Binomial Cut-level Content Validity method to determine the critical number of respondents needed for questionnaire item validation, addressing paradoxes and probability inaccuracies inherent in traditional Content Validation Ratio (CVR) approaches.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a chef trying to create the perfect new recipe for a soup. You have a list of 20 potential ingredients (like carrots, spices, and mystery powders). Before you serve this soup to a whole city, you need to ask a group of expert food critics: "Is this ingredient essential for the soup, or is it just junk?"
This is exactly what researchers do when they design a questionnaire. They have a list of questions (ingredients) and need to know which ones are "essential" to ask and which ones should be thrown out.
For decades, the standard way to do this has been a method called CVR (Content Validity Ratio). Think of CVR as an old, slightly broken scale that the critics use to weigh the ingredients. It has been used for 50 years, but the authors of this paper argue that the scale is a bit wobbly and sometimes gives confusing results.
Here is a simple breakdown of their new idea, which they call BCV (Binomial Cut-level Validity).
The Problem with the Old Scale (CVR)
The old method (CVR) has three main flaws, which the authors compare to these scenarios:
The "Coin Flip" Mistake:
Imagine the critics have three choices for every ingredient:- A) Keep it (Essential)
- B) It's okay, but not needed (Important but not essential)
- C) Throw it away (Unnecessary)
The old method assumes that if a critic picks an answer randomly, they have a 50/50 chance of picking "Keep it" or "Throw it away." But that's like flipping a coin! In reality, there are three options. If you pick randomly, you only have a 1-in-3 chance of picking "Keep it." The old math was using the wrong odds, making the results slightly inaccurate.
The "One-Sided" Blind Spot:
The old method only checks if enough people said "Keep it." It never checks if enough people said "Throw it away."- The Analogy: Imagine a jury that only asks, "Is the defendant guilty?" They never ask, "Is the defendant innocent?"
- The Risk: You might end up keeping an ingredient because 11 people said "Keep it," but 12 people actually said "Throw it away." The old method misses this conflict.
The "Paradox" Trap:
Sometimes, a group of critics is so confused that they say an ingredient is both "Essential" AND "Unnecessary" at the same time.- The Analogy: It's like a group of people voting on whether a bridge is safe. 60% say "It's safe!" and 60% say "It's unsafe!" (because some people voted twice or the group is split down the middle).
- The old method doesn't have a way to catch this "paradox." It just gives a result, even if the data is contradictory.
The New Solution: The "Double-Check" Scale (BCV)
The authors propose a new, smarter scale called BCV. Here is how it works in plain English:
1. Fixing the Odds (The Dice Roll)
Instead of assuming a coin flip (50/50), the new method uses a three-sided die (or a four-sided die if you add an "I don't know" option). It calculates the math based on the actual number of choices the critics have. This makes the "Essential" threshold much more precise.
2. The "Double-Check" Validation
The new method asks two questions for every ingredient:
- Question A: Did enough people say "Keep it" to prove it's essential?
- Question B: Did enough people say "Throw it away" to prove it's junk?
3. Catching the Paradoxes
By answering both questions, the new method creates four possible outcomes, like a traffic light system:
- 🟢 Green Light (Keep It): Enough people said "Essential," and not enough said "Unnecessary." -> Keep the question.
- 🔴 Red Light (Toss It): Enough people said "Unnecessary," and not enough said "Essential." -> Delete the question.
- ⚠️ Yellow Light (The "Paradox" - Type A): Too many people said "Essential," AND too many people said "Unnecessary."
- What this means: The group is confused or divided. Maybe the question is worded poorly, or the experts disagree wildly. The system flags this as a "Strong Paradox" so you can investigate before making a decision.
- ⚪ White Light (The "Paradox" - Type B): Not enough people said "Essential," AND not enough people said "Unnecessary."
- What this means: Everyone is on the fence. The question is too vague or boring. The system flags this as a "Weak Paradox" (or "No Consensus") so you can rewrite it.
Why Does This Matter?
Think of the old method as a crude filter that lets some bad ingredients through and misses some good ones.
The new BCV method is like a smart, high-tech scanner. It doesn't just count votes; it checks for contradictions. It tells you:
- "This is definitely good."
- "This is definitely bad."
- "Wait, the data is fighting itself! We need to look closer."
The Bottom Line
The authors aren't just tweaking a formula; they are upgrading the entire process of deciding what questions to ask. By using better math (Binomial distribution instead of Normal approximation) and checking for "paradoxes," they ensure that the final questionnaire is not just a list of questions, but a reliable, clear, and logical tool that actually measures what it's supposed to measure.
It's the difference between guessing if a soup tastes good and actually tasting every single ingredient to make sure the recipe is perfect.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.