← Latest papers
📄 medicine

Inverted validity arguments in the evaluation of India’s national professionalism curriculum: a systematic review

This systematic review reveals that the literature evaluating India's national AETCOM professionalism curriculum is methodologically flawed and "inverted," offering large, confident claims of effectiveness based on weak designs and spin, while failing to provide the necessary evidence to determine if the curriculum actually improves real-world medical practice.

Original authors: H S Siddalingaiah, Preethi Ashok Masali

Published 2026-07-29
📖 5 min read🧠 Deep dive

Original authors: H S Siddalingaiah, Preethi Ashok Masali

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a chef trying to figure out if a new, government-mandated recipe for "Perfect Soup" is actually making people healthier. In the world of medical education, this recipe is called AETCOM—a national program in India designed to teach future doctors how to be kind, ethical, and good at talking to patients. But here is the tricky part: how do you know if the soup is actually working? You can't just ask the chef if they think it tastes good, and you can't just look at the ingredients list. You have to watch what happens when the soup is served to real customers in a real restaurant. This is where the science of "evaluation" comes in. It's the process of checking if a lesson actually changes behavior in the real world, not just in a classroom. To do this properly, scientists use a few big ideas: "Validity," which is like checking if your ruler is actually measuring length and not just temperature; "Transfer of Training," which asks if what you learn in the kitchen actually helps you cook better when the restaurant gets busy; and "Spin," which is when someone writes a report that makes a small success sound like a giant victory.

Now, imagine a team of detectives—researchers Siddalingaiah and Masali—deciding to investigate the pile of reports written about this "Perfect Soup" recipe. They didn't just ask, "Does the soup work?" They asked a much stranger question: "Are the reports we have even capable of answering that question?" They looked at dozens of studies from 2015 onwards that tried to test the AETCOM curriculum. What they found was a bit like walking into a room where everyone is shouting that the soup is amazing, but nobody is actually tasting the soup in the restaurant. Instead, they are tasting the raw ingredients in the kitchen, or just asking the chef how they feel about the recipe.

The detectives found that the "size" of the success reported in these studies depended entirely on how the test was set up, not on how good the soup actually was. When teachers graded their own students immediately after class, the results were huge and glowing, like a score of 3.49 on a scale where 1 is already great. But when the researchers looked for the only two studies that actually compared students who got the training against students who didn't, the magic vanished. The score dropped to a tiny 0.12. It's as if the first group of testers were using a magnifying glass that made everything look bigger, while the second group used a plain, honest mirror. The paper shows that the more "permissive" (or loose) the rules of the test were, the bigger the reported success seemed to be.

Even more interesting, the researchers noticed a strange pattern of "spin." They found that in 75% of the reports, the authors made claims that went way beyond what they actually measured. It's like a student getting an 'A' on a spelling test and then writing in their report, "This proves I will win the Nobel Prize for Literature." The studies measured things like "how students felt" or "what they knew on a test," but the reports claimed these studies proved that doctors were now treating patients better or that lawsuits would disappear. In fact, 89% of the reports claimed more than their evidence could support. They were confident about the results in the exact places where they had the least amount of proof.

The paper also points out that almost no one checked the "workplace conditions." Think of it like this: you can teach a driver to park perfectly in an empty parking lot, but if the real city streets are blocked, crowded, and the boss tells them to drive fast, they might crash anyway. The researchers found that only 5% of the studies checked if the doctors' supervisors supported the new behavior, and only 12% checked if the doctors even had a chance to practice what they learned. Without knowing if the "parking lot" was open, you can't blame the driver for not parking well, and you can't blame the driving school either.

So, what is the verdict? The paper isn't saying the AETCOM curriculum is bad or that it doesn't work. It's saying that the current way we are testing it is broken. The evidence base is "inverted": it is loudest and most confident where it is actually the weakest. The authors conclude that we can't tell if the curriculum is a success or a failure yet, simply because we haven't built a test that can actually see if the lessons stick when doctors are dealing with real patients in real hospitals. They suggest we need to stop using the "magnifying glass" tests and start building a system that watches what happens in the real world, checking if the environment allows the new skills to be used. Until we do that, the big claims about saving lives and fixing relationships remain just that—claims, not proven facts.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →