External validation and comparison of PREVENT and SCORE2 atherosclerotic cardiovascular risk scores in the MASHAD cohort study
This study externally validated the PREVENT and SCORE2 cardiovascular risk models in the Iranian MASHAD cohort, finding that while both models demonstrated acceptable discrimination, they significantly underestimated absolute risk and required local recalibration to provide accurate predictions for this Middle Eastern population.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have two very popular weather forecasters, PREVENT and SCORE2. These forecasters are famous in Europe and the United States for predicting the chance of a "storm" (a heart attack or stroke) hitting a person over the next 10 years. They use a standard recipe: they look at your age, blood pressure, cholesterol, and whether you smoke to give you a percentage chance of getting hit by the storm.
The researchers in this study wanted to see if these famous forecasters would work correctly in Mashhad, Iran. They treated the local population like a new, untested weather zone.
Here is what they found, explained simply:
1. The Problem: The "Foreign" Forecasters Were Too Optimistic
When the researchers used the original European and American formulas on the Iranian people, the forecasters made a big mistake: they were too optimistic.
- The Analogy: Imagine a weatherman in London telling a farmer in a desert, "There is a 1% chance of rain today." But in the desert, it actually rains 30% of the time. The farmer would get soaked because the forecast was wrong for his specific location.
- The Reality: Both PREVENT and SCORE2 told the Iranian people their risk of a heart event was very low. However, when the researchers looked at the actual data over 10 years, many more people had heart events than the models predicted. The models were "underestimating" the danger.
2. The Solution: "Local Tuning" (Recalibration)
The researchers didn't throw the models away. Instead, they performed a "local tune-up."
- The Analogy: Think of the models as a radio station. The music (the way they rank people from low risk to high risk) was good, but the volume was too quiet. They didn't need to change the song; they just needed to turn up the volume knob to match the local audience.
- The Action: They adjusted the math slightly to match the actual number of heart events happening in Mashhad. They kept the "ranking" system the same but shifted the numbers so the percentages reflected reality.
3. The Results: Who Was Better?
After the "tune-up," both models worked much better, but one had a slight edge.
- The Ranking Ability (Discrimination): Both models were good at sorting people. They could tell who was more likely to have an event than someone else. PREVENT was slightly better at this sorting than SCORE2, like a slightly sharper pair of glasses.
- The Accuracy (Calibration): Before the tune-up, the models were wrong about the amount of risk. After the tune-up, the predicted risk matched the actual risk almost perfectly.
- PREVENT was a bit more accurate overall. This is likely because PREVENT looks at more details, like kidney function and diabetes, which are very common in this specific group of people.
- SCORE2 is a simpler model (it doesn't look at diabetes), but it still worked well once it was tuned to the local "volume."
4. Why This Matters (The "Net Benefit")
The study used a concept called "Net Benefit" to see if the tuned models helped doctors make better decisions.
- The Analogy: Imagine a security guard checking bags.
- Original Model: The guard was too relaxed, letting many dangerous bags through (missing high-risk people).
- Tuned Model: The guard became more alert. He stopped more dangerous bags (catching more high-risk people), but he also stopped a few harmless bags by mistake (lowering "specificity").
- The Finding: The researchers found that the "tuned" models were worth it. Even though they flagged a few extra people who turned out to be safe, they successfully caught many more people who were actually at risk. In a health context, it is better to catch a potential heart attack early than to miss it.
Summary
The study concludes that you cannot simply copy-paste a risk calculator from Europe or America and expect it to work perfectly in Iran. The "recipe" for risk is different in different places.
However, the good news is that these tools can work if you "calibrate" them to the local population. Once the researchers adjusted the numbers to fit the reality of Mashhad, both PREVENT and SCORE2 became reliable tools for predicting heart risks, with PREVENT having a slight advantage because it considers more health factors.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.