Development and Validation of Performance Assessment Scales for Core Cardiac Surgical Procedures: A National Simulation-Based Study
This national simulation-based study successfully developed and validated four procedure-specific Performance Assessment of Surgical Skills (PASS) scales for median sternotomy, cardiopulmonary bypass, surgical aortic valve replacement, and internal mammary artery harvesting, demonstrating their high reliability and validity as objective tools for competency-based cardiac surgical training.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Becoming a heart surgeon is a journey that demands years of intense study and practice. It is a field where the margin for error is virtually non-existent, and the ability to perform delicate maneuvers on a beating organ is the difference between life and death. For decades, the path to mastery relied heavily on watching a mentor and then practicing on patients, a method that has slowly evolved to include high-fidelity simulation. In these simulated environments, trainees can refine their skills on realistic models without risk. However, a critical piece of the puzzle has often been missing: a way to measure exactly how well a trainee is doing. While instructors can offer general feedback, the medical community has long sought a standardized, objective tool that breaks down complex surgeries into specific, observable steps. Without such a tool, it is difficult to know if a student has truly mastered a technique or is simply repeating motions without understanding the underlying precision required.
A team of researchers in France set out to solve this problem by creating and testing a new set of scoring systems for the most fundamental procedures in cardiac surgery. They focused on four distinct but connected tasks: cutting through the breastbone to access the heart, connecting the patient to a heart-lung machine, replacing a diseased aortic valve, and harvesting a vital artery from the chest wall to use as a graft. The researchers called these new tools PASS scales, short for Performance Assessment of Surgical Skills. To build them, they did not simply guess what steps were important; they gathered a panel of thirteen national experts in cardiac surgery. These experts reviewed existing guidelines and debated every single step of the procedures until they reached a consensus on what constituted a correct action. They then tested these lists during simulation sessions, refining the language until every item was clear and unambiguous. The final result was a set of checklists where each step could be marked as not done, partially done, or completely done, turning a complex surgical performance into a series of measurable facts.
To see if these checklists actually worked, the researchers organized a massive training event involving forty-three surgical residents. These trainees, who were in the middle of their specialized training, performed a complete simulated heart valve replacement on a realistic model of a human body. This model was not a plastic toy; it was a donated human body that had been reconnected to a circulation system and a breathing machine, allowing the surgeons to work under conditions that felt and looked almost exactly like a real operation. As the residents worked, two independent expert surgeons watched them closely, each filling out the new scoring sheets without knowing what the other had written. The goal was to see if the scales could reliably distinguish between different levels of skill and if two different experts would agree on the score for the same performance.
The results showed that the new scales were highly effective. When the researchers analyzed the scores, they found that the items on each checklist hung together logically, measuring the same underlying skill set. For the most complex procedure, the full valve replacement, the consistency between the different steps was exceptionally high. More importantly, the two independent observers almost always agreed on the scores they gave. There was no significant difference between what one expert thought and what the other thought, and the statistical measures of their agreement were among the highest possible. This means that the tool is not dependent on the personal opinion of a single teacher; it provides a stable, objective standard that different experts can use to judge performance in the same way. The researchers also asked the trainees to grade themselves, and while the students were generally honest, they tended to be harsher on their own work than the experts were, highlighting the difficulty of self-assessment in such a high-stakes field.
The study did have its boundaries. The entire validation process took place within the controlled environment of a simulation laboratory, using a specific type of human model. The researchers noted that while the tools performed perfectly in this setting, their ability to measure performance in a real operating room with a living patient still needs to be confirmed. Additionally, the group of experts who helped design the scales came from a single national society, which might limit how well the tools apply to surgeons trained in different parts of the world. Despite these limitations, the findings offer a significant step forward. The researchers concluded that these new scales are valid and reliable instruments that can be used to objectively track the progress of heart surgeons. By providing a clear, step-by-step map of what good performance looks like, these tools can help trainees identify their specific weaknesses and allow educators to guide them more effectively toward the high level of competence required to operate on the human heart.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.