← Latest papers
📄 medicine

Evaluating the Effectiveness of an Information System-Assisted Competency Assessment for Residents: A Single-Center Study in China

This single-center study in China demonstrates that an information system-assisted competency assessment is feasible for large-scale residency training but reveals significant faculty leniency and resident self-assessment inflation, underscoring the critical need for rater calibration to ensure assessment accuracy.

Original authors: Lan Li, Qian Zhao, Chunyan Cheng

Published 2026-08-28
📖 4 min read☕ Coffee break read

Original authors: Lan Li, Qian Zhao, Chunyan Cheng

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of medical training, the goal has long been to ensure that every doctor graduating from a residency program is not just knowledgeable, but truly competent. For decades, this training relied heavily on how much time a student spent in a hospital and whether they passed written exams. However, the modern approach, known as competency-based medical education, shifts the focus to what a doctor can actually do. It breaks down the complex job of being a physician into specific skills, such as managing a patient's care, communicating with a family, or teaching a junior colleague. The challenge for hospitals is how to measure these skills fairly and accurately across hundreds of trainees. If the people grading the residents are too easy, or if the residents think they are better than they are, the system fails to identify who needs more help. Without a clear picture of a resident's true abilities, a hospital cannot guarantee that its doctors are ready to treat patients safely.

To solve this, researchers at West China Hospital in Sichuan, China, turned to technology to evaluate nearly 1,800 residents across 24 different medical specialties. They built a digital system to gather data on how these trainees were performing over several months. Instead of relying on a single opinion, the system collected three different types of feedback for each resident: a score from their direct faculty mentor, a score from a special committee of experts known as the Clinical Competency Committee, and a score from the residents themselves. The researchers wanted to see if this digital approach could work on such a large scale and, more importantly, to see if the different groups of evaluators agreed with one another. They were looking for the truth behind the numbers: did the scores rise as the residents gained experience, and did the people grading them see the same reality?

The results showed that the digital system worked well, successfully organizing a massive amount of data that would have been impossible to manage with paper forms. When the researchers looked at the scores given by the expert committee, a clear pattern emerged. As residents moved from their first year of training to their third, their scores improved significantly in every area of competence. This confirmed that the training was working and that the residents were indeed growing more skilled over time. However, one specific skill stood out as consistently weak. Across all three years of training, the residents received the lowest marks in teaching skills. Even the most experienced trainees struggled in this area, suggesting that while they were becoming excellent doctors, they were not being taught how to teach others.

The study also revealed a significant disconnect between who was doing the grading and how they graded. The faculty mentors, who work closely with the residents every day, gave scores that were generally in the same ballpark as the expert committee, but they were consistently too kind. In five out of the six skill areas, the mentors rated the residents higher than the committee did. This "leniency" meant that the mentors were not spotting the same weaknesses that the committee saw. The gap was even wider when it came to the residents' own views of their performance. When the residents rated themselves, they gave themselves scores that were dramatically higher than the expert committee's ratings. On average, a resident rated themselves about four points higher than the committee did. This suggests that the residents were not just slightly overconfident; they were largely unaware of the gap between their self-perception and their actual performance.

The researchers found that the digital system was a powerful tool for uncovering these hidden patterns. It allowed them to see that while the training program was successful in building clinical skills, it was failing to develop teaching abilities. It also highlighted a serious need for better training for the people doing the grading. The faculty mentors needed to learn how to be more objective, and the residents needed help in learning how to judge their own work accurately. The study did not claim to have solved these problems, but it proved that a computerized system could provide the clear, honest data needed to fix them. By moving away from fragmented paper records to a unified digital platform, the hospital created a way to measure progress that was both large-scale and detailed. This approach offers a path forward for medical education, where data can be used to ensure that every resident is truly ready to care for patients, rather than just having completed the required time.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →