← Latest papers
📄 medicine

Reliability and Feasibility of a Multi-Source, Competency-Based Evaluation System for Surgical Sub-Internship Students: A Single-Institution Pilot Study

This single-institution pilot study demonstrates that a brief, multi-source, competency-based evaluation system for surgical sub-interns achieves strong inter-rater reliability and feasibility with minimal evaluator burden, supporting its potential for high-stakes residency selection.

Original authors: Anthony Onde Morada, Jessica L Becker, Michael J Furey, Christian Hailey Summa, Kristen R Richards, Ajmal Baray, Megan L Gordon, Akshay Patel, Rebecca Michelle Jordan, Joseph P Bannon, Alexandra Falvo

Published 2026-07-01
📖 4 min read☕ Coffee break read

Original authors: Anthony Onde Morada, Jessica L Becker, Michael J Furey, Christian Hailey Summa, Kristen R Richards, Ajmal Baray, Megan L Gordon, Akshay Patel, Rebecca Michelle Jordan, Joseph P Bannon, Alexandra Falvo

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a new chef for a busy restaurant. Usually, you might just ask the head chef, "Did this person seem okay?" and make a decision based on that single opinion. But what if the head chef was too busy to watch the whole time, or what if they only saw the chef on a good day? You'd want a better system to see if the new hire is truly ready to work the dinner rush.

This paper is about testing a new "report card" system for medical students who are doing a final trial run (called a "Sub-Internship") before they become real surgeons. The goal was to see if a new, structured way of grading them works better than the old, informal way.

Here is the breakdown of what they did and found, using simple analogies:

The Problem: The "Guessing Game"

Right now, when medical schools pick residents, they often rely on a single letter of recommendation from one doctor. It's like hiring a pilot based on a single conversation with the airport manager, without checking the flight logs or asking the co-pilots. This can be unfair, inconsistent, and misses the full picture of how a student actually performs.

The Solution: The "360-Degree Scorecard"

The researchers built a new digital report card. Instead of just one person grading the student, they asked three different groups to fill it out:

  1. The Bosses: Attending surgeons (like the head chefs).
  2. The Peers: Other residents (like the sous-chefs who work side-by-side with the student).
  3. The Support Staff: Nurses and techs (like the dishwashers and waiters who see how the student treats the whole team).

The report card didn't just ask, "Is this student good?" (which is vague). Instead, it asked specific questions about six key areas, like "Can they handle a knife?" (Technical Skills), "Do they know the recipe?" (Clinical Knowledge), and "Are they polite under pressure?" (Professionalism). They used a simple 1-to-5 scale, similar to rating a movie.

The Test: A 5-Month Pilot

They tried this new system at one hospital for five months. Eight students went through their rotations, and the staff filled out 24 of these digital forms. The forms were designed to be quick, like ordering a coffee on a phone app, taking only about 3 minutes to complete.

The Results: Does the New System Work?

The researchers wanted to know two things:

  1. Reliability: If different people graded the same student, would they agree? (If one person says "5 stars" and another says "1 star," the system is broken).
  2. Feasibility: Was it too much work for the staff?

Here is what they found:

  • It was fast and easy: The staff finished the forms in about 3 minutes. It didn't slow down the hospital.
  • The "Big Picture" grades were super consistent: When it came to the most important questions—"Is this student ready to be a surgeon?" and "How good are their technical skills?"—the different evaluators (doctors, residents, and staff) agreed almost perfectly. It's like if you asked five different food critics to rate a dish, and they all gave it a 5-star review.
  • The "Knowledge" grades were okay: When it came to grading how much the student knew about medicine, the evaluators agreed moderately well.
  • The "Professionalism" grades were tricky: Interestingly, the scores for "being a good person" were so high for everyone that the system couldn't tell much difference between them. It's like a class where every student gets an "A" for attendance; the grades are all the same, so the math looks weird, but it just means everyone was on time! The system still worked because it had a special "red flag" section for serious problems.
  • No Bias: It didn't matter if a boss or a peer filled out the form; they gave similar scores. This means the system is fair regardless of who is doing the grading.
  • It caught the problems: The system successfully flagged students who were struggling. About 17% of the evaluations had a "concern" note, and the system made sure these were written down clearly so the hiring committee could see exactly what went wrong.

The Bottom Line

The researchers concluded that this new, multi-person report card is a solid, reliable tool. It gives a clear, consistent picture of whether a student is ready for the job, without taking up too much time from the busy medical staff.

They noted that this was just a small test (a "pilot") with a small group of students. To make it perfect, they say they need to try it at more hospitals and with more students to fine-tune the numbers. But for now, it proves that a structured, team-based way of grading is possible and works well.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →