MedSimAI: Simulation and Formative Feedback Generation to Enhance Deliberate Practice in Medical Education
Original authors: Yann Hicke, Jadon Geathers, Kellen Vu, Justin Sewell, Claire Cardie, Jaideep Talwalkar, Dennis Shung, Anyanate Gwendolyne Jack, Susannah Cornes, Mackenzi Preston, Rene Kizilcec
Original authors: Yann Hicke, Jadon Geathers, Kellen Vu, Justin Sewell, Claire Cardie, Jaideep Talwalkar, Dennis Shung, Anyanate Gwendolyne Jack, Susannah Cornes, Mackenzi Preston, Rene Kizilcec
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Technical Summary: MedSimAI
Problem Statement
Medical education faces significant challenges in scaling clinical skills training, particularly in history-taking and communication. Traditional simulation-based learning using Standardized Patients (SPs) is resource-intensive, expensive (e.g., high-fidelity mannequins), and often variable in feedback quality. While AI-powered simulations offer a potential solution, existing tools suffer from critical limitations: they often address only narrow competencies, lack comprehensive assessment frameworks, provide limited evidence of downstream clinical impact, and rarely integrate Self-Regulated Learning (SRL) principles. Furthermore, current implementations often fail to balance scalability with the need for realistic, culturally competent, and clinically authentic interactions.
Methodology
The authors developed MedSimAI, an AI-powered simulation platform, through a multi-phase co-design process with subject matter experts (SMEs) from three medical institutions. The system architecture and evaluation involved the following components:
1. System Architecture
- AI Standardized Patient (AI-SP): The core engine uses Large Language Models (LLMs), specifically GPT-4o for text and OpenAI's Realtime API for voice. Instructors define patient parameters (demographics, medical history, emotional state) via structured templates. System prompts guide the LLM to maintain clinical authenticity, avoid jargon, and exhibit consistent personality traits.
- Assessment and Feedback Framework: The system processes encounters using specialized evaluation prompts. It supports:
- Rubric-based scoring: Utilizing frameworks like the 28-item Master Interview Rating Scale (MIRS) on a 1–5 scale.
- Checklist-based scoring: Binary assessment of required history-taking elements (e.g., red-flag symptoms).
- Feedback Generation: Automated, quote-grounded feedback is generated within 2–3 minutes post-encounter, highlighting strengths, gaps, and specific excerpts from the dialogue.
- Learning Hub (SRL Interface): A dashboard integrating SRL strategies using clinical metaphors (e.g., "appointments" for scheduling practice, "patient charts" for review). Features include strategic planning, competency dashboards mapped to Kalamazoo essential elements, and reflection tools.
2. Study Design and Data Collection
- Deployment: A multi-institutional deployment across three U.S. medical schools (Institution A: private research-intensive; Institution B: large public; Institution C: private Ivy League-affiliated).
- Participants: 410 unique learners generated 1,024 completed simulated encounters.
- Data Sources:
- Platform Telemetry: Session duration, modality (voice/text), dialogue turns, automated scores (MIRS, checklists), and SRL artifacts (840 learner reflections).
- Surveys: Baseline and exit surveys (n=19 at Institution B) measuring confidence, anxiety, and perceived value.
- Performance Metrics: Comparison of Objective Structured Clinical Examination (OSCE) history-taking scores at Institution A (quasi-experimental) and Demonstration of Clinical Skills (DOCS) scores at Institution B.
- Validation: A corpus of 104 OSCE transcripts from Institution C, previously rated by human evaluators, was re-scored by MedSimAI to benchmark automated scoring accuracy.
3. Analysis
- Quantitative: Mixed-effects models (random intercepts for learners) analyzed the association between usage patterns (dose, modality, dialogue turns) and performance outcomes. Kruskal-Wallis tests assessed institutional and case differences.
- Qualitative: Inductive thematic analysis was performed on 840 learner reflections and open-ended survey responses.
- Validation Metrics: Automated scoring was evaluated against human raters using Exact Accuracy, Off-by-One Accuracy, and Thresholded Accuracy (distinguishing proficiency).
Key Contributions
- Co-Designed Platform: An AI-SP platform developed through iterative collaboration with medical education experts, featuring instructor-authored cases and configurable realism.
- Flexible Assessment Layer: A system combining multi-rubric scoring (including MIRS) and dynamic checklists to generate quote-grounded formative feedback and SRL-oriented dashboards.
- Multi-Site Evaluation: A comprehensive evaluation involving 1,024 encounters and 410 learners, utilizing mixed-effects models to probe institution and case effects, alongside an inductive thematic analysis of learner reflections.
- Evidence of Validity and Impact:
- Quasi-experimental results: A significant improvement in OSCE history-taking scores at one institution.
- Automated Scoring Validation: Independent benchmarking showing high accuracy in identifying proficiency thresholds.
Results
Engagement and Usage
- Overall Engagement: 59.5% of learners engaged in repeated practice (completing ≥2 cases).
- Institutional Variation: Engagement varied significantly. Institution B showed the highest intensity (3.48 cases/student, 73.1% repeated practice), while Institution C showed minimal engagement (1.35 cases/student).
- High-Engagement Subgroup: 2.9% of students (12 individuals) completed 8+ cases, accounting for 12.1% of total platform usage.
Clinical Performance
- Institution A (Curriculum-Embedded): A quasi-experimental comparison showed a statistically significant improvement in OSCE history-taking scores for the cohort with MedSimAI access compared to a prior cohort without access. Scores rose from a mean of 82.8 to 88.8 (p<0.001, effect size d=0.75).
- Institution B (Voluntary Pilot): No statistically significant differences were found in DOCS scores between voluntary participants and non-participants (p>0.30), suggesting that self-directed use without curricular integration may not yield measurable exam gains.
Automated Scoring Validation
- Accuracy: The system achieved 32.5% exact accuracy, 64.1% off-by-one accuracy, and 87.0% thresholded accuracy in distinguishing proficient from under-performing learners on the MIRS.
- Utility: While exact score matching is imperfect, the system is deemed adequate for formative triage to flag students needing remediation.
Learner Reflections and Experience
- Thematic Analysis: Analysis of 840 reflections revealed top themes: missed/forgotten items (26.3%), organization/flow (23.0%), and review of systems (ROS) technique (22.4%).
- Perceived Value: Survey respondents reported reduced exam anxiety (M=5.38) and improved preparation (M=5.38).
- Challenges: Students cited technical glitches, a perceived lack of realism in voice interactions, and a tension between checklist completeness and conversational flow.
Significance and Claims
The paper positions MedSimAI as a scalable formative platform for history-taking and communication training. The authors claim that:
- Scalability: AI-simulated patients can overcome the resource barriers of traditional SPs, providing immediate, structured feedback to large cohorts.
- Curricular Integration is Critical: The disparity in results between Institution A (curriculum-embedded) and Institution B (voluntary) suggests that technical deployment alone is insufficient; successful implementation requires thoughtful curricular integration and structured practice schedules.
- Formative Utility: The high thresholded accuracy (87%) supports the use of automated scoring for formative assessment and triage, though human oversight remains necessary for summative evaluation.
- SRL Limitations: Despite the inclusion of SRL scaffolds, engagement with these features was low without explicit curricular embedding, indicating that tools alone do not guarantee behavior change.
- Context Matters: Institutional context, case difficulty, and implementation choices (e.g., integration vs. optional use) significantly influence outcomes, often outweighing raw practice volume in predicting performance gains.
The authors conclude that while MedSimAI shows promise for enhancing clinical skills training, its effectiveness is contingent on how it is integrated into the curriculum and the specific educational context, rather than the technology alone.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.
Get the best AI papers every week.
Trusted by researchers at Stanford, Cambridge, and the French Academy of Sciences.
Check your inbox to confirm your subscription.
Something went wrong. Try again?
No spam, unsubscribe anytime.