Large Language Models in Healthcare Simulation Education: A Bibliometric Analysis with AI-Assisted Screening
This bibliometric analysis reveals that while the field of large language models in healthcare simulation education is experiencing rapid growth, it remains heavily concentrated on communication and decision-making skills using proprietary models like ChatGPT, leaving critical gaps in team-based training, open-source model evaluation, and specialty-specific applications, all while demonstrating the feasibility of using multiple independent AI agents for reliable, scalable literature screening.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Imagine the world of medical training as a massive, bustling construction site. For decades, they've been building skills using blueprints, practice dummies, and role-playing actors. Now, a new, incredibly powerful tool has arrived: Large Language Models (LLMs) like ChatGPT. These are AI systems that can read, write, and talk like humans.
This paper is like a satellite map taken by a team of researchers to see exactly how this new tool is being used on that construction site. They didn't just look at a few buildings; they scanned over 100,000 potential blueprints (research papers) and, using a clever mix of keyword filters and 83 independent AI "inspectors," they narrowed it down to the 551 most relevant projects.
Here is what their map reveals, explained in simple terms:
1. The Explosion: A Firework, Not a Slow Burn
Before late 2022, there were almost no papers about using these AI tools for medical simulation. It was a quiet field. Then, ChatGPT was released, and it was like someone lit a firework.
- The Analogy: Imagine a pond that was still for years, and then suddenly, a giant rock is dropped in. The water (research) didn't just ripple; it exploded.
- The Fact: The number of papers grew by 109% every year. Almost all the research (99%) happened after ChatGPT arrived. It's a brand-new field growing at a breakneck speed.
2. The "One-Tool" Problem
If you walked into a workshop and saw 100 tools, you'd expect a variety. But in this research, almost half of the workers are using only one specific brand of hammer (ChatGPT).
- The Analogy: It's like a chef's competition where 46% of the contestants are only allowed to use a specific brand of knife, and almost no one is using the cheaper, open-source knives that anyone can build themselves.
- The Fact: ChatGPT appears in nearly half the papers. Open-source models (free, community-built tools) are almost invisible in the research. This means the whole field is relying on a single company's product, which could be risky if that product changes or costs too much.
3. The "Solo Act" vs. The "Team Sport"
Medical training isn't just about talking to a patient; it's about working as a team during a crisis. The researchers found that the AI is mostly being used for one-on-one conversations.
- The Analogy: Imagine a sports team practicing. Right now, the AI is great at helping a player practice their penalty kicks (communication) or strategy calls (decision-making). But it is almost completely absent when it comes to practicing team huddles, leadership, or handling a chaotic emergency (teamwork, leadership, situational awareness).
- The Fact: The research is heavily focused on "Communication" and "Decision-making." Critical skills like "Teamwork," "Leadership," and "Crisis Management" are barely studied. The AI is good at talking, but it hasn't been tested much on how to lead a team through a storm.
4. The "Virtual Patient" Takeover
The most common way this AI is being used is as a chatbot patient.
- The Analogy: Instead of hiring a human actor to pretend to be a sick person, researchers are using the AI to chat with students. It's like having a digital actor that never gets tired, never forgets its lines, and is available 24/7.
- The Fact: "Virtual Patient Chatbots" are the number one way AI is used in these studies. However, the researchers noted that we don't yet know if talking to a robot is exactly the same as talking to a real human actor, especially when it comes to reading emotions.
5. The "Urology" Blind Spot
The researchers specifically looked at Urology (a branch of surgery dealing with the urinary system), which is famous for having very structured, high-quality training camps (boot camps).
- The Analogy: It's like finding a massive, high-tech gym that is perfect for training, but the new AI equipment has been completely ignored there.
- The Fact: Out of 551 papers, only 6 were about Urology. Even in those 6, no one had tested the AI in the specific "boot camp" style training that urologists use. It's a huge missed opportunity because these training camps are perfect for testing new tech.
6. The "AI Inspector" Method
How did they find these 551 papers out of 100,000? They didn't just ask one human to read them all (that would take forever).
- The Analogy: They built a robot assembly line. First, they used a simple filter to remove the junk. Then, they hired 83 different AI "inspectors" to check the remaining papers. Each inspector worked alone.
- The Fact: These AI inspectors agreed with each other (and with human experts) almost perfectly. This proves that we can use AI to help us research AI, making the process faster and more reliable.
The Bottom Line
The map shows a field that is growing incredibly fast but is unbalanced.
- It's too fast (we don't have enough time to see if the methods actually work long-term).
- It's too narrow (everyone is using the same AI tool).
- It's too solo (it focuses on talking, not on team crisis management).
- It's missing key areas (like Urology training).
The paper concludes that while the technology is exciting and the growth is amazing, we need to stop just looking at the "one-on-one" chatbots and start figuring out how to use AI to help medical teams work together, lead better, and handle emergencies.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.