Polish-English medical knowledge transfer: A new benchmark and results
This paper introduces a new Polish-English medical benchmark derived from over 24,000 licensing and specialization exam questions to evaluate state-of-the-art LLMs against human performance, revealing that while top models approach human levels, significant challenges remain in cross-lingual translation and domain-specific understanding.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Artificial intelligence has learned to read and write with a fluency that often rivals human conversation. These systems, known as large language models, are trained on vast amounts of text from the internet, books, and articles, allowing them to answer questions, summarize stories, and solve problems. However, their knowledge is not evenly distributed. Because the majority of the data used to teach them comes from English-speaking sources, these models tend to understand English far better than other languages. This creates a significant gap when the technology is applied to specialized fields like medicine, where guidelines, diseases, and treatments can vary from country to country. If a doctor in Poland relies on an AI trained primarily on American or British medical data, the advice might not fit the local reality, potentially leading to errors in diagnosis or treatment.
Researchers in Poland set out to measure exactly how well these artificial minds perform when asked to think like a Polish doctor. They created a new testing ground based on the actual exams that medical students and specialists must pass to practice in their country. These exams are rigorous, covering everything from basic anatomy to complex surgical procedures, and they are written in Polish. The team gathered over 22,000 questions from these real-world tests, including some that had been professionally translated into English. By feeding these questions to various AI models, they could see if the machines could pass the same hurdles as human students, and whether the language of the question changed the outcome.
The results revealed a clear picture of where these models stand. The most advanced artificial intelligence systems, particularly one known as GPT-4o, performed remarkably well, answering nearly 90 percent of the questions correctly on the general medical licensing exam. This score is high enough that the model would pass the test alongside a typical medical student. However, the story changes when the questions become more difficult or specialized. When the researchers tested the models on exams for dental students and on advanced tests for medical specialists, the scores dropped significantly. Even the best-performing AI struggled with the dental exam, and on the specialist exams, it failed to outperform the average human doctor in more than 30 percent of the medical fields tested. In some specific areas, such as audiology and ear, nose, and throat medicine, the AI performed worse than every single human who took the test that year.
A crucial part of the study was comparing how the models handled the same questions in Polish versus English. The researchers found that for most models, the language made a huge difference. When a question was asked in English, the AI was much more likely to get it right than when the exact same question was asked in Polish. This suggests that the models are not truly understanding the medical concepts in a universal way; instead, they are relying on patterns they learned from English texts. As the models became larger and more powerful, this gap between languages began to shrink, but it did not disappear entirely. One model designed specifically for the Polish language performed better on Polish questions than on English ones, but it still could not match the raw power of the largest, general-purpose models when those were asked in English.
The study also highlighted that passing a multiple-choice test is not the same as being a competent doctor. The exams used in the study are written tests where a student selects the correct answer from a list of options. Real medical practice, however, involves talking to patients, listening to their stories, examining their bodies, and making decisions when there is no single correct answer. The researchers emphasized that while these AI tools are impressive, they cannot yet replace the human interaction and complex judgment required in a hospital. The models can be useful assistants, helping to organize information or draft reports, but they are not ready to take the place of a licensed physician.
Ultimately, the work serves as a warning and a guide for the future of medical technology. It shows that while artificial intelligence is getting better at answering medical questions, it is not yet reliable enough to be trusted blindly, especially when the language or the specific medical specialty changes. The findings suggest that before these tools are used in clinics, they must be tested rigorously in the specific languages and contexts where they will be used. The gap between English and Polish performance proves that a model trained on one set of data cannot simply be assumed to work perfectly in another. As the technology continues to evolve, the goal is to create systems that are not just smart, but also fair and accurate for patients everywhere, regardless of the language they speak.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.