← Latest papers
💬 NLP

Leveraging Few-Shot Learning and Large Language Models for Analyzing Blood Pressure Variations Across Biological Sex from Scientific Literature

This study evaluates the feasibility of using large language models (LLMs) and few-shot learning to automatically extract sex-specific blood pressure statistics from scientific literature, aiming to address the limitations of current demographic-blind diagnostic standards.

Original authors: Yuting Guo, Seyedeh Somayyeh Mousavi, Reza Sameni, Abeed Sarker

Published 2026-08-17
📖 5 min read🧠 Deep dive

Original authors: Yuting Guo, Seyedeh Somayyeh Mousavi, Reza Sameni, Abeed Sarker

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to bake the perfect cake for a crowd, but your recipe book was written fifty years ago by a very specific group of bakers who only ever baked for people who looked and lived exactly like them. You try to use that old recipe for everyone else, but the cake turns out too dry for some and too sweet for others. This is essentially the problem with how we currently measure blood pressure. For decades, doctors have used a single set of "normal" numbers to decide if someone's heart is under too much stress, often ignoring that a person's biology—like whether they are male or female—might mean they need a different set of rules.

To fix this, scientists are turning to a new kind of digital detective: Large Language Models (LLMs). Think of these models as super-smart robots that have read almost everything ever written on the internet. They are like a student who has memorized every library in the world and can instantly find a specific fact in a million-page book without getting tired. The big question researchers wanted to answer was: Can these AI robots read thousands of old medical studies, find the specific blood pressure numbers for men and women, and tell us if the old "one-size-fits-all" recipe is actually broken?


The Great Blood Pressure Treasure Hunt

In this study, a team of researchers decided to put these AI robots to the test. They wanted to see if the robots could dig through a mountain of scientific papers to find a very specific treasure: the average blood pressure numbers and how much those numbers vary, separated by biological sex (male vs. female).

The Setup: A Digital Gold Rush
The researchers started with a massive pile of data. They grabbed over 5.6 million scientific manuscripts from a giant online library called PubMed Central. Imagine trying to find a few specific needles in a haystack the size of a city. To make this easier, they built a special search engine (like a super-charged Google for scientists) that filtered the papers down to about 113,000 that talked about blood pressure. Then, they broke those papers into smaller chunks (paragraphs) and filtered again, ending up with about 50,000 relevant paragraphs.

From this huge pile, they picked a smaller, manageable group of 213 articles to test their AI tools on. They even had human experts read these articles and write down the correct numbers to use as a "answer key."

The Contest: Three Different Detectives
The researchers set up a race between three different methods to see which one could find the right numbers best:

  1. The Old School Rookie (DANN): This is a traditional computer program that needs to be shown a few examples (like "few-shot learning") before it can guess the rest. It's like a student who has to study a few flashcards before taking a test.
  2. The Famous Chatbot (GPT-3.5): This is a well-known AI that can chat and write stories. The researchers asked it to just read the text and find the numbers without giving it any special training first (this is called "zero-shot").
  3. The Open-Source Giant (LLaMA3): This is another powerful AI, similar to the chatbot but built by a different group. Like the chatbot, it was also asked to do the job without any special training, just using its general knowledge.

The Results: Who Won the Race?
The results were pretty clear. The "Old School Rookie" (DANN) struggled mightily. It only got about 30% of the answers right. It was like a student who guessed randomly on a test.

The "Famous Chatbot" (GPT-3.5) did better, getting about 67% of the answers correct. It was a solid middle-ground performer, but it sometimes missed things it should have found.

The winner was the "Open-Source Giant" (LLaMA3). It scored an impressive 0.85 (or 85%) on the accuracy scale. It was the most reliable detective, finding the right numbers for men and women with high precision and rarely missing a clue. The researchers found that this AI was so consistent that its scores didn't even wiggle around much when they ran the test multiple times.

What Did They Find in the Treasure?
Once the winning AI (LLaMA3) successfully pulled the numbers out of the papers, the researchers looked at the patterns. They created colorful maps (called heatmaps and contour plots) to visualize the data. These maps showed that, generally speaking, males tend to have higher blood pressure values than females.

This isn't a shocking discovery in itself—previous studies have hinted at this—but the cool part is how they found it. They didn't ask a human to read thousands of papers; they used an AI to do it automatically. This suggests that we can use these powerful robots to scan the entire history of medical literature to see how different groups of people are actually doing, rather than relying on old, generic rules.

The Bottom Line
The study concludes that these AI robots are ready for the big leagues. They can read scientific literature, separate the data by sex, and give us a clearer picture of blood pressure trends than we could get by just guessing or using old methods. While the researchers admit they only looked at sex and didn't check other factors like age or race in this specific run, the method works. It's like finally having a tool that can read the fine print in a million contracts to tell you if the rules are fair for everyone, not just the people the rules were originally written for.

The paper suggests that by using these tools, we might eventually be able to create more personalized, accurate health guidelines that fit real people, rather than forcing everyone into a box that doesn't quite fit.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →