← Latest papers
🤖 AI

Compass: Navigating Global Marine Lead Data Integration through Expert-Guided LLM Agent

The paper introduces Compass, an expert-guided LLM agent framework that successfully extracts 3,751 previously inaccessible marine lead records from over 230,000 scientific papers to create the largest integrated marine lead database, achieving 92% accuracy through a knowledge-tree-driven approach that bridges the gap between general-purpose AI and high-stakes scientific data integration.

Original authors: Yiming Liu, Bin Lu, Meng Jin, Ziyuan Sang, Shuo Jiang, Lei Zhou, Xinbing Wang, Chenghu Zhou, Jing Zhang

Published 2026-05-29
📖 4 min read☕ Coffee break read

Original authors: Yiming Liu, Bin Lu, Meng Jin, Ziyuan Sang, Shuo Jiang, Lei Zhou, Xinbing Wang, Chenghu Zhou, Jing Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Problem: A Library of Lost Treasure

Imagine the world's oceans are a giant, mysterious library. Inside this library, there are millions of books (scientific papers) written over the last 50 years. These books contain "treasure maps" showing exactly where lead pollution is in the ocean, how it moves, and how it changes over time.

However, there's a catch:

  1. The maps are hidden: The data isn't in a neat spreadsheet; it's buried inside paragraphs of text and messy tables in these books.
  2. The librarians are tired: Finding and copying this data by hand is like trying to read every single book in the library to find one specific sentence. It takes too long and costs too much money.
  3. The robots are confused: If you ask a standard AI (a "general-purpose robot") to read these books and find the data, it often gets it wrong. It might mix up units (like confusing miles with kilometers) or invent facts that don't exist (a problem called "hallucination"). In science, getting the numbers wrong is dangerous.

The Solution: "Compass"

The authors built a new tool called Compass. Think of Compass not as a robot that just guesses, but as a highly trained detective who works with a specialized rulebook.

1. The Expert Rulebook (The Knowledge Tree)

Instead of letting the AI guess, the scientists (who are experts in ocean chemistry) sat down with the AI developers to write a "Knowledge Tree."

  • The Analogy: Imagine a flowchart for a detective.
    • Step 1: Is this book about the ocean? If yes, go to Step 2. If no, stop.
    • Step 2: Is the table talking about lead? If yes, go to Step 3.
    • Step 3: Does the number look physically possible? (e.g., Lead can't be negative). If yes, keep it. If no, throw it out.
  • This "tree" forces the AI to follow strict, logical steps designed by human experts. It prevents the AI from making wild guesses.

2. The Three-Step Detective Work

Compass uses this rulebook to do three things in order:

  1. Collection: It scans over 230,000 open-access scientific papers to find the ones that might have the treasure.
  2. Extraction: It reads the specific tables and text in those papers, using the rulebook to pull out the exact numbers for lead concentration and isotope ratios.
  3. Aggregation: It takes all those scattered numbers and cleans them up, making sure they all speak the same "language" (same units, same format) so they can be put into one giant database.

The Results: Finding Hidden Gold

By using Compass, the team achieved something huge:

  • The Hunt: They processed 230,000 papers.
  • The Catch: They found 3,751 new records of lead data that had never been put into a global database before.
  • The Accuracy: When human ocean scientists checked the work, they found it was 92% accurate. This is a massive improvement over standard AI, which often fails at these specific scientific tasks.
  • The Map: They created the largest map of marine lead data to date. It fills in huge blank spots on the map, especially in areas like the East China Sea and the Southern Ocean, which were previously "data deserts."

Why This Matters

Think of the ocean as a giant puzzle. For decades, scientists only had a few scattered pieces. They couldn't see the whole picture.

  • Before Compass: Scientists had to guess what the missing pieces looked like.
  • With Compass: They now have thousands of new pieces. They can finally see how lead pollution travels from factories, through the air, and into the deep ocean. They can see how the ocean is recovering after we stopped using leaded gasoline.

What They Did NOT Do

  • They did not build a new type of robot that learns on its own without help.
  • They did not claim this tool can diagnose diseases or help with medical treatments.
  • They did not say this solves all ocean problems; it is specifically designed to find and organize lead data from scientific papers.

In a Nutshell

The authors built a smart, rule-following assistant that acts like a bridge between messy, old scientific books and a clean, modern database. By giving the AI a strict "rulebook" written by ocean experts, they were able to rescue thousands of lost data points, creating the most complete map of ocean lead pollution the world has ever seen.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →