← Latest papers
🌿 ecology

Large language models possess some ecologicalknowledge, but how much?

This study evaluates the ecological knowledge of Gemini 1.5 Pro and GPT-4o using a new benchmark dataset, revealing that while these models outperform naive baselines in tasks like species presence prediction, their overall performance remains significantly below expert levels, indicating a critical need for domain-specific fine-tuning to effectively integrate LLMs into ecological science.

Original authors: Dorm, F., Millard, J., Purves, D., Harfoot, M., Mac Aodha, O.

Published 2026-02-04
📖 3 min read☕ Coffee break read

Original authors: Dorm, F., Millard, J., Purves, D., Harfoot, M., Mac Aodha, O.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine Large Language Models (LLMs) like Gemini 1.5 Pro and GPT-4o as incredibly well-read librarians who have read almost every book in the world. They are famous for being able to answer questions about history, math, and pop culture with ease. But the researchers behind this paper wanted to know: How good are these librarians at talking about nature?

To find out, the team treated the AI models like students taking a very specific "Ecology Exam." They didn't just ask general questions; they gave them five tough practical tasks:

  1. Guessing who lives where: Predicting if a specific animal or plant is present in a certain area.
  2. Drawing maps: Creating a map showing where a species' home range is.
  3. Listing the endangered: Naming species that are on the brink of extinction.
  4. Identifying dangers: Classifying what threats are hurting these species.
  5. Describing traits: Estimating physical characteristics of different species.

The researchers created a new "answer key" based on real expert data to grade the AI's homework.

Here is how the AI students performed:

  • The Good News: They weren't total beginners. When it came to guessing if a species was present in an area, the AI scored about 20 percentage points higher than someone just guessing randomly. It's like a student who knows the basics of the textbook and can pass the easy questions.
  • The Reality Check: When the tasks got harder, the AI struggled.
    • When asked to draw range maps (showing exactly where an animal lives), the AI only managed to get about one-third of the accuracy that a human expert would achieve. It's like trying to draw a detailed map of a city from memory and getting the major streets right, but missing all the neighborhoods.
    • When asked to classify threats, the AI only improved its score by about 10 points over random guessing. It's barely better than flipping a coin to decide what is hurting a species.

The Bottom Line:
The paper concludes that while these AI models have some knowledge of ecology, it's not deep or precise enough to replace human experts yet. They are like a tourist who has read a travel guide about a forest but hasn't actually walked the trails.

To fix this, the authors suggest that these models need specialized training (fine-tuning) specifically on ecological data, rather than just relying on their general knowledge. The main goal of this paper was simply to set up a fair "test" (a benchmark) so that future researchers can measure exactly how much these AI tools improve over time.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →