MineralBench: an evaluation benchmark of large language models for mineral resources
This paper introduces MineralBench, a comprehensive evaluation benchmark featuring three specialized tasks and four metrics to systematically assess and compare the performance of 13 large language models in the mineral resources domain, revealing significant performance gaps between general and domain-specific capabilities and the task-dependent nature of fine-tuning.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Earth holds a vast, silent library of information written in stone. For centuries, geologists have read this library by hand, sifting through mountains of reports, maps, and field notes to find where valuable minerals lie. This manual work is slow, expensive, and prone to human error, especially when the data piles up. In recent years, a new kind of computer program has emerged that can read and understand human language with surprising fluency. These programs, known as large language models, have shown they can answer questions, summarize texts, and solve problems across many fields. But while they are excellent at general knowledge, no one knew for sure if they could handle the specific, technical language of geology. The question was not just whether these programs could read a geology book, but whether they could actually help a geologist find a new mine or understand a complex rock formation without making dangerous mistakes.
To answer this, a team of researchers from the China University of Geosciences and the Chinese Academy of Geological Sciences built a new testing ground called MineralBench. Think of this as a specialized exam designed specifically for artificial intelligence in the world of minerals. The researchers did not just ask the computers to guess random facts; they created three distinct types of challenges that mimic real work. First, they tested if the models could understand the definitions of specific mineral terms, like knowing exactly what "pyrite decay" means. Second, they asked the models to identify specific properties of minerals, such as listing the chemical elements that make up a rock called Actinolite. Finally, they gave the models long, unstructured paragraphs of geological text and asked them to pull out specific names of minerals and rock layers, a task that requires finding needles in a haystack of words.
The team put thirteen different artificial intelligence models through this gauntlet. Some were massive, cloud-based systems accessed over the internet, while others were smaller versions that could run on a single computer in a lab. The results revealed a clear divide. The models were generally quite good at the general science questions, often scoring high on standard tests about math and basic science. However, when the questions shifted to the specific world of minerals, their performance dropped significantly. No single model was the best at everything. One model might be the fastest at finding names in a text, while another was better at listing chemical elements, and a third might be the most consistent at giving the same answer every time it was asked the same question. This suggests that there is no single "perfect" geology AI yet; instead, different models have different strengths, much like a team of specialists where each person excels at a different part of the job.
The researchers also looked at how stable the answers were. They asked the same questions multiple times to see if the models would give the same answer or if they would change their minds. They found that while some models were very consistent, they were sometimes consistently wrong, repeating the same incorrect answer over and over. Others were more variable but occasionally hit the right answer. This discovery is crucial because it shows that simply asking a computer once is not enough; for serious work, you need to know if the answer is both correct and reliable. The study also tested what happened when they tried to "teach" a model more about geology by fine-tuning it with extra data. They found that this process was a double-edged sword: the model got much better at one specific task, like extracting names from text, but it actually got worse at other tasks, like solving math problems. This indicates that teaching an AI to be an expert in one narrow field can sometimes make it forget how to do other things well.
Ultimately, this work provides a clear map for anyone trying to use artificial intelligence in the mineral industry. It shows that while these tools have great potential, they are not yet ready to replace human experts on their own. The best approach appears to be a careful selection of the right tool for the specific job, understanding that a model that is fast might not be accurate, and a model that is accurate might be slow. By establishing these standards and revealing the specific strengths and weaknesses of current technology, the researchers have given the scientific community a way to measure progress and choose the right digital assistant for the difficult work of exploring the Earth's resources.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.