← Latest papers
💬 NLP

A Comparative Benchmark of Large Language Models for Labelling Wind Turbine Maintenance Logs

This paper introduces an open-source framework to benchmark diverse Large Language Models on classifying unstructured wind turbine maintenance logs, revealing that while top models show high alignment and calibration, a Human-in-the-Loop approach is the most effective near-term strategy to leverage their capabilities for improving maintenance data quality.

Original authors: Max Malyi, Jonathan Shek, Alasdair McDonald, Andre Biscaya

Published 2026-04-14
📖 5 min read🧠 Deep dive

Original authors: Max Malyi, Jonathan Shek, Alasdair McDonald, Andre Biscaya

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a wind farm as a giant, complex machine that never sleeps. To keep it running, technicians climb the towers and write down what they see and do in "maintenance logs." Think of these logs as the diary of the wind turbine.

However, there's a big problem: these diaries are written in a messy, chaotic way. One technician might write "hydraulic leak," another might scribble "oil spill in pump," and a third might just write "fix pump." Because the writing is so unstructured and full of abbreviations, computers can't read them. This means the valuable data inside is trapped, like a library where all the books are written in a secret code that no one can decipher.

This paper is about teaching Large Language Models (LLMs)—super-smart AI chatbots—to read these messy diaries and turn them into neat, organized data.

Here is a simple breakdown of what the researchers did and what they found:

1. The Challenge: The "Messy Desk" Problem

The researchers had a pile of 400+ maintenance logs. Some were in English, some in Portuguese, and all were written in a mix of technical jargon, typos, and shorthand.

  • The Goal: Get an AI to read a log like "Stops with error sludge pitch hydraulic" and correctly label it as: Component: Hydraulic System and Issue: Sludge buildup.
  • The Old Way: Humans had to read every single note and type it into a spreadsheet. It was slow, expensive, and different humans often disagreed on what a note meant.

2. The Experiment: The AI "Talent Show"

The researchers set up a massive "talent show" to see which AI models were best at this job. They invited two types of contestants:

  • The "Big Brands" (Proprietary Models): These are the famous, powerful AIs from companies like OpenAI (GPT-5) and Google (Gemini). They are like master chefs who can cook anything but charge a lot for their ingredients.
  • The "Local Heroes" (Open-Source Models): These are free models that can run on a regular laptop. They are like home cooks who are cheaper and keep their recipes secret, but might sometimes burn the toast.

They tested these models on two specific tasks:

  1. The "What?" Task (Objective): Identifying which part broke (e.g., "The blade"). This is easy, like spotting a red car in a parking lot.
  2. The "Why?" Task (Subjective): Identifying what happened to it (e.g., "Was it a repair, an inspection, or a replacement?"). This is harder, like trying to guess if a chef is "cooking" or "prepping" just by looking at a photo.

3. The Results: Who Won?

The study found a clear hierarchy, much like a sports league:

  • The Champions: The biggest, most expensive models (like GPT-5 and GPT-o3) were the most accurate. They rarely made mistakes and were very good at understanding the messy text.
  • The Budget Runners: The smaller, free models were decent but made more mistakes. Sometimes, they would "hallucinate"—which is like a student making up an answer because they are too confident they know the material, even when they don't.
  • The Speed vs. Cost Trade-off: The fastest models were often the cheapest, but they were slightly less reliable. The most reliable models were slower and cost a bit more money to run.

The Big Surprise: Even the best AI wasn't perfect. When the task was subjective (guessing the type of maintenance), even the smartest AIs disagreed with each other. This proved that the problem wasn't just the AI; the human notes were just too vague to have a single "correct" answer.

4. The "Trust Meter" (Calibration)

One of the most important findings was about confidence.

  • Some AIs were like overconfident students: They would say, "I'm 100% sure this is a broken blade!" when they were actually wrong. This is dangerous in engineering.
  • The best models were like honest experts: If they said, "I'm 90% sure," they were usually right. If they said, "I'm only 50% sure," they knew they were guessing.

5. The Final Verdict: The "Human-in-the-Loop"

The researchers concluded that we shouldn't let the AI work alone yet. It's not ready to be the sole boss.

Instead, they propose a Human-in-the-Loop system. Think of it like a GPS navigation system:

  • The AI (the GPS) does the heavy lifting. It reads the messy note and suggests, "This looks like a hydraulic repair."
  • The Human (the driver) just glances at the suggestion and hits "Confirm" or "Correct."

Why is this the best approach?

  • Speed: It's 100x faster than a human reading every note from scratch.
  • Accuracy: The human is still the final judge, so mistakes are caught.
  • Cost: It costs very little (a few dollars to process hundreds of logs) compared to the massive savings gained by having clean data to make better maintenance decisions.

Summary

This paper is a guide for wind energy companies. It says: "Don't try to replace your human experts with AI. Instead, give your experts a super-smart AI assistant that does the boring paperwork, so they can focus on keeping the wind turbines running safely and cheaply."

By using this new framework, wind farms can finally unlock the secrets hidden in their messy notebooks, leading to fewer breakdowns and cheaper electricity for everyone.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →