← Latest papers
⚡ electrical engineering

Overview of Publicly Available Degradation Data Sets for Tasks within Prognostics and Health Management

This paper provides an overview of publicly available degradation data sets that are essential for analyzing the evolving health conditions, failure modes, and performance trends of engineering systems within the field of prognostics and health management.

Original authors: Fabian Mauthe, Christopher Braun, Julian Raible, Peter Zeiler, Marco F. Huber

Published 2026-02-06
📖 5 min read🧠 Deep dive

Original authors: Fabian Mauthe, Christopher Braun, Julian Raible, Peter Zeiler, Marco F. Huber

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to predict when a machine is going to break. To do this, the robot needs to "study" from a library of stories about machines that have actually broken down in the past. These stories are called degradation data sets.

This paper is essentially a massive, organized library catalog for that specific type of story. The authors, a team of researchers from Germany, have gone through the public internet to find every available "story" (data set) where engineers have recorded how machines get sick and eventually fail. They found 110 different data sets and organized them so that anyone can easily find the right one for their needs.

Here is a breakdown of what they did, using some everyday analogies:

1. The Problem: The "Lost in the Library" Dilemma

In the world of machine health (called Prognostics and Health Management or PHM), researchers need data to train their AI. But finding the right data is like trying to find a specific book in a library where the books are thrown in a pile, the shelves are empty, and the catalog is missing.

  • The Paper's Goal: They wanted to stop researchers from wasting time searching. They built a clear, organized map (a taxonomy) to show exactly what data is out there, what it's about, and what kind of "illness" it tracks.

2. The Four Stages of Machine Health

The paper explains that tracking a machine's health is like a doctor's visit, which happens in four steps. The data sets are sorted by which step they help with:

  • Fault Detection (The "Is it sick?" check): Like a smoke alarm. It just tells you, "Hey, something is wrong!" It doesn't know what or why.
  • Diagnosis (The "What's wrong?" check): Like a doctor asking, "Is it a broken leg or a sprained ankle?" This identifies the specific cause of the problem.
  • Health Assessment (The "How bad is it?" check): Like a doctor saying, "The leg is broken, and it's 50% healed." It measures the current state of damage.
  • Prognosis (The "When will it fail?" check): Like a doctor predicting, "You will need a cast for six more weeks." This predicts the Remaining Useful Life (RUL)—how much time is left before the machine stops working.

3. The Library Catalog (The Results)

The authors sorted their 110 data sets into different categories, much like organizing a library by genre and topic.

The "Genres" (Domains):
They found that most stories are about Mechanical Parts (like gears and bearings) and Electrical Parts (like batteries and motors).

  • The Stars of the Show: Batteries and Bearings are the most popular subjects. There are 16 data sets about batteries and 17 about bearings. It's like having 17 different books about "How to fix a bicycle wheel" and 16 books about "How to fix a car battery."
  • The Obscure Genres: There are also stories about aircraft engines, filtration systems, and even subway doors, but there are very few of them. Some categories only have one or two data sets, making it hard to learn from them.

The "Plot Points" (Signals):
To tell the story of a machine failing, you need to record specific things, called signals. Think of these as the vital signs of the machine.

  • Vibration: The most common signal (34 data sets). It's like listening to the engine rumble to hear if it's knocking.
  • Current & Temperature: The second most common (33 data sets each). This is like checking the machine's heart rate and body temperature.
  • Voltage: Also very common (25 data sets).
  • The Missing Links: The paper notes that for many specific machines, we don't have enough "vital sign" recordings. For example, while we have lots of data on how batteries fail, we have very little data on how specific types of robots or production lines fail.

4. The Big Takeaway

The paper concludes that while we have a great "library" for a few popular machines (like batteries and bearings), the rest of the library is quite empty.

  • What this means: If you want to train an AI to predict when a battery will die, you have a huge library of books to choose from. If you want to train an AI to predict when a specific type of industrial robot will break, you might only find one or two thin pamphlets.
  • The Future: The authors plan to keep updating this catalog as new data is published, ensuring that researchers don't have to start from scratch every time they want to study a new machine.

In short, this paper is a guidebook that says: "Here is everything we know about machines getting sick, organized by what kind of machine it is and what kind of sickness it has, so you don't have to hunt for it yourself."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →