Nigeria Machinery: A Low-Resource Industrial Dataset with a Domain-Grounded Reasoning Layer
This paper introduces the Nigeria Machinery Usage and Failures Dataset, a small but meticulously sourced collection of industrial records from 2006–2025, accompanied by a domain-grounded reasoning layer that ensures language model prompts accurately reflect real-world numeric values and sources.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine trying to teach a super-smart robot how to fix a broken oil refinery in Nigeria. You'd want to show it real stories: "In 2020, the Port Harcourt refinery stopped working for three days because of X," or "The manufacturing sector used only 40% of its machines." But here's the catch: all those real stories are locked away in dusty, giant PDF reports, scattered across different government buildings, and written in tables that robots can't read. It's like having a library of treasure maps where every map is written in a different language and hidden in a different attic.
That's exactly the problem Gospel Bassey and Vincent Fakiyesi tackled. They didn't just find the maps; they built a tiny, super-detailed treasure chest called the Nigeria Machinery Usage and Failures Dataset.
The Treasure Chest: 89 Real Clues
Think of this dataset as a collection of 89 specific, real-world clues about machines in Nigeria, stretching from 2006 to 2025. These clues cover two main neighborhoods: Industrial Manufacturing (38 clues) and Oil and Gas (51 clues).
Inside, you'll find facts about things like how much fuel is being made, how many machines are sitting idle (downtime), and how much money is spent on repairs. But here's the most important part: every single clue has a "receipt." The authors didn't guess or make things up. If a number says "40% capacity," they can point you to the exact page in a government report where that number was printed.
However, the authors are very honest about the size of their chest. They call it a "seed" dataset, not a giant forest. Why? Because 17 of the 28 types of clues they found only have one single entry. It's like having a dictionary where most words only have one example sentence. You can learn the meaning of those words, but you can't write a whole novel with them yet. It's a starting point, a reference guide, not a massive training manual for a robot to learn everything on its own.
The "Fake Expert" Problem
Now, here is the clever part of their story. The authors realized that just having the numbers isn't enough. You also need to teach the robot how to think about them. This is where they ran into a funny but dangerous trap.
Imagine you have a robot that is really good at math but terrible at context. You give it a number, say 35.5, and ask it to explain it.
- The Bad Way: The robot might say, "Ah, 35.5 is the temperature of a patient's fever!" It got the number right, but the story is completely wrong. It's like a detective solving a murder case by guessing the victim was a fish.
- The Old Release: In an earlier version of their work, they had 78 of these "thinking" examples. But 77 of them were like the fever example. They had the right numbers, but the stories were about medical sensors or abstract math, not Nigerian oil refineries. Only 1 out of 78 actually talked about the real industrial world.
The authors argue that this is a big mistake. If you teach a robot about medical sensors using Nigerian oil data, the robot learns nothing about how to fix a refinery. It's like teaching a pilot to fly a plane using a video game about driving a car. The controls might look similar, but the reality is totally different.
The Fix: Real Stories, Real Thinking
To fix this, the authors (with help from Adaption Labs) rebuilt the "thinking" part from scratch. They created 94 new examples where the robot is forced to act like a real engineer.
Instead of asking, "What is the log base 2 of 512?", they ask, "What was the capacity utilization of the Warri Refinery in 2015, according to the NUPRC report?"
- The Result: Every single one of the 94 new examples is grounded in the real world.
- The Math: They also added 6 examples where the robot has to do real, multi-step math (like calculating the cost per day of a machine being broken), rather than just looking up a number.
- The Proof: They set up a strict "bouncer" at the door. If a story didn't mention a real industry term, or if the math didn't match the source perfectly, it got kicked out. Because of this, 100% of the final 94 rows passed the test. Every answer matches the source document exactly.
What This Means for You
The authors are very clear: this isn't a magic wand that will instantly fix all of Nigeria's machines. They explicitly rule out the idea that this small dataset can train a robot to predict failures on a massive scale right now. There just aren't enough numbers for that.
Instead, think of this as a blueprint and a starter kit.
- The Data: It proves that you can build a reliable, real-world dataset from scratch, even when the information is scattered and hard to find.
- The Method: It shows that if you want AI to be useful in the real world, you can't just feed it numbers; you have to feed it the story behind the numbers.
- The Future: They hope this "seed" will grow. They want other researchers to use this method to build similar "starter kits" for other industries and countries that are currently ignored by big tech.
In short, the paper says: "We found the real numbers, we fixed the fake stories, and we made a small but honest toolbox. It's not the whole house yet, but it's a solid foundation for building one."
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.