Enhancing Deep Neural Network Resilience Through Integrated Reliability-Aware Design and Fault-Aware Training
This paper introduces ResilientNet, a reliability-aware framework that enhances Deep Neural Network resilience against hardware faults through four synergistic design and training strategies—bounded activations, layer reordering, NaN filtering, and curriculum-based fault training—demonstrated to significantly reduce error propagation and maintain high classification accuracy with negligible inference overhead across extensive fault injection experiments.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Deep neural networks are the invisible engines powering some of the most critical decisions in modern life, from helping airplanes avoid collisions to assisting doctors in diagnosing diseases. These systems are designed to recognize patterns and make predictions with a level of accuracy that rivals human experts. However, the hardware that runs these complex calculations is not perfect. The tiny chips inside computers, which operate at scales so small they are measured in billionths of a meter, are vulnerable to invisible threats. Cosmic rays and other natural radiation can occasionally strike a memory cell or a processing unit, causing a momentary glitch. In a standard computer, such a glitch might cause a single number to flip, but in a deep neural network, that single error can ripple through the system, corrupting the final result without ever triggering an alarm. This phenomenon, known as silent data corruption, is particularly dangerous because the system continues to operate, delivering a wrong answer that looks completely correct to the user.
For years, the solutions to this problem have been difficult trade-offs. Engineers could build special, expensive hardware that is immune to these glitches, but this requires custom-made silicon that is not available in standard computers. Alternatively, they could add software checks that catch errors, but these checks slow down the system significantly, making it too slow for real-time tasks like guiding a vehicle or managing air traffic. Researchers at Puducherry Technological University have now proposed a different path. They developed a new framework called ResilientNet, which strengthens the neural network itself against these hardware failures without needing special chips or slowing down the processing. Their approach combines four specific design changes that work together to stop errors before they can spread, effectively teaching the network to ignore the noise caused by radiation.
The core of this new method involves changing how the network handles the numbers it processes. In a typical neural network, if a radiation strike causes a number to become impossibly large, the system passes that huge number along to the next stage, where it can overwhelm the entire calculation. The researchers replaced the standard way of handling these numbers with a method that sets a strict upper limit. If a value tries to exceed this limit, it is simply capped, preventing it from growing out of control. However, this alone was not enough. The researchers found that the order in which the network performs its calculations mattered just as much as the limits themselves. By rearranging the sequence of operations so that the capping happens before the data is normalized, they prevented the system from amplifying errors. This rearrangement, combined with a filter that instantly replaces impossible mathematical values with zeros, creates a barrier that stops corruption from spreading.
To ensure these changes actually worked, the team did not rely on computer simulations or theoretical models. Instead, they ran a massive series of real-world tests. They took six different versions of a neural network, ranging from a basic setup to the fully hardened ResilientNet, and subjected them to nearly fifteen thousand separate fault injections. In each test, they deliberately corrupted the data to mimic the exact patterns of errors seen in real radiation experiments, including single-bit flips, rows of corrupted data, and blocks of damaged information. They measured how often these errors led to a wrong prediction, a metric known as the silent data corruption rate. The results were striking. In the unmodified network, a specific type of error caused by invalid numbers led to a corruption rate of ninety-two percent. With the new filtering and design changes in place, that rate dropped to just seventeen percent. Furthermore, the network that had been trained to expect these errors actually performed better at its primary job than the original network, achieving a classification accuracy of ninety-six point six seven percent compared to the baseline of ninety-two point five percent.
The study also revealed that not all parts of the network are equally vulnerable. The researchers mapped out where errors were most likely to cause a failure and found that the earliest layers of the network were the most sensitive, with a vulnerability rate nearly one and a half times higher than the later layers. This suggests that in resource-constrained environments, protection efforts could be focused on the beginning of the processing chain. Crucially, all these improvements came with zero added time cost during operation. The changes to the code did not require extra steps that would delay the results, and the only minor increase was in the number of internal registers used by the processor, a change that had no measurable impact on speed. While the tests were conducted on a relatively small dataset of handwritten digits, the principles were verified using real code execution rather than approximations, offering a promising blueprint for making safety-critical systems more reliable without the need for expensive hardware upgrades.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.