← Latest papers
💻 computer science

Research on Rolling Bearing Degradation Modeling and Remaining Useful Life Prediction Based on TCN and an Improved Transformer Algorithm

This paper proposes a novel hybrid framework combining a Temporal Convolutional Network (TCN) with an improved Transformer featuring segmented sparse attention and MAML-based meta-learning to achieve high-accuracy, computationally efficient, and robust small-sample remaining useful life prediction for rolling bearings under complex operating conditions.

Original authors: Jian Zhao, Jun Huang

Published 2026-09-21
📖 1 min read☕ Coffee break read

Original authors: Jian Zhao, Jun Huang

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Technical Summary: Rolling Bearing Degradation Modeling and RUL Prediction Based on TCN and an Improved Transformer

Problem Statement
The paper addresses three critical challenges in predicting the Remaining Useful Life (RUL) of rolling bearings under complex operating conditions:

  1. Computational Complexity: Standard Transformer models exhibit quadratic computational complexity (O(T2)O(T^2)) relative to sequence length, creating efficiency bottlenecks when modeling ultra-long degradation trajectories typical of full-life-cycle bearing data.
  2. Local Feature Insensitivity: Transformer architectures, relying on global self-attention, often lack the inductive bias necessary to capture short-term temporal fluctuations, transient impacts, and local degradation patterns, which are crucial for early fault detection.
  3. Weak Generalization in Small-Sample Scenarios: In industrial settings, labeled data for new operating conditions (e.g., different loads or speeds) is often scarce. Standard deep learning models struggle to adapt rapidly to these unseen conditions without extensive retraining, leading to performance degradation.

Methodology
The authors propose a hybrid framework integrating a Temporal Convolutional Network (TCN) with an Improved Transformer algorithm, augmented by a Model-Agnostic Meta-Learning (MAML) strategy.

  • TCN Integration for Local Features: A one-dimensional convolutional layer is embedded at the input stage. Utilizing causal and dilated convolutions, this module extracts local temporal features and short-term fluctuation patterns, enhancing the model's sensitivity to transient shock events before data enters the Transformer.
  • Improved Transformer Architecture:
    • Segmented Sparse Attention: To mitigate the O(T2)O(T^2) complexity, the input sequence is divided into non-overlapping segments. Dense self-attention is applied within segments to capture fine-grained local evolution. Between segments, a compressed memory matrix facilitates sparse interactions. This reduces computational complexity to near-linear (O(S2+SL)O(S^2 + SL), where SS is the number of segments) while preserving global trend perception.
    • Locally Enhanced Convolution: A Temporal Convolution Enhancement (TCE) module is introduced to compensate for the Transformer's lack of local inductive bias, ensuring short-term dependencies are modeled effectively.
  • Meta-Learning Strategy (MAML): To address small-sample adaptation, the method employs MAML. Data from various operating conditions are treated as meta-tasks. The model learns a task-general initialization parameter through bi-level optimization, enabling rapid adaptation to new target domains with only a few gradient update steps and minimal labeled data.

Key Contributions

  1. Unified Framework: The paper presents a novel architecture that unifies the local inductive bias of TCN with the global modeling capability of an improved Transformer, effectively balancing local detail capture with long-range dependency modeling.
  2. Efficient Long-Sequence Modeling: The proposed segmented sparse attention mechanism successfully reduces computational complexity from quadratic to linear, making full-life-cycle modeling feasible for industrial deployment.
  3. Small-Sample Adaptation: By integrating MAML, the method demonstrates the ability to learn generalizable initialization parameters, allowing for rapid transfer learning to new operating conditions with limited data.

Experimental Results
The method was validated using 2,400 full-life-cycle vibration datasets collected from a self-built accelerated life test rig under three distinct operating conditions (varying speeds and loads).

  • Prediction Accuracy: Under the target-domain operating condition, the proposed method achieved a Mean Absolute Error (MAE) of 0.13, Root Mean Square Error (RMSE) of 0.22, and Mean Squared Error (MSE) of 0.05, with a coefficient of determination (R2R^2) of 0.99.
  • Comparison: These results represent an improvement of more than two orders of magnitude in accuracy compared to baseline CNN, LSTM, and GRU models. Specifically, error metrics were reduced by 96.24% (MAE), 94.01% (RMSE), and 99.63% (MSE) compared to the best-performing baseline (LSTM).
  • Efficiency: The training time was 34.36 seconds, only slightly higher than the baselines (24–28 seconds), indicating a marginal increase in computational overhead for a significant gain in accuracy.
  • Generalization: Cross-condition transfer tests confirmed the model's ability to adapt to new operating scenarios using only small support sets (5%–20% of target data).

Significance and Claims
The authors claim that this research provides a high-accuracy and strongly generalizable solution for rolling bearing RUL prediction in complex industrial environments. By integrating local feature extraction with efficient global modeling and meta-learning, the proposed method overcomes the trade-offs between computational cost, local sensitivity, and data scarcity. The study suggests that this approach is particularly suitable for intelligent maintenance scenarios where labeled data is limited and operating conditions vary, offering a unified framework for efficient computation, long-range perception, and rapid adaptation.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →