Research on Spur Gear Fault Diagnosis Method Based on Vibration Signal

In the context of the rapid development of modern industry, traditional mechanical equipment is undergoing a profound transformation toward informatization, intelligence, and high efficiency. As the core transmission component in rotating machinery, the spur gear pair directly determines the reliability and safety of the entire mechanical system. However, due to the harsh operating environment, high load, and alternating stress, spur gears are prone to various forms of damage such as wear, pitting, and tooth fracture. Early gear faults are often weak and submerged in strong background noise, making their feature extraction extremely difficult. To address this issue, I conducted a systematic study on vibration-based fault diagnosis methods for spur gears. This paper focuses on the diagnosis of early weak faults in spur gears under non-stationary operating conditions, proposing an integrated framework that combines improved signal decomposition, multi-domain feature extraction and fusion, and optimized support vector machine classification. The effectiveness of the proposed method is verified through a self-designed gear fault diagnosis experimental platform.

1. Introduction to Gear Fault Diagnosis

Gearboxes are widely used in manufacturing, transportation, aerospace, and other industrial fields. Their failure not only reduces production efficiency but also may cause severe safety accidents. According to statistical data, gear failure accounts for approximately 60% of the total failures in rotating machinery. Therefore, condition monitoring and fault diagnosis of gear transmission systems are of great significance for ensuring equipment reliability and reducing maintenance costs. Vibration signal analysis is the most commonly used method for gear fault diagnosis because vibration signals are sensitive to structural changes and contain rich fault-related information. However, the vibration signals of spur gears are typically non-stationary and non-linear, especially when the gear operates under variable speed or variable load conditions. The early weak fault features are often masked by strong ambient noise, which makes fault diagnosis a challenging task.

The diagnosis process can be divided into three main stages: signal preprocessing and denoising, fault feature extraction, and pattern recognition. In the first stage, effective signal processing algorithms are needed to suppress noise and highlight fault impulses. In the second stage, representative features are extracted from multiple domains to characterize the gear health state. In the third stage, a reliable classifier is constructed to automatically identify different fault types. In this paper, I focus on all three stages and propose improvements to achieve more accurate and efficient spur gear fault diagnosis.

The main contributions and organization of this study are as follows: (1) A best decomposition layer calculation method for Variational Mode Decomposition (VMD) based on singular value kurtosis difference spectrum is proposed; (2) The Maximum Correlated Kurtosis Deconvolution (MCKD) algorithm is optimized by the Dung Beetle Optimizer (DBO) using the sample entropy minimum principle to enhance weak fault pulses; (3) Multi-domain features are extracted from time, frequency, and time-frequency domains, and the Random Forest (RF) algorithm combined with Principal Component Analysis (PCA) is used for feature selection and fusion; (4) A hybrid kernel support vector machine model optimized by DBO is constructed for gear fault classification.

2. Vibration Mechanism and Fault Characteristics of Spur Gears

To establish a solid theoretical basis for fault diagnosis, I first analyzed the vibration mechanism of spur gear pairs. The dynamic model of a gear pair can be simplified as a spring-damper system. The gear meshing process is accompanied by periodic changes in meshing stiffness, which is one of the main internal excitations of gear vibration. When a single pair of teeth meshes, the stiffness is relatively low, while during double-tooth meshing, the stiffness increases significantly. This periodic variation in meshing stiffness produces a parametric excitation, leading to the generation of meshing vibration. The dynamic equation of the gear pair can be expressed as:

$$M_m \ddot{x} + C \dot{x} + K(t) x = F(t)$$

where \(M_m\) is the equivalent mass of the gear pair, \(C\) is the meshing damping coefficient, \(K(t)\) is the time-varying meshing stiffness, \(x\) is the relative displacement along the line of action, and \(F(t)\) is the dynamic meshing force. By further considering the gear deformation and transmission error, the dynamic equation can be written as:

$$M_m \ddot{x} + C \dot{x} + K(t)x = K(t)(E_1 + E_2(t))$$

where \(E_1\) represents the deformation caused by the static load and \(E_2(t)\) represents the relative displacement caused by gear errors and faults. The meshing stiffness can be calculated based on the potential energy method. The bending stiffness \(k_b\), shear stiffness \(k_s\), axial compression stiffness \(k_a\), and Hertzian contact stiffness \(k_h\) can be derived from the cantilever beam theory. Taking into account the gear foundation flexibility, the total meshing stiffness of a single tooth pair can be expressed as:

$$\frac{1}{k_m} = \frac{1}{k_{b1}} + \frac{1}{k_{s1}} + \frac{1}{k_{a1}} + \frac{1}{k_{f1}} + \frac{1}{k_{b2}} + \frac{1}{k_{s2}} + \frac{1}{k_{a2}} + \frac{1}{k_{f2}} + \frac{1}{k_h}$$

In the spectrum analysis of gear vibration signals, the gear meshing frequency and its harmonics are the dominant components. The meshing frequency \(f_c\) is determined by the shaft rotational frequency \(f_r\) and the number of gear teeth \(N\):

$$f_c = N \cdot f_r$$

When a local fault occurs on a gear tooth, such as a tooth crack, pitting, or breakage, the fault produces periodic impact pulses during meshing. These pulses excite the natural resonance of the gear-shaft-bearing system, resulting in amplitude modulation and frequency modulation phenomena in the vibration signal. Considering the combined effects of both AM and FM, the fault vibration signal model can be expressed as:

$$x(t) = X_i \left[ 1 + a_i(t) \right] \cos \left[ 2\pi N f_r t + \varphi_i + b_i(t) \right]$$

where \(a_i(t)\) is the amplitude modulation function and \(b_i(t)\) is the frequency modulation function. The characteristic sidebands appear around the meshing frequency and its harmonics at intervals equal to the shaft rotational frequency. Thus, the sideband structure of the spectrum provides valuable clues for identifying the existence and location of gear faults.

The most common gear fault modes include tooth wear, tooth pitting, tooth breakage, and their combinations. Tooth breakage accounts for about 41% of gear failures, tooth pitting accounts for about 31%, and tooth wear accounts for about 20%. Different fault modes produce different vibration characteristics. For example, tooth fracture produces periodic high-energy shocks with large amplitude; pitting causes an increase in high-frequency components and sideband energy; wear produces broad-spectrum noise and reduces the meshing stiffness. The gear fault modes and their corresponding vibration characteristics are summarized in the following table.

Fault Mode Proportion Main Cause Vibration Characteristics
Tooth Breakage 41% Fatigue crack, overload Large periodic impact pulses, dense sidebands around meshing frequency
Tooth Pitting 31% Surface fatigue, cyclic contact stress Increased high-frequency energy, extra spectral peaks
Tooth Wear 20% Abrasive particles, insufficient lubrication Sidebands clustered in a frequency range, new spectral components
Compound Faults Combination of the above factors Complex sideband structures, strong modulation

3. Vibration Signal Processing Based on Improved VMD

In order to extract weak fault features from vibration signals polluted by strong background noise, I adopted VMD as the core signal decomposition method. VMD is an adaptive non-recursive signal decomposition technique that can decompose a complex signal into a series of band-limited intrinsic mode functions (IMF). Compared with EMD and its improved versions, VMD has solid mathematical foundations and can effectively avoid the mode mixing problem. The VMD algorithm defines the constrained variational problem as:

$$\min_{\{u_k\},\{\omega_k\}} \left\{ \sum_{k=1}^{K} \left\| \partial_t \left[ \left(\delta(t) + \frac{j}{\pi t}\right) * u_k(t) \right] e^{-j\omega_k t} \right\|_2^2 \right\} \quad \text{s.t.} \quad \sum_{k=1}^{K} u_k = f(t)$$

where \(\{u_k\} = \{u_1, u_2, \dots, u_K\}\) represent the decomposed IMFs, \(\{\omega_k\} = \{\omega_1, \omega_2, \dots, \omega_K\}\) are their corresponding center frequencies, \(K\) is the total number of modes, and \(f(t)\) is the original input signal. To solve this constrained optimization problem, a quadratic penalty term \(\alpha\) and the Lagrangian multiplier \(\lambda(t)\) are introduced to transform it into an unconstrained problem:

$$L(\{u_k\}, \{\omega_k\}, \lambda) = \alpha \sum_{k=1}^K \left\| \partial_t \left[ \left(\delta(t) + \frac{j}{\pi t}\right) * u_k(t) \right] e^{-j\omega_k t} \right\|_2^2 + \left\| f(t) – \sum_{k=1}^K u_k(t) \right\|_2^2 + \langle \lambda(t), f(t) – \sum_{k=1}^K u_k(t) \rangle$$

The alternating direction method of multipliers (ADMM) is employed to solve the above problem iteratively. In the frequency domain, the update equations for the modes and center frequencies are:

$$\hat{u}_k^{n+1}(\omega) = \frac{\hat{f}(\omega) – \sum_{i \neq k} \hat{u}_i(\omega) + \hat{\lambda}(\omega)/2}{1 + 2\alpha(\omega – \omega_k)^2}$$

$$\omega_k^{n+1} = \frac{\int_0^{\infty} \omega |\hat{u}_k(\omega)|^2 d\omega}{\int_0^{\infty} |\hat{u}_k(\omega)|^2 d\omega}$$

The selection of the decomposition layer number \(K\) is critical for VMD performance. If \(K\) is too small, the signal cannot be fully decomposed, while if \(K\) is too large, over-decomposition occurs. In this study, I proposed a new method to determine the optimal \(K\) value based on the singular value kurtosis difference spectrum. The original signal is first used to construct a Hankel matrix. Then, the singular value decomposition is performed to obtain a singular value matrix. The kurtosis difference spectrum is calculated to locate the maximum mutation position, which corresponds to the separation between useful signal components and noise. The flowchart of the proposed optimal \(K\) selection method is as follows:

Step 1: Construct the Hankel matrix of the original signal and apply SVD to obtain singular values;

Step 2: Calculate the singular value difference spectrum and use kurtosis to identify the maximum abrupt change point;

Step 3: Take the position of this abrupt change as the initial K value;

Step 4: Apply VMD to decompose the original signal into K IMFs;

Step 5: Calculate the center frequency distance between adjacent IMFs and judge whether over-decomposition exists;

Step 6: If over-decomposition is detected, decrease K and repeat the process until the optimal K is found.

After obtaining the IMFs, it is necessary to select the informative components for signal reconstruction. I used the correlation coefficient and kurtosis as the selection criteria. The correlation coefficient \(r_{xy}\) measures the similarity between each IMF and the original signal:

$$r_{xy} = \frac{\sum_{i=1}^N (x_i – \bar{x})(y_i – \bar{y})}{\sqrt{\sum_{i=1}^N (x_i – \bar{x})^2} \sqrt{\sum_{i=1}^N (y_i – \bar{y})^2}}$$

Kurtosis reflects the impulsiveness of the signal and is defined as:

$$k = \frac{E[(x-\mu)^4]}{\sigma^4}$$

According to experience, the kurtosis of a healthy gear vibration signal is close to 3. When a fault exists, the kurtosis value will be significantly larger than 3 due to the presence of periodic impulses. Therefore, the IMFs that satisfy both the correlation coefficient criterion and the kurtosis criterion are selected for signal reconstruction.

4. Fault Feature Enhancement Based on Optimized MCKD

To further enhance the weak periodic impulses in the reconstructed signal, I applied the Maximum Correlated Kurtosis Deconvolution (MCKD) method. MCKD is an effective deconvolution technique that maximizes the correlated kurtosis of the output signal to recover the periodic impact components from a noisy signal. The output signal \(y_n\) is obtained by convolving the input signal \(x_n\) with a finite impulse response filter:

$$y_n = \sum_{i=1}^{L} h_i x_{n-i+1}$$

where \(L\) is the filter length and \(h_i\) are the filter coefficients. The correlated kurtosis of the output signal is defined as:

$$CK_M(T) = \frac{\sum_{n=1}^{N} \left( \prod_{m=0}^{M} y_{n-mT} \right)^2}{\left( \sum_{n=1}^N y_n^2 \right)^{M+1}}$$

where \(T\) is the period of the impulse signal, \(M\) is the number of shifts, and \(N\) is the signal length. The filter coefficients are obtained by maximizing the correlated kurtosis. The period \(T\) can be calculated based on the sampling frequency \(f_s\) and the fault characteristic frequency \(f_i\):

$$T = \frac{f_s}{f_i}$$

Since the selection of the filter length \(L\) and the number of shifts \(M\) greatly affects the enhancement effect, I used the Dung Beetle Optimizer (DBO) to determine the optimal combination of \(L\) and \(M\). The DBO algorithm is inspired by the behaviors of dung beetles, including rolling, dancing, foraging, breeding, and stealing. The population is divided into four sub-populations, and each sub-population performs a different search strategy. The positions are updated according to a set of equations. In this study, the sample entropy is used as the fitness function, and the objective is to minimize the sample entropy. Sample entropy is a measure of signal complexity and is defined as:

$$SampEn(m, r, N) = -\ln \left( \frac{A^m(r)}{B^m(r)} \right)$$

where \(A^m(r)\) is the number of template vector pairs having a distance smaller than the tolerance \(r\) in dimension \(m+1\), and \(B^m(r)\) is the number of matching vector pairs in dimension \(m\). The smaller the sample entropy, the more regular the signal, which implies a better enhancement of fault impulses.

To validate the effectiveness of the proposed signal processing method, I constructed a simulated signal model:

$$x(t) = \sum_i A h(t – iT_p) + n(t)$$

where \(h(t) = \exp(-Ct) \cos(2\pi f_n t)\) is the impulse response function, \(f_n = 4000\) Hz is the resonance frequency, \(C = 700\) is the attenuation coefficient, \(T_p = 1/120\) s is the impulse period corresponding to a fault frequency of 120 Hz, and \(n(t)\) is white Gaussian noise with a signal-to-noise ratio of \(-14\) dB. The sampling frequency is 16 kHz. The simulation results showed that the proposed method successfully recovered the fault impulses and made the fault frequency and its harmonics clearly detectable in the envelope spectrum. The DBO optimization curve demonstrated fast convergence, and the best parameter combination was found to be \(L=243\), \(M=1\).

Parameter Range Optimal Value
Filter length \(L\) [100, 500] 243 (simulation) / 273 (experiment)
Number of shifts \(M\) [1, 7] 1
Population size 100
Maximum iterations 10 6 (simulation) / 4 (experiment)
Fitness function Sample entropy Minimum value

5. Multi-Domain Feature Extraction and Feature Fusion

After the signal preprocessing stage, I proceeded to extract fault features from multiple domains. A single type of feature may not be sufficient to characterize different fault modes accurately. Therefore, I extracted a total of 57 features from the time domain, frequency domain, and time-frequency domain. In the time domain, 15 statistical features were extracted, including standard deviation, variance, peak value, root mean square, crest factor, clearance factor, impulse factor, skewness, kurtosis, and so on. In the frequency domain, 10 features were calculated, including mean frequency, root mean square frequency, standard deviation frequency, power spectral density, sample entropy, approximate entropy, fuzzy entropy, singular entropy, and center of gravity frequency. In the time-frequency domain, a four-layer wavelet packet decomposition was performed to obtain 32 features. The table below lists the features extracted from the time and frequency domains.

Domain Extracted Features
Time Domain Clearance factor, standard deviation, maximum value, minimum value, mean value, peak value, variance, crest factor, skewness, peak-to-peak value, root mean square, pulse factor, margin factor, kurtosis, pulse count
Frequency Domain Mean frequency, spectral entropy, power spectral density, root mean square frequency, frequency standard deviation, singular entropy, sample entropy, center-of-gravity frequency, fuzzy entropy, approximate entropy

However, some features contain redundant information that may reduce classification performance and increase computational complexity. To eliminate redundant features and retain the most discriminative ones, I used the Random Forest (RF) algorithm for feature selection. RF is an ensemble learning method based on decision trees. It can evaluate the contribution of each feature by computing the Gini index. The Gini index is defined as:

$$G = \sum_{c=1}^{C} P_c (1 – P_c) = 1 – \sum_{c=1}^{C} P_c^2$$

where \(P_c\) is the proportion of samples belonging to class \(c\) and \(C\) is the total number of classes. The contribution of the \(j\)-th feature is calculated from the average decrease of the Gini index caused by the feature at all splitting nodes. In this study, the RF algorithm was trained with 60% of the data and tested with 40% of the data. The number of decision trees was set to 1000. After ranking all 57 features, I selected nine features with the highest contribution: clearance factor, standard deviation, peak value, root mean square value, amplitude factor, singular entropy, quantile (0.75), energy, and energy entropy. The RF model accuracy with the full set of 57 features was 99.41%, while the model with the selected 9 features achieved an accuracy of 99.58%. Meanwhile, the time consumption was reduced by 10%. This proved the effectiveness of the RF feature selection method.

After feature selection, I further adopted Principal Component Analysis (PCA) to fuse the selected features. PCA is a linear transformation technique that projects the data onto a lower-dimensional subspace while maximizing the variance. The original feature set is first normalized using the z-score method:

$$X_{std} = \frac{X – \mu}{\sigma}$$

The covariance matrix is then computed as:

$$C = \frac{1}{m-1} X_{std}^T X_{std}$$

Then eigenvalue decomposition is performed on the covariance matrix:

$$C v_i = \lambda_i v_i$$

where \(\lambda_i\) is the i-th eigenvalue and \(v_i\) is the corresponding eigenvector. The first \(k\) eigenvectors with the largest eigenvalues are selected to form the projection matrix \(W\). The fused feature matrix is obtained by:

$$Y = X_{std} W$$

In this study, the first two principal components accounted for 95% of the total variance, so I chose the two-dimensional feature vector as the final input to the classifier. The two-dimensional fused features showed well-separated clusters for the six health states, which demonstrated the excellent discriminative ability of the extracted features.

6. Hybrid Kernel SVM Based on DBO Optimization

Support Vector Machine (SVM) is a powerful supervised machine-learning algorithm widely used for fault classification. The basic idea of SVM is to find an optimal hyperplane that maximizes the margin between different classes. For nonlinear classification problems, SVM uses kernel functions to map the input data into a high-dimensional feature space. The four commonly used kernel functions are:

The linear kernel function:

$$K(x, x_i) = x \cdot x_i$$

The polynomial kernel function:

$$K(x, x_i) = (x \cdot x_i + 1)^d$$

The radial basis function (RBF) kernel:

$$K(x, x_i) = \exp\left( -\gamma \|x – x_i\|^2 \right)$$

The Sigmoid kernel function:

$$K(x, x_i) = \tanh(v \cdot (x \cdot x_i) + c)$$

In practice, different kernel functions have different advantages. Polynomial kernels have good global generalization ability, while RBF kernels have strong local learning ability. To achieve complementary advantages, I constructed a hybrid kernel function by combining the polynomial kernel and the RBF kernel:

$$K_{mix}(x, x_i) = \alpha K_{poly}(x, x_i) + (1 – \alpha) K_{RBF}(x, x_i)$$

where \(\alpha \in [0,1]\) is the weight coefficient. By adjusting the value of \(\alpha\), the balance between the global and local kernels can be optimized.

The classification performance of SVM is also significantly affected by the penalty parameter \(C\) and the kernel parameter \(\gamma\). In order to obtain the optimal parameter combination, I utilized the DBO algorithm for parameter optimization. The penalty parameter \(C\), the kernel function parameter \(\gamma\), and the mixing weight \(\alpha\) are the three optimization targets. The DBO algorithm iteratively searches for the parameter set that maximizes the classification accuracy. The flowchart of the DBO-based hybrid kernel SVM fault diagnosis is described as follows:

Step 1: Input the fused two-dimensional feature vectors;

Step 2: Determine the search ranges of \(C\), \(\gamma\), and \(\alpha\);

Step 3: Initialize the DBO population within the search space;

Step 4: For each individual in the population, train the hybrid kernel SVM model and calculate the classification accuracy as the fitness value;

Step 5: Update the position of each dung beetle according to its behavior;

Step 6: Repeat steps 4 and 5 until the maximum number of iterations is reached;

Step 7: Output the best parameter combination and train the final model.

In order to verify the effectiveness of the DBO algorithm, I compared the unoptimized SVM model with the DBO-optimized model. The classification results are discussed in the experimental section.

7. Experimental Setup

To validate the proposed fault diagnosis method, I designed and constructed a gear fault diagnosis experimental platform. The platform consists of a mechanical system, a hardware circuit, and an upper computer software system. The mechanical part includes a DC brushless motor, a gearbox containing a spur gear pair, couplings, shafts, and a rotational speed sensor. The gear parameters used in the experiment are shown in the following table:

Parameter Driving Gear Driven Gear
Number of teeth 24 22
Module (mm) 2.54 2.54
Helix angle (deg) 0 0
Pressure angle (deg) 17.5 17.5
Root circle diameter (mm) 55.41 55.41
Pitch circle diameter (mm) 60.96 55.88
Base circle diameter (mm) 58.14 53.30
Shaft hole diameter (mm) 15.80 16.04
Profile shift coefficient 0.557 0.557

The hardware circuit design includes the STC15W4K60S4 microcontroller unit, the ADXL345 accelerometer sensor module, the WS5500 Ethernet communication module, and the W25Q128FV Flash memory chip. The motor speed controller is the ZM-6405E brushless driver. The upper computer software was developed using Qt, providing functions of user login, data acquisition parameter configuration, data storage, real-time display, and data analysis. The gear rotational speed during the experiment was 880 r/min, and the sampling frequency was 5120 Hz. Vibration signals were collected using the 352C33 ICP piezoelectric acceleration sensor. A total of six gear health states were tested: healthy gear, tooth wear, tooth pitting, tooth breakage, wear combined with pitting, and wear combined with tooth breakage. For each health state, 120 samples were collected, and the total number of samples was 720. The training-to-testing ratio was set to 2:1. The dataset description is shown in the following table:

Health Condition Training Samples Testing Samples Class Label
Healthy gear 80 40 1
Tooth wear 80 40 2
Tooth pitting 80 40 3
Tooth breakage 80 40 4
Wear plus pitting 80 40 5
Wear plus breakage 80 40 6

8. Experimental Results and Analysis

8.1 Signal Preprocessing Results

I first applied the proposed signal processing method to the collected vibration signal of the gear with tooth wear. In the original time-domain waveform, the signal was very noisy, and no obvious fault impulses could be identified. In the frequency spectrum, the fault frequency and its harmonics were almost invisible, and in the time-frequency representation, the signal was contaminated by dense noise components. After the SVD-based noise reduction, the optimal decomposition layer number was determined as \(K=10\). The VMD algorithm decomposed the signal into 10 IMFs. The correlation coefficients and kurtosis values of all IMFs were calculated, and the criterion of correlation coefficient greater than the mean value of 0.3388 and kurtosis greater than 3 was used for IMF selection. The following table shows the correlation coefficients and kurtosis values of each IMF:

IMF Index Correlation Coefficient Kurtosis
IMF 1 0.1551 2.4361
IMF 2 0.2800 3.7678
IMF 3 0.3787 3.1598
IMF 4 0.3796 13.1226
IMF 5 0.5182 10.6509
IMF 6 0.4658 12.2336
IMF 7 0.3810 12.8252
IMF 8 0.3424 12.2671
IMF 9 0.2669 9.3492
IMF 10 0.2206 14.6701

According to the selection criteria, IMF3, IMF4, IMF5, IMF6, IMF7, and IMF8 were preserved for signal reconstruction. After signal reconstruction, the fault impulses became more visible in the time-domain waveform. The period \(T\) of MCKD was calculated as 155 based on the fault characteristic frequency. The DBO algorithm was used to optimize the filter length \(L\) and the shift number \(M\). The optimal parameter combination was \(L=273\), \(M=1\). After applying the optimized MCKD to the reconstructed signal, the fault characteristic frequency at 322 Hz and its harmonics were significantly enhanced in the envelope spectrum, demonstrating the effectiveness of the proposed signal processing approach.

8.2 Feature Extraction and Fusion Results

After signal processing, 57 original features were extracted from the time domain, frequency domain, and time-frequency domain. The RF algorithm was used to rank the feature contributions. The nine features with the highest contributions were selected. The RF model accuracy and time consumption before and after feature selection were compared. The results are shown in the following table:

Feature Set Accuracy (%) Time Consumption (s)
All 57 features 99.41
Selected 9 features 99.58 10% reduction

PCA was then applied to the selected feature subset. The first two principal components contributed 95% of the cumulative variance. The two-dimensional fused features formed distinguishable clusters in the feature space, which allowed the SVM classifier to identify different gear health states accurately.

8.3 Classification Results of Different Kernel Functions

To verify the superiority of the hybrid kernel SVM, I compared it with SVM models using three single kernel functions (Sigmoid, polynomial, RBF). All models were trained and tested using the same two-dimensional fused features. The diagnostic accuracy and time consumption are shown in the following table:

Kernel Function Accuracy (%) Time (s)
SVM (Sigmoid) 57.92 60.28
SVM (Polynomial) 83.75 85.53
SVM (RBF) 87.91 76.73
SVM (Hybrid) 91.67 89.06

From the table, it can be observed that the hybrid kernel SVM achieved the highest classification accuracy of 91.67%, outperforming all three single-kernel SVM models. Although the hybrid kernel consumed slightly more time, its accuracy improvement was significant. This validates the advantage of combining global and local kernels for gear fault classification.

8.4 Comparison between SVM and DBO-SVM

After verifying the effectiveness of the hybrid kernel, I further used the DBO algorithm to optimize the parameters of the SVM model, including the penalty parameter \(C\), the kernel width parameter \(\gamma\), and the mixing coefficient \(\alpha\). The optimized model was compared with the unoptimized hybrid kernel SVM. The diagnostic results are shown in the following table:

Diagnostic Method Accuracy (%) Time (s)
Improved VMD + SVM 87.21 100.28
Improved VMD + DBO-SVM 97.08 80.56

It can be observed that the DBO-optimized SVM achieved a higher diagnostic accuracy of 97.08% and a shorter diagnostic time of 80.56 seconds, compared to the unoptimized SVM model, which achieved an accuracy of 87.21% and took 100.28 seconds. This indicates that the DBO algorithm effectively found better hyperparameters and improved the generalization capability of the SVM model, which is of great importance for early weak gear fault diagnosis under non-stationary conditions. To further investigate the classification details, I selected 20 samples from each of the six gear health states for comparison. The DBO-SVM model correctly classified all the test samples, while the unoptimized SVM produced several misclassifications, particularly confusing the compound fault states with the single fault states. These results further demonstrate the superiority of the proposed DBO-optimized hybrid kernel SVM in distinguishing different fault types of spur gears.

9. Conclusion

In this paper, I presented a systematic study on vibration-based fault diagnosis methods for spur gears under non-stationary operating conditions. The main conclusions of this work are summarized as follows:

(1) A signal preprocessing method combining the singular value kurtosis difference spectrum-based optimal VMD decomposition layer selection and the DBO-optimized MCKD was proposed. The experimental results proved that this method can effectively suppress noise, enhance weak fault impulses, and successfully extract the fault characteristic frequency of spur gears. The proposed optimal \(K\) selection algorithm can adaptively determine the number of decomposition layers without manual experience, avoiding both under-decomposition and over-decomposition. Moreover, the DBO algorithm can efficiently find the optimal filter length and shift number of the MCKD algorithm, which solves the difficulty of parameter selection.

(2) A multi-domain feature extraction and fusion framework was established. By extracting 57 features from the time, frequency, and time-frequency domains, and then applying RF-based feature selection and PCA-based feature fusion, I reduced the feature dimension to a two-dimensional feature vector without losing discriminative information. The accuracy of the RF model improved from 99.41% to 99.58% after feature selection, while the time consumption was reduced by 10%. This confirms that the selected features have strong representative ability and the fused features are well-separable for different health states of spur gears.

(3) A hybrid kernel support vector machine optimized by the DBO algorithm was proposed for gear fault pattern recognition. By combining the local and global advantages of polynomial and RBF kernels, the hybrid kernel achieved higher classification accuracy than any single kernel. The DBO algorithm effectively optimized the penalty parameter \(C\), the kernel parameter \(\gamma\), and the mixing weight \(\alpha\), enabling the SVM to adaptively match the data characteristics. The final diagnostic accuracy reached 97.08% on the experimental platform, demonstrating the feasibility and superiority of the proposed method for spur gear fault diagnosis. Compared with the unoptimized SVM, both the diagnostic accuracy and the diagnostic time were significantly improved.

In summary, this study provides a complete solution for the early weak fault diagnosis of spur gears, from signal preprocessing, feature extraction, and feature fusion to fault classification. The proposed method achieves high diagnostic accuracy and efficiency, offering reliable technical support for health monitoring and condition-based maintenance of gear transmission systems.

Scroll to Top