Enhancing Cloud Job Failure Prediction with a Novel Multilayer Voting-Based Framework
DOI:
https://doi.org/10.62643/Abstract
To achieve high performance, efficient resource utilization and enhanced tolerance to failures, it is essential for the cloud data center to correctly predict job failures. The purpose of this paper is to offer a novel multilayer ensemble-based cloud task failure prediction methodology based upon the Google Cluster 2019 dataset. This framework contains many classification algorithms such as DT, KNN, ANN, Extreme Gradient Boosting and Adaptive Boosting. These models are combined in a hard Voting Classifier to increase the stability and robustness of the prediction. Furthermore, a Stacking Classifier is constructed with the base learners: Random Forest, KNN and MLP and the meta estimator: LR to improve the predictive accuracy further. The experimental results demonstrate a very good performance. The Voting Classifier has 99.98% accuracy and the Stacking Classifier has 100% accuracy in forecasting the cloud job results. The use of explainable AI methods, such as LIME and SHAP, ensures transparency and reliability, and helps to explain the prediction and promote the value of the features. The trained models are integrated with a Flask powered web application for real-world deployment, which includes secure user signin and sign up using SQLite and real-time processing and presentation of the user inputs. The system delivers an understandable judgement of “job will successfully complete” or “job failure predicted” for reliable, interpretable and user friendly cloud job failure prediction support.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.













