Identification of Spambots and Fake Followers on Social Network via Interpretable AI-Based Machine Learning
DOI:
https://doi.org/10.62643/Abstract
Social networking platforms such as Twitter enable broad human interaction but are increasingly targeted by automated accounts that mimic human behavior, spreading misinformation and manipulating public opinion. Detecting such spambots is critical for maintaining information integrity, yet many conventional approaches rely on opaque, black-box models, limiting transparency and interpretability. This study employs the Cresci-15 and Cresci-17 datasets to investigate interpretable machine learning techniques for identifying spambots and fake followers. Both feature-based and text-based data are utilized, with preprocessing steps including normalization, tokenization, and removal of irrelevant content. Recursive Feature Elimination (RFE) reduces feature dimensionality, while resampling strategies such as SMOTE and SMOTEENN address class imbalance. Multiple machine learning algorithms, including Decision Tree, Random Forest, SVM, XGBoost, AdaBoost, Stacking Classifier, and Voting Classifier, are evaluated. Results demonstrate that Stacking Classifier achieves superior performance, reaching 99.9% accuracy on Cresci-15 and 99.5% accuracy on Cresci-17 datasets. Additionally, explainable AI methods such as LIME and SHAP provide clear insights into feature importance, enhancing model transparency and supporting informed decision-making. These findings highlight the effectiveness of combining feature selection, advanced resampling, and ensemble learning strategies with interpretable techniques for robust detection of automated accounts in social networks.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.













