Dark Side of the Web:Dark Web Classification Based on Text CNN and Topic Modeling Weight
Abstract
The rapid expansion of the dark web has created significant challenges for monitoring and analyzing illicit activities due to its anonymous and unstructured nature. Dark web content often includes illegal trade, cybercrime discussions, and harmful communications, making effective classification essential for digital forensics and cybersecurity. However, traditional text classification methods struggle to capture the complex semantics and hidden patterns present in dark web data. This paper proposes a novel approach for dark web classification by combining Text Convolutional Neural Networks (Text CNN) with topic modeling-based weighting techniques. The proposed system leverages the feature extraction capability of Text CNN to learn local and contextual patterns from textual data, while topic modeling methods such as Latent Dirichlet Allocation (LDA) are used to assign semantic weights to words based on underlying topics. By integrating topic-based weights into the neural network framework, the model enhances its ability to focus on meaningful and domain-relevant features. Experimental results demonstrate that the hybrid approach improves classification accuracy, precision, and robustness compared to standalone deep learning or topic modeling methods. The study highlights the effectiveness of combining deep learning with probabilistic topic modeling for analyzing complex and sensitive dark web content, providing a scalable solution for law enforcement and cybersecurity applications.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.













