A Robust Transformer Framework for Identifying Large-Scale Disease Spread Patterns
DOI:
https://doi.org/10.62643/ijerst.2026.v22.n2(1).2640Keywords:
Electronic Health Records (EHR), ELECTRA (Efficiently Learning an Encoder that Classifies Token Replacements Accurately), Transformer-based Embeddings, Natural Language Processing (NLP).Abstract
The rapid growth of digital healthcare data, particularly unstructured clinical text such as electronic health records and medical transcriptions, has opened new opportunities for developing intelligent healthcare systems. These datasets contain rich information that can assist in identifying medical specialties and improving clinical workflows. However, extracting meaningful insights from such data remains challenging due to its complex structure, specialized medical terminology, and high dimensionality. This work focuses on the problem of automatically classifying clinical text into appropriate medical specialties, which is crucial for enhancing patient care, optimizing healthcare resources, and supporting accurate clinical decision-making. Traditional approaches rely on manual annotation and rule-based methods, which are time-consuming, error-prone, and lack scalability. Additionally, conventional machine learning techniques depend heavily on handcrafted features and struggle to capture deep semantic relationships within domain-specific text. These limitations become more evident when handling large-scale and imbalanced healthcare datasets. To overcome these challenges, the proposed system integrates transformer-based embeddings with advanced machine learning models. Efficiently Learning an Encoder that Classifies Token Replacements Accurately (ELECTRA) is employed to generate deep contextual representations that effectively capture semantic relationships in clinical text. To address class imbalance, the Synthetic Minority Over-sampling Technique (SMOTE) is applied for data augmentation. Multiple classifiers, including Adaptive Boosting (AB), Random Forest (RF), Tree Alternating Optimization Tree (TT), and Extra Trees (ET), are evaluated, with the Extra Trees classifier selected as the optimal model due to its superior performance.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.













