Semantic Aware Text Summarization Through LaBSE Feature Learning with Omni Score Classification
DOI:
https://doi.org/10.62643/Abstract
The need for text summarization has become a significant task in many real-world applications such as news monitoring, academic literature review, business analytics, legal document analysis, healthcare information management, and digital content recommendation. To address these drawbacks, this study introduces a new text summarization system built on several benchmark datasets, such as BBC and XSUM, to achieve a more general approach and performance for various document types and styles. The framework starts with the Natural Language Processing (NLP) pre-processing phase, which involves cleaning, denoising, and normalizing the textual data and converting it into an appropriate format for further analysis. After preprocessing, the Language agnostic BERT Sentence Encoder (LaBSE) is used to produce high-dimensional semantic vector representations of both sentences that accurately represent the linguistic and textual relationships. Sentiment Named Entity Recognition (SNER) Feature Ranking is proposed to better select the important features of the information, which combines sentiment characteristics and named entity information while identifying and ranking the features. Lastly, the Weighted Omni Score (WOS) Classifier assesses and ranks the extracted content in terms of its informational value, resulting in the production of coherent, concise and contextually appropriate summaries. The proposed framework incorporates the integration of semantic understanding, sentiment appreciation, and intelligent feature prioritization to enhance the relevance of the summaries, the information retained, the readability, and summarization effectiveness in various types of textual datasets.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.













