EMOTION DETECTION USING SPEECH RECOGNITION AND FACIAL EXPRESSION
DOI:
https://doi.org/10.62643/Abstract
Emotion detection is an important research area in Artificial Intelligence, Machine Learning, and Human–Computer Interaction, as understanding human emotions can enable computers to interact with users in a more natural and intelligent manner. This project, “Emotion Detection Using Speech Recognition and Facial Expression,” proposes a multimodal emotion recognition system that combines speech and facialexpression information to improve emotion classification accuracy and reliability. Speech signals are processed to extract emotional features such as pitch, tone, intensity, frequency, speech rate, and Mel-Frequency Cepstral Coefficients (MFCCs), while facial images are processed to identify visual features such as facial landmarks, eye movements, eyebrow positions, and mouth expressions. Deep learning techniques, particularly Convolutional Neural Networks (CNN) and Long Short-Term Memory (LSTM) networks, are applied for feature learning and emotion classification. The extracted information from both modalities is fused to identify emotional states such as happiness, sadness, anger, fear, surprise, disgust, and neutrality. The proposed multimodal approach reduces the limitations of single-modal systems caused by background noise, illumination variations, and individual differences. The system is designed to support real-time and human-friendly interaction and can be applied in healthcare, education, security, customer service, and intelligent virtualagent applications. Keywords: Emotion Detection, Speech Recognition, Facial Expression Recognition, Multimodal Emotion Recognition, Artificial Intelligence, Machine Learning, Deep Learning, Convolutional Neural Network
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.













