Voice Tone-Based Emotion Detection Using Deep Learning: A Hybrid Transformer–CNN–BiLSTM Framework with Multi-Feature Fusion
This paper presents a hybrid deep learning framework that integrates CNN, Transformer, and BiLSTM layers with multi-feature fusion to achieve state-of-the-art speech emotion recognition performance across five benchmark datasets, demonstrating robust generalization and significant accuracy improvements over existing baselines.