Implementasi Multimodal Emotion Recognition Menggunakan CNN-BiLSTM-Attention pada Audio dan Mobilenetv2 pada Ekspresi Wajah


  • Ai Solihah * Mail Universitas Nusa Putra, Sukabumi, Indonesia
  • Alun Sujjada Universitas Nusa Putra, Sukabumi, Indonesia
  • Ivana Lucia Kharisma Universitas Nusa Putra, Sukabumi, Indonesia
  • (*) Corresponding Author
Keywords: Multimodal Emotion Recognition; CNN-BiLSTM-Attention; MobileNetV2; Late Fusion; Facial Expression Recognition; Audio Processing; Deep Learning

Abstract

The advancement of artificial intelligence technology has significantly contributed to the development of emotion recognition systems capable of understanding human emotional states more accurately. However, unimodal approaches that rely on a single data source, such as speech or facial expressions, still face limitations when dealing with dynamic environmental conditions. This study aims to implement a Multimodal Emotion Recognition system by integrating audio and visual modalities to improve emotion detection accuracy. The audio modality is processed using a CNN-BiLSTM-Attention architecture with Mel-Spectrogram representations as input, while the visual modality is analyzed using MobileNetV2 to recognize facial expressions. The integration of both modalities is performed using the Late Fusion method at the decision level through a weighted fusion approach, assigning weights of 0.4 and 0.6 to the audio and visual modalities, respectively. This research employs an experimental method with a quantitative approach. The research stages include literature review, data collection, audio and video preprocessing, model design, model training, multimodal integration, and system evaluation. Audio data are transformed into Mel-Spectrogram features, while video data are processed using MediaPipe Face Detection to extract facial regions prior to emotion classification. The proposed system is designed to recognize three emotional classes: neutral, happy, and angry. Model performance is evaluated using a Confusion Matrix, accuracy, precision, recall, and F1-score metrics. The results indicate that multimodal integration using the Late Fusion method effectively leverages the strengths of each modality, resulting in more accurate and robust emotion detection compared to unimodal approaches. Furthermore, the developed system is capable of performing real-time emotion classification through a camera and microphone, making it potentially applicable in various domains, including education, customer service, human-computer interaction systems, and user emotional state monitoring. The primary contribution of this research is the development of a multimodal emotion recognition system that integrates a CNN-BiLSTM-Attention model for the audio modality and MobileNetV2 for the visual modality using a late fusion method. The proposed system leverages emotional information from both modalities simultaneously to enhance the accuracy and stability of real-time emotion detection, thereby offering an alternative solution for the development of AI-based human-computer interaction systems.

Downloads

Download data is not yet available.

References

Almulla, M. A. (2025). A multimodal emotion recognition system using deep convolution neural networks. Journal of Engineering Research, 13(2), 721–729. https://doi.org/10.1016/j.jer.2024.03.021

Ardiansyah, I., Huda, F. Al, & Yudistira, N. (2024). Pengembangan Aplikasi Mobile Klasifikasi Ekspresi Wajah Manusia pada Platform Android Menggunakan Arsitektur MobileNet. 1(1).

Boitel, E., Mohasseb, A., & Haig, E. (2025). MIST : Multimodal emotion recognition using DeBERTa for text , Semi-CNN for speech , ResNet-50 for facial , and 3D-CNN for motion analysis. Expert Systems With Applications, 270(April 2024), 126236. https://doi.org/10.1016/j.eswa.2024.126236

Budiono, L. A., & Masing, M. (2022). INNOVATIVE : Volume 2 Nomor 1 Tahun 2022 Research & Learning in Primary Education Emosi Dalam Perspektif Lintas Budaya. 2(1995), 579–584.

Cai, X., Yuan, J., Zheng, R., Huang, L., & Church, K. (2021). Speech Emotion Recognition with Multi-task Learning. 4508–4512.

Caldwell, J. E., & Beaulieu, M. (2024). Cross-Modal Knowledge Transfer for Handling Missing Modalities in Affective Computing.

Dhimas, D., Putra, P., Anaga, G. K., Fitriyana, W. T., & Informatika, T. (2024). Implementasi Algoritma Convolutional Neural Network Arsitektur Mobilenetv2 Untuk Klasifikasi Ekspresi Wajah Pada Dataset FER. 3, 291–297.

Ekspresi, D. A. N., & Sosial, D. A. N. R. (2025). Kecerdasan Emosional Perempuan :

Februari, V. N., Hidayat, M. M., Darell, E., Genio, G., Astutik, A. D., & Taufik, A. (2025). Dike Jurnal Ilmu Multidisiplin Analisis Ekspresi Wajah Untuk Deteksi Emosi Dan Menerapkannya Dalam Game Dan Survei Kepuasan Pengguna Dike Jurnal Ilmu Multidisiplin. 3, 31–35.

Guntoro, A. L. S., Julianto, E., & Budiyanto, D. (2022). Pengenalan Ekspresi Wajah Menggunakan Convolutional Neural Network. 155–160.

Hadi, M. N., & Sari, R. W. (2025). Multi-Label Emotion Detection for Mental Health Monitoring Using Deep CNN and Visual Attention. 5(2), 597–605. https://doi.org/10.30811/jaise.v5i2.6961

Islam, M., Nooruddin, S., Karray, F., & Muhammad, G. (2024). Enhanced multimodal emotion recognition in healthcare analytics : A deep learning based model-level fusion approach. Biomedical Signal Processing and Control, 94(November 2023), 106241. https://doi.org/10.1016/j.bspc.2024.106241

Kamal, A. (2025). Pengenalan Emosi Oleh AI dalam komunikasi dan pengambilan keputusan . Kemampuan untuk mengenali sosial dan profesional . Dengan perkembangan teknologi kecerdasan buatan menarik dan memiliki berbagai aplikasi potensial 1 . Pengenalan emosi oleh AI melibatk. 1(1), 1–9.

Khubrani, M. (2026). Article in Press Multimodal Stress-Aware Ensemble Learning fusion for real-time proxy-based stress recognition integrating typing , facial , and voice IN IN.

Li, Y., Wei, J., Liu, Y., Kauttonen, J., & Zhao, G. (2022). Deep Learning for Micro-Expression Recognition : A Survey. IEEE Transactions on Affective Computing, 13(4), 2028–2046. https://doi.org/10.1109/TAFFC.2022.3205170

Liu, L., Luo, Q., Zhang, W., Zhang, M., & Zhai, B. (2025). Multimodal emotion recognition method in complex dynamic scenes. Journal of Information and Intelligence, 3(3), 257–274. https://doi.org/10.1016/j.jiixd.2025.02.004

Mehendale, N. (2020). Facial emotion recognition using convolutional neural networks ( FERC ). SN Applied Sciences, 2(3), 1–8. https://doi.org/10.1007/s42452-020-2234-1

N.A., M. I. (2025). Pengembangan Speech Emotion Recognition. 1–2.

Nurjihan, S. W., Nurbadillah, N., Faturrahman, N., Wiguna, I. M., Lasardi, E. M., Giri, E. P., & Mindara, G. P. (2024). indonesia Pengenalan Pola Ekspresi Wajah Untuk Pengolahan Citra Menggunakan Metode Convolutional Neural Network. JATISI, 11(4).

Razaqa, D., Rezka, M., Maghribi, A., Gunasti, N. D., & Wati, T. (2024). Analisis Etika dan Dampak Penggunaan Sistem Pengenalan Wajah untuk Manajemen Kehadiran di Lingkungan Sekolah. 4221, 50–57.

Riyandi, R., & Sumarsono, S. (2025). Analisis Efektivitas Fusi Fitur Multimodal dalam Klasifikasi Citra Daun Herbal. Jurnal Teknik Informatika Dan Sistem Informasi, 11(3), 463–474. https://doi.org/10.28932/jutisi.v11i3.12262

Roihan, A., Dwi, R., & Astriyani, E. (2025). Analisis Konseptual Fusi Multimodal Wajah dan Visual Speech untuk Autentikasi Biometrik Non-Vokal terhadap Deepfake. Journal of Big Data Analytic and Artificial Intelligence, 8(2), 54–61. https://doi.org/10.71302/jbidai.v8i2.84

S Budiawan, ST Alvianus Dengen, R. A. (2025). Dialog Digital: Memahami Komunikasi Manusia Dan Mesin Dalam Era Interkoneksi.

Shane, K., Prakoso, B., & Prasetio, B. H. (2025). Implementasi Speech Recognition berbasis Raspberry Pi 5 pada Ekosistem Smart-Home menggunakan Algoritma Gated Recurrent Unit ( GRU ). 9(3), 1–10.

Singla, C., Singh, S., Sharma, P., Mittal, N., & Gared, F. (2024). Emotion recognition for human – computer interaction using high level descriptors. Scientific Reports, 1–12. https://doi.org/10.1038/s41598-024-59294-y

Sutarti, & Fariza Syaqialloh. (2025). Klasifikasi dan Pengenalan Emosi dari Ekspresi Wajah Menggunakan CNN-BiLSTM dengan Teknik Data Augmentation. Decode: Jurnal Pendidikan Teknologi Informasi, 5(1), 79–91. https://doi.org/10.51454/decode.v5i1.1038

Syaqialloh, F. (2025). Klasifikasi dan Pengenalan Emosi dari Ekspresi Wajah Menggunakan CNN-BiLSTM dengan Teknik Data Augmentation. 5(1), 79–91.

Wijaya, H. (2024). Teknologi Pengenalan Suara tentang Metode , Bahasa dan Tantangan : Systematic Literature Review. 7(2). https://doi.org/10.32877/bt.v7i2.1888

Wu, R., Wang, H., Chen, H., & Carneiro, G. (n.d.). Deep Multimodal Learning with Missing Modality: A Survey. 1(1).


Bila bermanfaat silahkan share artikel ini

Berikan Komentar Anda terhadap artikel Implementasi Multimodal Emotion Recognition Menggunakan CNN-BiLSTM-Attention pada Audio dan Mobilenetv2 pada Ekspresi Wajah

Dimensions Badge
Article History
Published: 2026-07-22
Abstract View: 25 times
PDF Download: 18 times
Issue
Section
Articles