Implementasi Multimodal Emotion Recognition Menggunakan CNN-BiLSTM-Attention pada Audio dan Mobilenetv2 pada Ekspresi Wajah
Abstract
The advancement of artificial intelligence technology has significantly contributed to the development of emotion recognition systems capable of understanding human emotional states more accurately. However, unimodal approaches that rely on a single data source, such as speech or facial expressions, still face limitations when dealing with dynamic environmental conditions. This study aims to implement a Multimodal Emotion Recognition system by integrating audio and visual modalities to improve emotion detection accuracy. The audio modality is processed using a CNN-BiLSTM-Attention architecture with Mel-Spectrogram representations as input, while the visual modality is analyzed using MobileNetV2 to recognize facial expressions. The integration of both modalities is performed using the Late Fusion method at the decision level through a weighted fusion approach, assigning weights of 0.4 and 0.6 to the audio and visual modalities, respectively. This research employs an experimental method with a quantitative approach. The research stages include literature review, data collection, audio and video preprocessing, model design, model training, multimodal integration, and system evaluation. Audio data are transformed into Mel-Spectrogram features, while video data are processed using MediaPipe Face Detection to extract facial regions prior to emotion classification. The proposed system is designed to recognize three emotional classes: neutral, happy, and angry. Model performance is evaluated using a Confusion Matrix, accuracy, precision, recall, and F1-score metrics. The results indicate that multimodal integration using the Late Fusion method effectively leverages the strengths of each modality, resulting in more accurate and robust emotion detection compared to unimodal approaches. Furthermore, the developed system is capable of performing real-time emotion classification through a camera and microphone, making it potentially applicable in various domains, including education, customer service, human-computer interaction systems, and user emotional state monitoring. The primary contribution of this research is the development of a multimodal emotion recognition system that integrates a CNN-BiLSTM-Attention model for the audio modality and MobileNetV2 for the visual modality using a late fusion method. The proposed system leverages emotional information from both modalities simultaneously to enhance the accuracy and stability of real-time emotion detection, thereby offering an alternative solution for the development of AI-based human-computer interaction systems.
Downloads
References
Almulla, M. A. (2025). A multimodal emotion recognition system using deep convolution neural networks. Journal of Engineering Research, 13(2), 721–729. https://doi.org/10.1016/j.jer.2024.03.021
Ardiansyah, I., Huda, F. Al, & Yudistira, N. (2024). Pengembangan Aplikasi Mobile Klasifikasi Ekspresi Wajah Manusia pada Platform Android Menggunakan Arsitektur MobileNet. 1(1).
Boitel, E., Mohasseb, A., & Haig, E. (2025). MIST : Multimodal emotion recognition using DeBERTa for text , Semi-CNN for speech , ResNet-50 for facial , and 3D-CNN for motion analysis. Expert Systems With Applications, 270(April 2024), 126236. https://doi.org/10.1016/j.eswa.2024.126236
Budiono, L. A., & Masing, M. (2022). INNOVATIVE : Volume 2 Nomor 1 Tahun 2022 Research & Learning in Primary Education Emosi Dalam Perspektif Lintas Budaya. 2(1995), 579–584.
Cai, X., Yuan, J., Zheng, R., Huang, L., & Church, K. (2021). Speech Emotion Recognition with Multi-task Learning. 4508–4512.
Caldwell, J. E., & Beaulieu, M. (2024). Cross-Modal Knowledge Transfer for Handling Missing Modalities in Affective Computing.
Dhimas, D., Putra, P., Anaga, G. K., Fitriyana, W. T., & Informatika, T. (2024). Implementasi Algoritma Convolutional Neural Network Arsitektur Mobilenetv2 Untuk Klasifikasi Ekspresi Wajah Pada Dataset FER. 3, 291–297.
Ekspresi, D. A. N., & Sosial, D. A. N. R. (2025). Kecerdasan Emosional Perempuan :
Februari, V. N., Hidayat, M. M., Darell, E., Genio, G., Astutik, A. D., & Taufik, A. (2025). Dike Jurnal Ilmu Multidisiplin Analisis Ekspresi Wajah Untuk Deteksi Emosi Dan Menerapkannya Dalam Game Dan Survei Kepuasan Pengguna Dike Jurnal Ilmu Multidisiplin. 3, 31–35.
Guntoro, A. L. S., Julianto, E., & Budiyanto, D. (2022). Pengenalan Ekspresi Wajah Menggunakan Convolutional Neural Network. 155–160.
Hadi, M. N., & Sari, R. W. (2025). Multi-Label Emotion Detection for Mental Health Monitoring Using Deep CNN and Visual Attention. 5(2), 597–605. https://doi.org/10.30811/jaise.v5i2.6961
Islam, M., Nooruddin, S., Karray, F., & Muhammad, G. (2024). Enhanced multimodal emotion recognition in healthcare analytics : A deep learning based model-level fusion approach. Biomedical Signal Processing and Control, 94(November 2023), 106241. https://doi.org/10.1016/j.bspc.2024.106241
Kamal, A. (2025). Pengenalan Emosi Oleh AI dalam komunikasi dan pengambilan keputusan . Kemampuan untuk mengenali sosial dan profesional . Dengan perkembangan teknologi kecerdasan buatan menarik dan memiliki berbagai aplikasi potensial 1 . Pengenalan emosi oleh AI melibatk. 1(1), 1–9.
Khubrani, M. (2026). Article in Press Multimodal Stress-Aware Ensemble Learning fusion for real-time proxy-based stress recognition integrating typing , facial , and voice IN IN.
Li, Y., Wei, J., Liu, Y., Kauttonen, J., & Zhao, G. (2022). Deep Learning for Micro-Expression Recognition : A Survey. IEEE Transactions on Affective Computing, 13(4), 2028–2046. https://doi.org/10.1109/TAFFC.2022.3205170
Liu, L., Luo, Q., Zhang, W., Zhang, M., & Zhai, B. (2025). Multimodal emotion recognition method in complex dynamic scenes. Journal of Information and Intelligence, 3(3), 257–274. https://doi.org/10.1016/j.jiixd.2025.02.004
Mehendale, N. (2020). Facial emotion recognition using convolutional neural networks ( FERC ). SN Applied Sciences, 2(3), 1–8. https://doi.org/10.1007/s42452-020-2234-1
N.A., M. I. (2025). Pengembangan Speech Emotion Recognition. 1–2.
Nurjihan, S. W., Nurbadillah, N., Faturrahman, N., Wiguna, I. M., Lasardi, E. M., Giri, E. P., & Mindara, G. P. (2024). indonesia Pengenalan Pola Ekspresi Wajah Untuk Pengolahan Citra Menggunakan Metode Convolutional Neural Network. JATISI, 11(4).
Razaqa, D., Rezka, M., Maghribi, A., Gunasti, N. D., & Wati, T. (2024). Analisis Etika dan Dampak Penggunaan Sistem Pengenalan Wajah untuk Manajemen Kehadiran di Lingkungan Sekolah. 4221, 50–57.
Riyandi, R., & Sumarsono, S. (2025). Analisis Efektivitas Fusi Fitur Multimodal dalam Klasifikasi Citra Daun Herbal. Jurnal Teknik Informatika Dan Sistem Informasi, 11(3), 463–474. https://doi.org/10.28932/jutisi.v11i3.12262
Roihan, A., Dwi, R., & Astriyani, E. (2025). Analisis Konseptual Fusi Multimodal Wajah dan Visual Speech untuk Autentikasi Biometrik Non-Vokal terhadap Deepfake. Journal of Big Data Analytic and Artificial Intelligence, 8(2), 54–61. https://doi.org/10.71302/jbidai.v8i2.84
S Budiawan, ST Alvianus Dengen, R. A. (2025). Dialog Digital: Memahami Komunikasi Manusia Dan Mesin Dalam Era Interkoneksi.
Shane, K., Prakoso, B., & Prasetio, B. H. (2025). Implementasi Speech Recognition berbasis Raspberry Pi 5 pada Ekosistem Smart-Home menggunakan Algoritma Gated Recurrent Unit ( GRU ). 9(3), 1–10.
Singla, C., Singh, S., Sharma, P., Mittal, N., & Gared, F. (2024). Emotion recognition for human – computer interaction using high level descriptors. Scientific Reports, 1–12. https://doi.org/10.1038/s41598-024-59294-y
Sutarti, & Fariza Syaqialloh. (2025). Klasifikasi dan Pengenalan Emosi dari Ekspresi Wajah Menggunakan CNN-BiLSTM dengan Teknik Data Augmentation. Decode: Jurnal Pendidikan Teknologi Informasi, 5(1), 79–91. https://doi.org/10.51454/decode.v5i1.1038
Syaqialloh, F. (2025). Klasifikasi dan Pengenalan Emosi dari Ekspresi Wajah Menggunakan CNN-BiLSTM dengan Teknik Data Augmentation. 5(1), 79–91.
Wijaya, H. (2024). Teknologi Pengenalan Suara tentang Metode , Bahasa dan Tantangan : Systematic Literature Review. 7(2). https://doi.org/10.32877/bt.v7i2.1888
Wu, R., Wang, H., Chen, H., & Carneiro, G. (n.d.). Deep Multimodal Learning with Missing Modality: A Survey. 1(1).
Bila bermanfaat silahkan share artikel ini
Berikan Komentar Anda terhadap artikel Implementasi Multimodal Emotion Recognition Menggunakan CNN-BiLSTM-Attention pada Audio dan Mobilenetv2 pada Ekspresi Wajah
Pages: 1065-1080
Copyright (c) 2026 Ai Solihah, Alun Sujjada, Ivana Lucia Kharisma

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under Creative Commons Attribution 4.0 International License that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (Refer to The Effect of Open Access).













