Analisis Sentimen Ulasan Pengguna Halodoc Menggunakan TextCNN pada Dataset Tidak Seimbang Berbasis NLP


  • Fitriyani Fitriyani * Mail Universitas Esa Unggul, Jakarta, Indonesia
  • Budi Tjahjono Universitas Esa Unggul, Jakarta, Indonesia
  • (*) Corresponding Author
Keywords: Sentiment Analysis; TextCNN; Natural Language Processing; Digital Healthcare Services; Halodoc

Abstract

This study analyzes user sentiment in Halodoc application reviews using a TextCNN architecture on an Indonesian-language dataset with extreme class imbalance. The dataset consists of 2,000 reviews collected from the Google Play Store and manually labeled into three classes: positive, negative, and neutral. The class distribution comprises 1,800 positive reviews (90%), 175 negative reviews (8.75%), and 25 neutral reviews (1.25%). To reduce bias toward the majority class, the TextCNN model was trained using class weighting and evaluated using accuracy, precision, recall, F1-score, and a confusion matrix. The model achieved an accuracy of 93.75% and a weighted F1-score of 0.94, while the macro F1-score was only 0.57. The recall for the neutral class was 0%, indicating that the model was unable to recognize the minority class effectively. The contributions of this study are threefold: an empirical evaluation of TextCNN on Indonesian Halodoc reviews with extreme class imbalance; an analysis of the effect of class distribution using class-level metrics and error analysis; and an identification of the practical implications of the classification results for monitoring digital healthcare service quality. These findings demonstrate that high accuracy alone is insufficient to represent model performance on imbalanced datasets; therefore, strategies such as oversampling, data augmentation, or hybrid approaches should be considered in future research.

Downloads

Download data is not yet available.

References

F. Binsar, Mts. Arief, V. U. Tjhin, and I. Susilowati, “Exploring consumer sentiments in telemedicine and telehealth services: Towards an integrated framework for innovation,” Journal of Open Innovation: Technology, Market, and Complexity, vol. 11, no. 1, p. 100453, 2025, doi: https://doi.org/10.1016/j.joitmc.2024.100453.

A. Jerfy, O. Selden, and R. Balkrishnan, “The Growing Impact of Natural Language Processing in Healthcare and Public Health,” INQUIRY: The Journal of Health Care Organization, Provision, and Financing, vol. 61, p. 469580241290095, Aug. 2024, doi: 10.1177/00469580241290095.

Q. and K. C. and W. J. and L. G. and F. X. and L. X. and Y. G. Li Xiadong and Shu, “An Intelligent System for Classifying Patient Complaints Using Machine Learning and Natural Language Processing: Development and Validation Study,” J Med Internet Res, vol. 27, p. e55721, Jan. 2025, doi: 10.2196/55721.

T. G. W. M. Sidabutar and D. Juardi, “Analisis Sentimen Masyarakat Terhadap Penggunaan Halodoc Sebagai Layanan Telemedicine Di Indonesia,” Jurnal Informatika dan Teknik Elektro Terapan, vol. 13, no. 1, Jan. 2025, doi: 10.23960/jitet.v13i1.5682.

H. Imaduddin, F. Yusfida A’la, and Y. S. Nugroho, “Sentiment Analysis in Indonesian Healthcare Applications using IndoBERT Approach.” [Online]. Available: www.ijacsa.thesai.org

S. Al-Hadhrami, T. Vinko, T. Al-Hadhrami, F. Saeed, and S. Qasem, “Deep learning-based method for sentiment analysis for patients’ drug reviews,” PeerJ Comput. Sci., vol. 10, p. e1976, Aug. 2024, doi: 10.7717/peerj-cs.1976.

C. Ouni, E. Benmohamed, and H. Ltifi, “Deep learning-based Soft word embedding approach for sentiment analysis,” Procedia Comput. Sci., vol. 246, pp. 1355–1364, 2024, doi: https://doi.org/10.1016/j.procs.2024.09.720.

Y. Ye, X. Chang, and A. Fan, “NE ZHA-TextCNN Method for Multi-Label Long Text Classification,” Procedia Comput. Sci., vol. 262, pp. 313–319, 2025, doi: https://doi.org/10.1016/j.procs.2025.05.058.

H. Imaduddin, F. A’la, and Y. Nugroho, “Sentiment Analysis in Indonesian Healthcare Applications using IndoBERT Approach,” International Journal of Advanced Computer Science and Applications, vol. 14, Jan. 2023, doi: 10.14569/IJACSA.2023.0140813.

F. Gräßer, H. Malberg, S. Kallumadi, and S. Zaunseder, “Aspect-Based Sentiment Analysis of Drug Reviews Applying Cross-Domain and Cross-Data Learning,” ACM International Conference Proceeding Series, 2018, pp. 121–125, doi: 10.1145/3194658.3194677.

X. Jiang, C. Song, Y. Xu, Y. Li, and Y. Peng, “Research on Sentiment Classification for Netizens Based on the BERT-BiLSTM-TextCNN Model,” PeerJ Computer Science, vol. 8, 2022, e1005, doi: 10.7717/peerj-cs.1005.

S. Natu and M. Aparicio, “Analyzing Knowledge Sharing Behaviors in Virtual Teams: Practical Evidence from Digitalized Workplaces,” Journal of Innovation and Knowledge, vol. 7, no. 4, 2022, article 100248, doi: 10.1016/j.jik.2022.100248

T. G. W. M. Sidabutar and D. Juardi, “Analisis Sentimen Masyarakat Terhadap Penggunaan Halodoc sebagai Layanan Telemedicine di Indonesia,” Jurnal Informatika dan Teknik Elektro Terapan, vol. 13, no. 1, 2025, doi: 10.23960/jitet.v13i1.5682.

R. Stewart, J. Chaturvedi, and A. Roberts, “Natural Language Processing—Relevance to Patient Outcomes and Real-World Evidence,” Expert Review of Pharmacoeconomics & Outcomes Research, vol. 24, no. 1, 2024, pp. 5–9, doi: 10.1080/14737167.2023.2275670.

A. Rawat, M. A. Wani, M. ElAffendi, A. S. Imran, Z. Kastrati, and S. M. Daudpota, “Drug Adverse Event Detection Using Text-Based Convolutional Neural Networks (TextCNN) Technique,” Electronics (Switzerland), vol. 11, no. 20, pp. 1–21, 2022, doi: 10.3390/electronics11203336.

K. Denecke and D. Reichenpfader, “Sentiment Analysis of Clinical Narratives: A Scoping Review,” Journal of Biomedical Informatics, vol. 140, 2023, article 104336, doi: 10.1016/j.jbi.2023.104336.

P. K. Nag, A. Bhagat, R. Vishnu Priya, and D. K. Khare, “Emotional Intelligence Through Artificial Intelligence: NLP and Deep Learning in the Analysis of Healthcare Texts,” in 2023 International Conference on Artificial Intelligence for Innovations in Healthcare Industries (ICAIIHI), 2023, pp. 1–7, doi: 10.1109/ICAIIHI57871.2023.10489117.

Y. Ling, “Bio+Clinical BERT, BERT Base, and CNN Performance Comparison for Predicting Drug-Review Satisfaction,” arXiv preprint arXiv:2308.03782, 2023, doi: 10.48550/arXiv.2308.03782.

A. Ghosh, S. Umer, B. C. Dhara, and G. G. M. N. Ali, “A Multimodal Pain Sentiment Analysis System Using Ensembled Deep Learning Approaches for IoT-Enabled Healthcare Framework,” Sensors, vol. 25, no. 4, 2025, article 1223, doi: 10.3390/s25041223.

M. Prasetya, M. Wulandari, and S. A. Nikmah, “Implementasi NLP (Natural Language Processing) Dasar pada Analisis Sentiment Review Spotify,” STAIN (Seminar Nasional Teknologi & Sains), vol. 3, no. 1, 2024, pp. 145–153, doi: 10.29407/stains.v3i1.4166.

M. Prasetya, M. Wulandari, and S. A. Nikmah, “Implementasi NLP (Natural Language Processing) Dasar pada Analisis Sentiment Review Spotify,” Stains (Seminar Nasional Teknologi & Sains), vol. 3, no. 1, pp. 145–153, 2024.

D. Suhartono, K. Purwandari, N. H. Jeremy, S. Philip, P. Arisaputra, and I. H. Parmonangan, “Deep Neural Networks and Weighted Word Embeddings for Sentiment Analysis of Drug Product Reviews,” Procedia Computer Science, vol. 216, 2023, pp. 664–671, doi: 10.1016/j.procs.2022.12.182.


Bila bermanfaat silahkan share artikel ini

Berikan Komentar Anda terhadap artikel Analisis Sentimen Ulasan Pengguna Halodoc Menggunakan TextCNN pada Dataset Tidak Seimbang Berbasis NLP

Dimensions Badge
Article History
Submitted: 2026-02-01
Published: 2026-08-31
Abstract View: 66 times
PDF Download: 17 times
Section
Articles