Klasifikasi Sentimen Publik terhadap Isu Toleransi Agama Menggunakan Algoritma Random Forest
Abstract
Social media, particularly Twitter, has evolved into a dynamic arena for discussions on religious issues in Indonesia. Interfaith tolerance is one of the topics that most frequently elicits a wide range of responses, from support to hate speech. This study designs a five-class sentiment scheme, consisting of Positive/Neutral, Neutral-Abusive, and Negative tweets divided into three intensity levels (Weak, Moderate, and Strong), and classifies them using the Random Forest algorithm. The dataset used is the Indonesian Abusive and Hate Speech Twitter Text available on the Kaggle platform, consisting of 13,169 tweets with dual labels. Sentiment labels were created based on a combination of the HS and HS_Religion columns and hate speech intensity levels: Weak, Moderate, and Strong. Tweets without hate speech and unrelated to religion are considered positive or neutral, while tweets with HS_Religion=1 are classified as negative and grouped into three intensity levels. Prior to modeling, the text undergoes slang normalization, removal of inappropriate words, Nazief-Adriani stemming, and feature extraction using TF-IDF bigrams. Results from 10-fold cross-validation show an accuracy of 66.0%, macro precision of 52.1%, macro recall of 56.1%, and macro F1-Score of 51.3%, comparable to SVM (F1 52.7%) and Naive Bayes (31.8%), with differences between models assessed statistically using the McNemar test.
Downloads
References
M. L. Tewu, D. Destine, and I. Gunawan, “Analysis of Social Media User Growth and Its Implications for Digital Marketing Strategies in Indonesia 2024,” International Journal of Management Studies and Social Science Research (IJMSSSR), vol. 7, no. 3, pp. 236–245, 2025, doi: 10.56293/IJMSSSR.2025.5623.
F. Ihsan, I. Iskandar, N. S. Harahap, and S. Agustian, “Algoritme Decision Tree untuk Mendeteksi Ujaran Kebencian dan Bahasa Kasar Multilabel pada Twitter Berbahasa Indonesia,” Jurnal Teknologi dan Sistem Komputer, vol. 9, no. 4, pp. 199–204, 2021, doi: 10.14710/jtsiskom.2021.13907.
M. O. Ibrohim and I. Budi, “Multi-Label Hate Speech and Abusive Language Detection in Indonesian Twitter,” in Proceedings of the Third Workshop on Abusive Language Online, 2019, pp. 46–57. doi: 10.18653/v1/W19-3506.
S. Cahyawijaya et al., “NusaCrowd: Open Source Initiative for Indonesian NLP Resources,” in Findings of the Association for Computational Linguistics: ACL 2023, 2023, pp. 13745–13818. doi: 10.18653/v1/2023.findings-acl.868.
C.-H. Lin and U. Nuha, “Sentiment Analysis of Indonesian Datasets Based on a Hybrid Deep-Learning Strategy,” J. Big Data, vol. 10, no. 1, p. 88, 2023, doi: 10.1186/s40537-023-00782-9.
A. F. Aji et al., “One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 7226–7249. doi: 10.18653/v1/2022.acl-long.500.
F. Koto, J. H. Lau, and T. Baldwin, “IndoBERTweet: A Pretrained Language Model for Indonesian Twitter with Effective Domain-Specific Vocabulary Initialization,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021, pp. 10660–10668. doi: 10.18653/v1/2021.emnlp-main.833.
A. D. Sanya and L. H. Suadaa, “Handling Imbalanced Dataset on Hate Speech Detection in Indonesian Online News Comments,” in 2022 10th International Conference on Information and Communication Technology (ICoICT), IEEE, 2022, pp. 380–385. doi: 10.1109/ICoICT55009.2022.9914883.
A. P. J. Dwitama, D. H. Fudholi, and S. Hidayat, “Indonesian Hate Speech Detection Using Bidirectional Long Short-Term Memory (Bi-LSTM),” Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi), vol. 7, no. 2, pp. 302–309, 2023, doi: 10.29207/resti.v7i2.4642.
L. Wang, M. Han, X. Li, N. Zhang, and H. Cheng, “Review of Classification Methods on Unbalanced Data Sets,” IEEE Access, vol. 9, pp. 64606–64628, 2021, doi: 10.1109/ACCESS.2021.3074243.
G. James, D. Witten, T. Hastie, R. Tibshirani, and J. Taylor, “‘Statistical Learning,’ in An Introduction to Statistical Learning: With Applications in Python,” in An introduction to statistical learning: With applications in Python, Springer, 2023, pp. 15–67. doi: 10.1007/978-3-031-38747-0_2.
A. Géron, Hands-On Machine Learning with Scikit-Learn, Keras, and Tensorflow, 3rd ed. "O’Reilly Media, Inc., 2022.
I. D. Mienye and Y. Sun, “A Survey of Ensemble Learning: Concepts, Algorithms, Applications, and Prospects,” IEEE Access, vol. 10, pp. 99129–99149, 2022, doi: 10.1109/ACCESS.2022.3207287.
A. Tharwat, “Classification Assessment Methods,” Applied Computing and Informatics, vol. 17, no. 1, pp. 168–192, 2021, doi: 10.1016/j.aci.2018.08.003.
O. Rainio, J. Teuho, and R. Klén, “Evaluation Metrics and Statistical Tests for Machine Learning,” Scientific Reports (Sci.Rep.), vol. 14, no. 1, p. 6086, 2024, doi: 10.1038/s41598-024-56706-x.
D. Elreedy, A. F. Atiya, and F. Kamalov, “A Theoretical Distribution Analysis of Synthetic Minority Oversampling Technique (SMOTE) for Imbalanced Learning,” Mach. Learn., vol. 113, no. 7, pp. 4903–4923, 2024, doi: 10.1007/s10994-022-06296-4.
H. Zhou, “Research of Text Classification Based on TF-IDF and CNN-LSTM,” J. Phys. Conf. Ser., vol. 2171, no. 1, p. 012021, 2022, doi: 10.1088/1742-6596/2171/1/012021.
N. A. Semary, W. Ahmed, K. Amin, P. Pławiak, and M. Hammad, “Enhancing Machine Learning-Based Sentiment Analysis Through Feature Extraction Techniques,” PLoS One, vol. 19, no. 2, p. e0294968, 2024, doi: 10.1371/journal.pone.0294968.
M. Aria, C. Cuccurullo, and A. Gnasso, “A Comparison Among Interpretative Proposals for Random Forests,” Machine Learning with Applications, vol. 6, p. 100094, 2021, doi: 10.1016/j.mlwa.2021.100094.
Bila bermanfaat silahkan share artikel ini
Berikan Komentar Anda terhadap artikel Klasifikasi Sentimen Publik terhadap Isu Toleransi Agama Menggunakan Algoritma Random Forest
Pages: 21-29
Copyright (c) 2026 Flienschy Faith Maxy Tamaka, Theresia Sheren Medea, Ivana Julia Poli

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under Creative Commons Attribution 4.0 International License that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (Refer to The Effect of Open Access).













