Klasifikasi Sentimen Publik terhadap Isu Toleransi Agama Menggunakan Algoritma Random Forest


  • Flienschy Faith Maxy Tamaka * Mail Universitas Prisma, Indonesia
  • Theresia Sheren Medea Universitas Prisma, Manado, Indonesia
  • Ivana Julia Poli Universitas Prisma, Manado, Indonesia
  • (*) Corresponding Author
Keywords: Sentiment Classification; Religious Tolerance; Random Forest; Twitter Indonesia; TF-IDF

Abstract

Social media, particularly Twitter, has evolved into a dynamic arena for discussions on religious issues in Indonesia. Interfaith tolerance is one of the topics that most frequently elicits a wide range of responses, from support to hate speech. This study designs a five-class sentiment scheme, consisting of Positive/Neutral, Neutral-Abusive, and Negative tweets divided into three intensity levels (Weak, Moderate, and Strong), and classifies them using the Random Forest algorithm. The dataset used is the Indonesian Abusive and Hate Speech Twitter Text available on the Kaggle platform, consisting of 13,169 tweets with dual labels. Sentiment labels were created based on a combination of the HS and HS_Religion columns and hate speech intensity levels: Weak, Moderate, and Strong. Tweets without hate speech and unrelated to religion are considered positive or neutral, while tweets with HS_Religion=1 are classified as negative and grouped into three intensity levels. Prior to modeling, the text undergoes slang normalization, removal of inappropriate words, Nazief-Adriani stemming, and feature extraction using TF-IDF bigrams. Results from 10-fold cross-validation show an accuracy of 66.0%, macro precision of 52.1%, macro recall of 56.1%, and macro F1-Score of 51.3%, comparable to SVM (F1 52.7%) and Naive Bayes (31.8%), with differences between models assessed statistically using the McNemar test.

Downloads

Download data is not yet available.

References

M. L. Tewu, D. Destine, and I. Gunawan, “Analysis of Social Media User Growth and Its Implications for Digital Marketing Strategies in Indonesia 2024,” International Journal of Management Studies and Social Science Research (IJMSSSR), vol. 7, no. 3, pp. 236–245, 2025, doi: 10.56293/IJMSSSR.2025.5623.

F. Ihsan, I. Iskandar, N. S. Harahap, and S. Agustian, “Algoritme Decision Tree untuk Mendeteksi Ujaran Kebencian dan Bahasa Kasar Multilabel pada Twitter Berbahasa Indonesia,” Jurnal Teknologi dan Sistem Komputer, vol. 9, no. 4, pp. 199–204, 2021, doi: 10.14710/jtsiskom.2021.13907.

M. O. Ibrohim and I. Budi, “Multi-Label Hate Speech and Abusive Language Detection in Indonesian Twitter,” in Proceedings of the Third Workshop on Abusive Language Online, 2019, pp. 46–57. doi: 10.18653/v1/W19-3506.

S. Cahyawijaya et al., “NusaCrowd: Open Source Initiative for Indonesian NLP Resources,” in Findings of the Association for Computational Linguistics: ACL 2023, 2023, pp. 13745–13818. doi: 10.18653/v1/2023.findings-acl.868.

C.-H. Lin and U. Nuha, “Sentiment Analysis of Indonesian Datasets Based on a Hybrid Deep-Learning Strategy,” J. Big Data, vol. 10, no. 1, p. 88, 2023, doi: 10.1186/s40537-023-00782-9.

A. F. Aji et al., “One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 7226–7249. doi: 10.18653/v1/2022.acl-long.500.

F. Koto, J. H. Lau, and T. Baldwin, “IndoBERTweet: A Pretrained Language Model for Indonesian Twitter with Effective Domain-Specific Vocabulary Initialization,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021, pp. 10660–10668. doi: 10.18653/v1/2021.emnlp-main.833.

A. D. Sanya and L. H. Suadaa, “Handling Imbalanced Dataset on Hate Speech Detection in Indonesian Online News Comments,” in 2022 10th International Conference on Information and Communication Technology (ICoICT), IEEE, 2022, pp. 380–385. doi: 10.1109/ICoICT55009.2022.9914883.

A. P. J. Dwitama, D. H. Fudholi, and S. Hidayat, “Indonesian Hate Speech Detection Using Bidirectional Long Short-Term Memory (Bi-LSTM),” Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi), vol. 7, no. 2, pp. 302–309, 2023, doi: 10.29207/resti.v7i2.4642.

L. Wang, M. Han, X. Li, N. Zhang, and H. Cheng, “Review of Classification Methods on Unbalanced Data Sets,” IEEE Access, vol. 9, pp. 64606–64628, 2021, doi: 10.1109/ACCESS.2021.3074243.

G. James, D. Witten, T. Hastie, R. Tibshirani, and J. Taylor, “‘Statistical Learning,’ in An Introduction to Statistical Learning: With Applications in Python,” in An introduction to statistical learning: With applications in Python, Springer, 2023, pp. 15–67. doi: 10.1007/978-3-031-38747-0_2.

A. Géron, Hands-On Machine Learning with Scikit-Learn, Keras, and Tensorflow, 3rd ed. "O’Reilly Media, Inc., 2022.

I. D. Mienye and Y. Sun, “A Survey of Ensemble Learning: Concepts, Algorithms, Applications, and Prospects,” IEEE Access, vol. 10, pp. 99129–99149, 2022, doi: 10.1109/ACCESS.2022.3207287.

A. Tharwat, “Classification Assessment Methods,” Applied Computing and Informatics, vol. 17, no. 1, pp. 168–192, 2021, doi: 10.1016/j.aci.2018.08.003.

O. Rainio, J. Teuho, and R. Klén, “Evaluation Metrics and Statistical Tests for Machine Learning,” Scientific Reports (Sci.Rep.), vol. 14, no. 1, p. 6086, 2024, doi: 10.1038/s41598-024-56706-x.

D. Elreedy, A. F. Atiya, and F. Kamalov, “A Theoretical Distribution Analysis of Synthetic Minority Oversampling Technique (SMOTE) for Imbalanced Learning,” Mach. Learn., vol. 113, no. 7, pp. 4903–4923, 2024, doi: 10.1007/s10994-022-06296-4.

H. Zhou, “Research of Text Classification Based on TF-IDF and CNN-LSTM,” J. Phys. Conf. Ser., vol. 2171, no. 1, p. 012021, 2022, doi: 10.1088/1742-6596/2171/1/012021.

N. A. Semary, W. Ahmed, K. Amin, P. Pławiak, and M. Hammad, “Enhancing Machine Learning-Based Sentiment Analysis Through Feature Extraction Techniques,” PLoS One, vol. 19, no. 2, p. e0294968, 2024, doi: 10.1371/journal.pone.0294968.

M. Aria, C. Cuccurullo, and A. Gnasso, “A Comparison Among Interpretative Proposals for Random Forests,” Machine Learning with Applications, vol. 6, p. 100094, 2021, doi: 10.1016/j.mlwa.2021.100094.


Bila bermanfaat silahkan share artikel ini

Berikan Komentar Anda terhadap artikel Klasifikasi Sentimen Publik terhadap Isu Toleransi Agama Menggunakan Algoritma Random Forest

Dimensions Badge
Article History
Submitted: 2026-05-21
Published: 2026-07-31
Abstract View: 7 times
PDF Download: 5 times
Section
Articles