Deteksi Cyberbullying pada Komentar Instagram Berbahasa Indonesia Menggunakan Fine-Tuning IndoBERT dengan Analisis Kesalahan Model


  • Dewi Irma Afriayanti Universitas Esa Unggul, Jakarta, Indonesia
  • Budi Tjahjono * Mail Universitas Esa Unggul, Jakarta, Indonesia
  • (*) Corresponding Author
Keywords: Cyberbullying; Instagram; IndoBERT; NLP; Text Classification; Transformer

Abstract

Cyberbullying on social media is difficult to detect automatically because comments may contain informal language, sarcasm, implicit body shaming, and meanings that depend on context. This study evaluates the ability of IndoBERT to detect cyberbullying in Indonesian-language Instagram comments and analyzes the model’s classification errors. The dataset consists of 1,050 comments collected from public Instagram posts and manually labeled into two classes: cyberbullying and non-cyberbullying. The IndoBERT-base-uncased model was fine-tuned and evaluated using a confusion matrix, accuracy, precision, recall, and F1-score. The experimental results show an accuracy of 84.76% and an F1-score of 0.8462. For the cyberbullying class, precision reached 0.9294, while recall was 0.7524, indicating that a portion of bullying comments were still missed as false negatives. Error analysis shows that difficult cases were mainly associated with sarcasm, implicit body shaming, informal language, and ambiguous comments. The contributions of this study are an empirical evaluation of IndoBERT on Indonesian Instagram comments, a clear identification of the research gap between conventional machine-learning approaches and the need for contextual Indonesian-language modeling, and an error analysis that provides practical evidence for improving AI-based content moderation. The findings indicate that IndoBERT has promising performance, but larger datasets, conversational context, and pragmatic language modeling are still required.

Downloads

Download data is not yet available.

References

F. Ö. Öztürk, M. Tamaddon, and A. Tezel, “Cyberbullying, psychosocial problems and affecting factors among adolescents,” Arch. Psychiatr. Nurs., vol. 54, pp. 12–17, 2025, doi: https://doi.org/10.1016/j.apnu.2024.12.001.

G. Gohal et al., “Prevalence and related risks of cyberbullying and its effects on adolescent,” BMC Psychiatry, vol. 23, no. 1, p. 39, 2023, doi: 10.1186/s12888-023-04542-0.

P. Yi and A. Zubiaga, “Session-based cyberbullying detection in social media: A survey,” Online Soc. Netw. Media, vol. 36, p. 100250, 2023, doi: https://doi.org/10.1016/j.osnem.2023.100250.

B. Akdeniz and A. Doğan, “Cyberbullying: Definition, Prevalence, Effects, Risk and Protective Factors,” Psikiyatride Güncel Yaklaşımlar, vol. 16, no. 3, pp. 425–438, 2024, doi: 10.18863/pgy.1325195.

Ditch the Label, “Cyberbullying: The impact, prevalence, and ways to address it - 2024 report.” Accessed: Jun. 13, 2025. [Online]. Available: https://www.ditchthelabel.org/cyberbullying-report-2024

M. S. Jahan and M. Oussalah, “A systematic review of hate speech automatic detection using natural language processing,” Neurocomputing, vol. 546, p. 126232, 2023, doi: https://doi.org/10.1016/j.neucom.2023.126232.

A. M. El Koshiry, E. H. I. Eliwa, T. Abd El-Hafeez, and M. Khairy, “Detecting cyberbullying using deep learning techniques: a pre-trained glove and focal loss technique,” PeerJ Comput. Sci., vol. 10, p. e1961, Mar. 2024, doi: 10.7717/peerj-cs.1961.

S. Minaee, N. Kalchbrenner, E. Cambria, N. Nikzad, M. Chenaghlu, and J. Gao, “Deep Learning–based Text Classification: A Comprehensive Review,” ACM Comput. Surv., vol. 54, no. 3, Apr. 2021, doi: 10.1145/3439726.

A. Vaswani et al., “Attention Is All You Need,” 2023.

H. Wang, J. Li, H. Wu, E. Hovy, and Y. Sun, “Pre-Trained Language Models and Their Applications,” Engineering, vol. 25, pp. 51–65, 2023, doi: https://doi.org/10.1016/j.eng.2022.04.024.

B. Ogunleye and B. Dharmaraj, “The Use of a Large Language Model for Cyberbullying Detection,” Analytics, vol. 2, no. 3, pp. 694–707, Sep. 2023, doi: 10.3390/analytics2030038.

M. K. Ojha, N. M. Patil, and M. Joshi, “A Comparative Study of Classification Models for Cyberbullying Detection,” in 2024 International Conference on Inventive Computation Technologies (ICICT), 2024, pp. 1531–1535. doi: 10.1109/ICICT60155.2024.10544792.

H. Pen, N. Anne Huiying TEO, Z. Wang, N. Anne Huiying, and N. Teo, “Comparative analysis of hate speech detection: Traditional vs. Comparative analysis of hate speech detection: Traditional vs. deep learning approaches deep learning approaches Part of the Artificial Intelligence and Robotics Commons, and the Databases and Information Systems Commons Citation Citation Comparative Analysis of Hate Speech Detection: Traditional vs. Deep Learning Approaches,” 2024. [Online]. Available: https://ink.library.smu.edu.sg/sis_research

A. Perera and P. Fernando, “Cyberbullying Detection System on Social Media Using Supervised Machine Learning,” Procedia Comput. Sci., vol. 239, pp. 506–516, 2024, doi: https://doi.org/10.1016/j.procs.2024.06.200.

R. Kulkarni, S. Chakrabarti, S. Salunke, T. Wagh, and A. Thool, “A Comprehensive Analysis of Cyberbullying Detection Using Various Machine Learning Approaches,” 2024, pp. 15–25. doi: 10.1007/978-981-97-6678-9_2.

F. R. Sayed, E. H. Elnashar, and F. A. Omara, “Cyberbullying detection in social media using natural language processing,” Jun. 01, 2025, Elsevier B.V. doi: 10.1016/j.sciaf.2025.e02713.

B. Joshi, B. K. Joshi, S. Pant, A. Kumar, and H. K. Sharma, “An Efficient Method for Detecting Cyberbullying Using Supervised Machine Learning Techniques,” in Procedia Computer Science, Elsevier B.V., 2025, pp. 1254–1261. doi: 10.1016/j.procs.2025.04.359.

B. S, H. D, and S. S, Cyberbullying Detection Using Fine-Tuned BERT. 2025. doi: 10.1109/ICTEST64710.2025.11042412.

J. Peng and K. Han, “Survey of Pre-trained Models for Natural Language Processing,” in 2021 International Conference on Electronic Communications, Internet of Things and Big Data (ICEIB), 2021, pp. 277–280. doi: 10.1109/ICEIB53692.2021.9686420.

J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” CoRR, vol. abs/1810.04805, 2018, [Online]. Available: http://arxiv.org/abs/1810.04805

P. K. Nag, A. Bhagat, R. Vishnu Priya, and D. K. Khare, “Emotional Intelligence Through Artificial Intelligence: NLP and Deep Learning in the Analysis of Healthcare Texts,” in 2023 International Conference on Artificial Intelligence for Innovations in Healthcare Industries (ICAIIHI), 2023, pp. 1–7. doi: 10.1109/ICAIIHI57871.2023.10489117.

A. Mansoori et al., “Detection of Sarcasm in News Headlines Using NLP and Machine Learning,” 2025, pp. 503–517. doi: 10.1007/978-3-031-89175-5_31.


Bila bermanfaat silahkan share artikel ini

Berikan Komentar Anda terhadap artikel Deteksi Cyberbullying pada Komentar Instagram Berbahasa Indonesia Menggunakan Fine-Tuning IndoBERT dengan Analisis Kesalahan Model

Dimensions Badge
Article History
Submitted: 2026-02-01
Published: 2026-08-31
Abstract View: 78 times
PDF Download: 13 times
Section
Articles