Evaluasi Pseudo-Labeling IndoRoBERTa dan InSet Lexicon dengan SVM pada Komentar TikTok


  • Yehezkiel Juandro Metta * Mail Universitas Pelita Bangsa, Bekasi, Indonesia
  • Aswan Supriyadi Sunge Universitas Pelita Bangsa, Bekasi, Indonesia
  • Asep Suprianto Universitas Pelita Bangsa, Bekasi, Indonesia
  • (*) Corresponding Author
Keywords: Sentiment Analysis; Pseudo-Labeling; IndoRoBERTa; InSet Lexicon; SVM; SMOTE

Abstract

TikTok generates large volumes of public comments that can be used to identify trends in public opinion, including responses to the accident involving a vehicle used in the Free Nutritious Meal program. However, informal social media language and imbalanced class distributions may affect sentiment labeling and classification performance. This study evaluates sentiment pseudo-labels generated by IndoRoBERTa and InSet Lexicon through Support Vector Machine classification with the application of the Synthetic Minority Over-sampling Technique. A quantitative experimental approach was applied to 10,309 TikTok comments collected through Apify scraping. Since no human annotations were used as ground truth, the positive, negative, and neutral labels produced by both methods were treated as pseudo-labels. The research stages included filtering, preprocessing, pseudo-label generation, Term Frequency-Inverse Document Frequency feature extraction, an 80:20 data split, SMOTE application to the training data, SVM classification, and evaluation using accuracy, precision, recall, and F1-score. The results show that SVM reproduced the InSet Lexicon pseudo-labels most effectively, achieving an accuracy of 0.865 without SMOTE. After SMOTE was applied, precision increased to 0.870 while the F1-score remained at 0.865, and neutral-class recall increased from 0.767 to 0.807. These findings indicate that InSet Lexicon produced pseudo-labels that were more consistently learned by SVM on this dataset, while SMOTE primarily improved minority-class recognition.

Downloads

Download data is not yet available.

References

Andrianto, F., Fadlil, A., & Riadi, I. (2024). Linear kernel optimization of Support Vector Machine algorithm on online marketplace sentiment analysis. Komputasi: Jurnal Ilmiah Ilmu Komputer Dan Matematika, 21(1), 68–82. https://doi.org/10.33751/komputasi.v21i1.9266

Azril, M., & Crisnawati, G. (2026). Klasifikasi sentimen komentar pengguna TikTok mengenai program Makan Bergizi Gratis (MBG) dengan Support Vector Machine (SVM). RIGGS: Journal of Artificial Intelligence and Digital Business, 5(1), 8859–8866. https://doi.org/10.31004/riggs.v5i1.7237

Darusman, D., & Gata, W. (2025). Perbandingan kinerja machine learning dan deep learning untuk analisis sentimen Fufufafa. Information System for Educators and Professionals: Journal of Information System, 10(1), 1–12. https://doi.org/10.51211/isbi.v10i1.3333

Dinata, R. M., Marhaeni, M., Atmadja, K., Rayhana, E., Hadi, V., & Al Kaf, U. (2025). Analisis komprehensif kinerja model klasifikasi sentimen: Evaluasi lintas metrik pada dataset tweet film bahasa Indonesia. Jurnal Rekayasa Informasi, 14(1), 38–47. https://repository.istn.ac.id/13543/

Firdaus, R., & Herdiani, A. (2021). Lexicon-based sentiment analysis of Indonesian language student feedback evaluation. Jurnal Computer, 6(1), 1–12. https://doi.org/10.34818/indojc.2021.6.1.408

Hairani, H., Widiyaningtyas, T., & Prasetya, D. D. (2024). Addressing class imbalance of health data: A systematic literature review on modified Synthetic Minority Oversampling Technique (SMOTE) strategies. International Journal of Informatics and Visualization, 8(3), 1310–1318. https://doi.org/10.62527/joiv.8.3.2283

Helmiyah, S., & Pramestiawan, R. (2025). Analisis komparatif algoritma machine learning dengan metrik akurasi, presisi, recall, dan F1-score pada dataset kacang kering. Jurnal Ilmu Komputer Dan Teknologi, 6(3), 152–159. https://doi.org/10.35960/ikomti.v6i3.2031

Iansyah, K., Nurlaili, A. L., & Al Haromainy, M. M. (2025). Comparative analysis of IndoBERT, IndoBERTweet, and XLM-RoBERTa for detecting online gambling comments on YouTube. Bit-Tech, 8(2), 2379–2390. https://doi.org/10.32877/bt.v8i2.3257

Indriyani, F. A., Fauzi, A., & Faisal, S. (2023a). Analisis sentimen aplikasi TikTok menggunakan algoritma Naive Bayes dan Support Vector Machine. TEKNOSAINS: Jurnal Sains, Teknologi Dan Informatika, 10(2), 176–184. https://doi.org/10.37373/tekno.v10i2.419

Indriyani, F. A., Fauzi, A., & Faisal, S. (2023b). Analisis sentimen aplikasi TikTok menggunakan algoritma Naive Bayes dan Support Vector Machine. TEKNOSAINS: Jurnal Sains, Teknologi Dan Informatika, 10(2), 176–184. https://doi.org/10.37373/tekno.v10i2.419

Istiqomah, A., Safaah, T. N., & Haryati, G. (2025). Variasi bahasa dalam komentar TikTok: Kajian sosiolinguistik digital. J-CEKI: Jurnal Cendekia Ilmiah, 4(6), 979–985. https://doi.org/10.56799/jceki.v4i6.10418

Junaedi, Gunawan, A. H., Kuswanto, V., & Jonathan. (2024). Tinjauan Support Vector Machine dalam text-mining untuk analisis sentimen di sektor pariwisata. Bit-Tech, 7(2), 323–330. https://doi.org/10.32877/bt.v7i2.1810

Kamalov, F., Choutri, S. E., & Atiya, A. F. (2025). Analytical formulation of Synthetic Minority Oversampling Technique (SMOTE) for imbalanced learning. Gulf Journal of Mathematics, 19(1), 400–415. https://doi.org/10.56947/gjom.v19i1.2639

Manalu, P. D., Simanjuntak, M., & Umri, C. (2025). Implementasi algoritma klasifikasi untuk analisis sentimen media sosial TikTok tahun 2025. Jurnal Teknik Informatika Dan Teknologi Informasi, 5(1), 488–504. https://doi.org/10.55606/jutiti.v5i1.5644

Musfiroh, D., Khaira, U., Utomo, P. E. P., & Suratno, T. (2021). Analisis sentimen terhadap perkuliahan daring di Indonesia dari Twitter dataset menggunakan InSet Lexicon. MALCOM: Indonesian Journal of Machine Learning and Computer Science, 1(1), 24–33. https://doi.org/10.57152/malcom.v1i1.20

Pateman, D., Prasetyo, T. F., & Sujadi, H. (2025). Sentiment analysis of government on TikTok and X platforms with SVM and SMOTE approach. JITK: Jurnal Ilmu Pengetahuan Dan Teknologi Komputer, 10(4), 900–908. https://doi.org/10.33480/jitk.v10i4.6645

Prasetya, H., Situmorang, Z., & Rosnelly, R. (2024). SVM optimization with kernel function for sentiment analysis on social media Twitter (X) in AFC U23 Asian Cup case study. Proceedings of the International Conference of Science Technology UISU, 227–233. https://doi.org/10.30743/wjxmmr59

Ramadhan, C., Atina, V., & Permatasari, H. (2025). Analisis perbandingan model CNN dan IndoBERT dalam sentimen berita politik Indonesia. Prosiding Seminar Nasional Teknologi Informasi Dan Bisnis 2025, 110–118. https://doi.org/10.47701/v1r9ka69

Rizki, A. S., Aristi, N. M., Ridha, N., Zulfahri, A. F., & Wibowo, D. A. (2023). Implementation of the Indonesian language stemming algorithm in Twitter data preprocessing: Case study Twitter Wargabanua and Instakalsel. Fidelity: Jurnal Teknik Elektro, 5(3), 175–183. https://doi.org/10.52005/fidelity.v5i3.170

Rufaida, A. S. R., Permanasari, A. E., & Setiawan, N. A. (2023). Lexicon-based sentiment analysis using InSet Dictionary: A systematic literature review. Proceedings of the 5th International Conference on Applied Engineering (ICAE 2022). https://doi.org/10.4108/eai.5-10-2022.2327474

Sidupa, B. C., & Dewi, C. (2025). Sentimen analisis terhadap aplikasi TikTok menggunakan Support Vector Classification. Jurnal Mnemonic, 8(1), 1–8. https://doi.org/10.36040/mnemonic.v8i1.12635

Suhaeni, C., Kamila, S. A., Fahira, F., Yusran, M., & Dito, G. A. (2025). Exploring a large language model on the ChatGPT platform for Indonesian text preprocessing tasks. Indonesian Journal of Statistics and Its Applications, 9(1), 100–116. https://doi.org/10.29244/ijsa.v9i1p100-116

Sulistiyono, M., Pristyanto, Y., Adi, S., & Gumelar, G. (2021). Implementasi algoritma Synthetic Minority Over-Sampling Technique untuk menangani ketidakseimbangan kelas pada dataset klasifikasi. Sistemasi, 10(2), 445–457. https://doi.org/10.32520/stmsi.v10i2.1303

Syafaah, L. A., & Haryanto, S. (2024). Slang semantic analysis on TikTok social media Generation Z. Proceeding ISETH: International Summit on Science, Technology, and Humanity, 476–484. https://doi.org/10.23917/iseth.3898

Wardhani, D., Astuti, R., & Saputra, D. D. (2024). Optimasi feature selection text mining: Stemming dan stopword untuk sentimen analisis aplikasi SatuSehat. Innovative: Journal of Social Science Research, 4(1), 7537–7548. https://doi.org/10.31004/innovative.v4i1.8759

Wijaya, I. N. S. W., Seputra, K. A., & Dewi, N. P. N. P. (2025). Fine tunning model IndoBERT untuk analisis sentimen berita pariwisata Indonesia. Jurnal Pendidikan Teknologi Dan Kejuruan, 22(2), 195–204. https://doi.org/10.23887/jptk-undiksha.v22i2.104056


Bila bermanfaat silahkan share artikel ini

Berikan Komentar Anda terhadap artikel Evaluasi Pseudo-Labeling IndoRoBERTa dan InSet Lexicon dengan SVM pada Komentar TikTok

Dimensions Badge
Article History
Published: 2026-06-30
Abstract View: 0 times
PDF Download: 0 times
Issue
Section
Articles