Evaluasi Pseudo-Labeling IndoRoBERTa dan InSet Lexicon dengan SVM pada Komentar TikTok
Abstract
TikTok generates large volumes of public comments that can be used to identify trends in public opinion, including responses to the accident involving a vehicle used in the Free Nutritious Meal program. However, informal social media language and imbalanced class distributions may affect sentiment labeling and classification performance. This study evaluates sentiment pseudo-labels generated by IndoRoBERTa and InSet Lexicon through Support Vector Machine classification with the application of the Synthetic Minority Over-sampling Technique. A quantitative experimental approach was applied to 10,309 TikTok comments collected through Apify scraping. Since no human annotations were used as ground truth, the positive, negative, and neutral labels produced by both methods were treated as pseudo-labels. The research stages included filtering, preprocessing, pseudo-label generation, Term Frequency-Inverse Document Frequency feature extraction, an 80:20 data split, SMOTE application to the training data, SVM classification, and evaluation using accuracy, precision, recall, and F1-score. The results show that SVM reproduced the InSet Lexicon pseudo-labels most effectively, achieving an accuracy of 0.865 without SMOTE. After SMOTE was applied, precision increased to 0.870 while the F1-score remained at 0.865, and neutral-class recall increased from 0.767 to 0.807. These findings indicate that InSet Lexicon produced pseudo-labels that were more consistently learned by SVM on this dataset, while SMOTE primarily improved minority-class recognition.
Downloads
References
Andrianto, F., Fadlil, A., & Riadi, I. (2024). Linear kernel optimization of Support Vector Machine algorithm on online marketplace sentiment analysis. Komputasi: Jurnal Ilmiah Ilmu Komputer Dan Matematika, 21(1), 68–82. https://doi.org/10.33751/komputasi.v21i1.9266
Azril, M., & Crisnawati, G. (2026). Klasifikasi sentimen komentar pengguna TikTok mengenai program Makan Bergizi Gratis (MBG) dengan Support Vector Machine (SVM). RIGGS: Journal of Artificial Intelligence and Digital Business, 5(1), 8859–8866. https://doi.org/10.31004/riggs.v5i1.7237
Darusman, D., & Gata, W. (2025). Perbandingan kinerja machine learning dan deep learning untuk analisis sentimen Fufufafa. Information System for Educators and Professionals: Journal of Information System, 10(1), 1–12. https://doi.org/10.51211/isbi.v10i1.3333
Dinata, R. M., Marhaeni, M., Atmadja, K., Rayhana, E., Hadi, V., & Al Kaf, U. (2025). Analisis komprehensif kinerja model klasifikasi sentimen: Evaluasi lintas metrik pada dataset tweet film bahasa Indonesia. Jurnal Rekayasa Informasi, 14(1), 38–47. https://repository.istn.ac.id/13543/
Firdaus, R., & Herdiani, A. (2021). Lexicon-based sentiment analysis of Indonesian language student feedback evaluation. Jurnal Computer, 6(1), 1–12. https://doi.org/10.34818/indojc.2021.6.1.408
Hairani, H., Widiyaningtyas, T., & Prasetya, D. D. (2024). Addressing class imbalance of health data: A systematic literature review on modified Synthetic Minority Oversampling Technique (SMOTE) strategies. International Journal of Informatics and Visualization, 8(3), 1310–1318. https://doi.org/10.62527/joiv.8.3.2283
Helmiyah, S., & Pramestiawan, R. (2025). Analisis komparatif algoritma machine learning dengan metrik akurasi, presisi, recall, dan F1-score pada dataset kacang kering. Jurnal Ilmu Komputer Dan Teknologi, 6(3), 152–159. https://doi.org/10.35960/ikomti.v6i3.2031
Iansyah, K., Nurlaili, A. L., & Al Haromainy, M. M. (2025). Comparative analysis of IndoBERT, IndoBERTweet, and XLM-RoBERTa for detecting online gambling comments on YouTube. Bit-Tech, 8(2), 2379–2390. https://doi.org/10.32877/bt.v8i2.3257
Indriyani, F. A., Fauzi, A., & Faisal, S. (2023a). Analisis sentimen aplikasi TikTok menggunakan algoritma Naive Bayes dan Support Vector Machine. TEKNOSAINS: Jurnal Sains, Teknologi Dan Informatika, 10(2), 176–184. https://doi.org/10.37373/tekno.v10i2.419
Indriyani, F. A., Fauzi, A., & Faisal, S. (2023b). Analisis sentimen aplikasi TikTok menggunakan algoritma Naive Bayes dan Support Vector Machine. TEKNOSAINS: Jurnal Sains, Teknologi Dan Informatika, 10(2), 176–184. https://doi.org/10.37373/tekno.v10i2.419
Istiqomah, A., Safaah, T. N., & Haryati, G. (2025). Variasi bahasa dalam komentar TikTok: Kajian sosiolinguistik digital. J-CEKI: Jurnal Cendekia Ilmiah, 4(6), 979–985. https://doi.org/10.56799/jceki.v4i6.10418
Junaedi, Gunawan, A. H., Kuswanto, V., & Jonathan. (2024). Tinjauan Support Vector Machine dalam text-mining untuk analisis sentimen di sektor pariwisata. Bit-Tech, 7(2), 323–330. https://doi.org/10.32877/bt.v7i2.1810
Kamalov, F., Choutri, S. E., & Atiya, A. F. (2025). Analytical formulation of Synthetic Minority Oversampling Technique (SMOTE) for imbalanced learning. Gulf Journal of Mathematics, 19(1), 400–415. https://doi.org/10.56947/gjom.v19i1.2639
Manalu, P. D., Simanjuntak, M., & Umri, C. (2025). Implementasi algoritma klasifikasi untuk analisis sentimen media sosial TikTok tahun 2025. Jurnal Teknik Informatika Dan Teknologi Informasi, 5(1), 488–504. https://doi.org/10.55606/jutiti.v5i1.5644
Musfiroh, D., Khaira, U., Utomo, P. E. P., & Suratno, T. (2021). Analisis sentimen terhadap perkuliahan daring di Indonesia dari Twitter dataset menggunakan InSet Lexicon. MALCOM: Indonesian Journal of Machine Learning and Computer Science, 1(1), 24–33. https://doi.org/10.57152/malcom.v1i1.20
Pateman, D., Prasetyo, T. F., & Sujadi, H. (2025). Sentiment analysis of government on TikTok and X platforms with SVM and SMOTE approach. JITK: Jurnal Ilmu Pengetahuan Dan Teknologi Komputer, 10(4), 900–908. https://doi.org/10.33480/jitk.v10i4.6645
Prasetya, H., Situmorang, Z., & Rosnelly, R. (2024). SVM optimization with kernel function for sentiment analysis on social media Twitter (X) in AFC U23 Asian Cup case study. Proceedings of the International Conference of Science Technology UISU, 227–233. https://doi.org/10.30743/wjxmmr59
Ramadhan, C., Atina, V., & Permatasari, H. (2025). Analisis perbandingan model CNN dan IndoBERT dalam sentimen berita politik Indonesia. Prosiding Seminar Nasional Teknologi Informasi Dan Bisnis 2025, 110–118. https://doi.org/10.47701/v1r9ka69
Rizki, A. S., Aristi, N. M., Ridha, N., Zulfahri, A. F., & Wibowo, D. A. (2023). Implementation of the Indonesian language stemming algorithm in Twitter data preprocessing: Case study Twitter Wargabanua and Instakalsel. Fidelity: Jurnal Teknik Elektro, 5(3), 175–183. https://doi.org/10.52005/fidelity.v5i3.170
Rufaida, A. S. R., Permanasari, A. E., & Setiawan, N. A. (2023). Lexicon-based sentiment analysis using InSet Dictionary: A systematic literature review. Proceedings of the 5th International Conference on Applied Engineering (ICAE 2022). https://doi.org/10.4108/eai.5-10-2022.2327474
Sidupa, B. C., & Dewi, C. (2025). Sentimen analisis terhadap aplikasi TikTok menggunakan Support Vector Classification. Jurnal Mnemonic, 8(1), 1–8. https://doi.org/10.36040/mnemonic.v8i1.12635
Suhaeni, C., Kamila, S. A., Fahira, F., Yusran, M., & Dito, G. A. (2025). Exploring a large language model on the ChatGPT platform for Indonesian text preprocessing tasks. Indonesian Journal of Statistics and Its Applications, 9(1), 100–116. https://doi.org/10.29244/ijsa.v9i1p100-116
Sulistiyono, M., Pristyanto, Y., Adi, S., & Gumelar, G. (2021). Implementasi algoritma Synthetic Minority Over-Sampling Technique untuk menangani ketidakseimbangan kelas pada dataset klasifikasi. Sistemasi, 10(2), 445–457. https://doi.org/10.32520/stmsi.v10i2.1303
Syafaah, L. A., & Haryanto, S. (2024). Slang semantic analysis on TikTok social media Generation Z. Proceeding ISETH: International Summit on Science, Technology, and Humanity, 476–484. https://doi.org/10.23917/iseth.3898
Wardhani, D., Astuti, R., & Saputra, D. D. (2024). Optimasi feature selection text mining: Stemming dan stopword untuk sentimen analisis aplikasi SatuSehat. Innovative: Journal of Social Science Research, 4(1), 7537–7548. https://doi.org/10.31004/innovative.v4i1.8759
Wijaya, I. N. S. W., Seputra, K. A., & Dewi, N. P. N. P. (2025). Fine tunning model IndoBERT untuk analisis sentimen berita pariwisata Indonesia. Jurnal Pendidikan Teknologi Dan Kejuruan, 22(2), 195–204. https://doi.org/10.23887/jptk-undiksha.v22i2.104056
Bila bermanfaat silahkan share artikel ini
Berikan Komentar Anda terhadap artikel Evaluasi Pseudo-Labeling IndoRoBERTa dan InSet Lexicon dengan SVM pada Komentar TikTok
Pages: 653-661
Copyright (c) 2026 Yehezkiel Juandro Metta, Aswan Supriyadi Sunge, Asep Suprianto

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under Creative Commons Attribution 4.0 International License that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (Refer to The Effect of Open Access).













