Random Forest, LSTM, and IndoBERT Comparison for TikTok App Sentiment Analysis


  • Imam Saputra Sekolah Tinggi Ilmu Manajemen Sukma, Medan, Indonesia
  • Mesran Mesran * Mail Sekolah Tinggi Ilmu Manajemen Sukma, Medan, Indonesia
  • Ruziana Mohamad Rasli Universiti Utara Malaysia, Kedah, Malaysia
  • (*) Corresponding Author
Keywords: Indonesian NLP; Sentiment Analysis; Class Imbalance; Random Forest; Long Short-Term Memory; IndoBERT; TikTok

Abstract

The rapid growth of social media platforms like TikTok has generated a massive volume of user reviews on the Google Play Store, serving as a critical indicator of application service quality. However, the unstructured nature of Indonesian social media text and the significant imbalance between sentiment classes pose substantial challenges for automated classification systems. Addressing this class imbalance is highly crucial for application developers, as critical negative and neutral feedback containing essential feature complaints is easily marginalized by the overwhelming majority of positive reviews, leading to biased operational insights. This research conducted a comprehensive comparative study of three distinct computational paradigms: Random Forest, Long Short-Term Memory (LSTM), and IndoBERT, to identify the most effective model for sentiment analysis. A dataset of 5,000 TikTok reviews was meticulously processed using a negation-aware preprocessing pipeline to preserve semantic integrity. To address class imbalance, architecture-specific techniques were deployed, including SMOTE for Random Forest, Class Weighting for LSTM, and Random OverSampling for IndoBERT. The experimental results demonstrate that IndoBERT significantly outperforms other models, achieving the highest global accuracy of 81% and a Macro F1-Score of 0.56. While Random Forest and LSTM yielded lower accuracies of 75% and 71%, respectively, they exhibited stability in predicting the majority class but struggled with the inherent ambiguity of neutral sentiments. The study concludes that IndoBERT’s bidirectional self-attention mechanism provides superior contextual understanding of Indonesian slang and non-formal syntax. This research contributes a robust framework for application developers to monitor public opinion objectively. Furthermore, the findings highlight that despite advanced balancing techniques, the "neutrality bottleneck" remains a challenge, suggesting that future research should explore aspect-based sentiment analysis to enhance classification granularity in the Indonesian NLP domain.

Downloads

Download data is not yet available.

References

J. B. Pick, F. Ren, and A. Sarkar, “Social media access and purposeful use in China: Geospatial patterns and socioeconomic and COVID-19 influences,” Telecomm. Policy, vol. 49, no. 7, pp. 1–18, 2025, doi: 10.1016/j.telpol.2025.103002.

O. Grljević, M. Marić, and R. Božić, “Exploring Mobile Application User Experience Through Topic Modeling,” Sustainability (Switzerland), vol. 17, no. 3, pp. 1–32, 2025, doi: 10.3390/su17031109.

M. Mustak, H. Hallikainen, T. Laukkanen, L. Plé, L. D. Hollebeek, and M. Aleem, “Using machine learning to develop customer insights from user-generated content,” Journal of Retailing and Consumer Services, vol. 81, no. March, pp. 1–14, 2024, doi: 10.1016/j.jretconser.2024.104034.

W. Chen, W. Hussain, F. Cauteruccio, and X. Zhang, “Deep Learning for Financial Time Series Prediction: A State-of-the-Art Review of Standalone and Hybrid Models,” CMES - Computer Modeling in Engineering and Sciences, vol. 139, no. 1, pp. 187–224, 2023, doi: 10.32604/cmes.2023.031388.

A. Daza, N. D. González Rueda, M. S. Aguilar Sánchez, W. F. Robles Espíritu, and M. E. Chauca Quiñones, “Sentiment Analysis on E-Commerce Product Reviews Using Machine Learning and Deep Learning Algorithms: A Bibliometric Analysisand Systematic Literature Review, Challenges and Future Works,” International Journal of Information Management Data Insights, vol. 4, no. 2, pp. 1–20, 2024, doi: 10.1016/j.jjimei.2024.100267.

E. C. M. Torres and L. G. de Picado-Santos, “Using Sentiment Analysis to Study the Potential for Improving Sustainable Mobility in University Campuses,” Sustainability (Switzerland), vol. 17, no. 14, pp. 1–22, 2025, doi: 10.3390/su17146645.

M. Kim, S. Kim, Y. Park, S. Bahn, S. H. Ahn, and B. NambiNarayanan, “User Sentiment Analysis Based on Securities Application Elements,” Behavioral Sciences, vol. 14, no. 9, pp. 1–19, 2024, doi: 10.3390/bs14090814.

K. Imtihan, L. Mutawali, W. Bagye, and A. Tantoni, “Automated Label Extraction for Sentiment Analysis in Indonesian Text,” Int. J. Adv. Sci. Eng. Inf. Technol., vol. 15, no. 3, pp. 718–728, 2025, doi: 10.18517/ijaseit.15.3.20602.

A. Abdillah, I. Widianingsih, R. A. Buchari, and H. Nurasa, “Big data security & individual (psychological) resilience: A review of social media risks and lessons learned from Indonesia,” Array, vol. 21, no. February, pp. 1–13, 2024, doi: 10.1016/j.array.2024.100336.

N. Merayo, R. Cotelo, R. Carratalá-Sáez, and F. J. Andújar, “Applying machine learning to assess emotional reactions to video game content streamed on Spanish Twitch channels,” Comput. Speech Lang., vol. 88, no. March, pp. 1–23, 2024, doi: 10.1016/j.csl.2024.101651.

M. O. Ibrohim and I. Budi, “Hate speech and abusive language detection in Indonesian social media: Progress and challenges,” Heliyon, vol. 9, no. 8, pp. 1–16, 2023, doi: 10.1016/j.heliyon.2023.e18647.

M. Abdalsalam, C. Li, A. Dahou, and N. Kryvinska, “Terrorism Attack Classification Using Machine Learning: The Effectiveness of Using Textual Features Extracted from GTD Dataset,” CMES - Computer Modeling in Engineering and Sciences, vol. 138, no. 2, pp. 1427–1467, 2024, doi: 10.32604/cmes.2023.029911.

B. R. Ermawan, M. B. Prayoga, A. R. Fadhillah, and E. Utami, “Optimizing Sentiment Classification Models for TikTok Comments using Emotion-Based Preprocessing and Grid Search,” vol. 10, no. 1, pp. 535–547, 2026.

Y. Aliyu, A. Sarlan, K. U. Danyaro, A. Sani abd Rahman, A. A. Muazu, and M. Y. Abubakar, “Deep learning techniques for sentiment analysis in code-switched Hausa-English tweets,” International Journal of Information Management Data Insights, vol. 5, no. 1, pp. 1–19, 2025, doi: 10.1016/j.jjimei.2025.100330.

A. M. Moslhi, H. H. Aly, and M. ElMessiery, “The Impact of Feature Extraction on Classification Accuracy Examined by Employing a Signal Transformer to Classify Hand Gestures Using Surface Electromyography Signals,” Sensors, vol. 24, no. 4, pp. 1–21, 2024, doi: 10.3390/s24041259.

C. Suhaeni and H. S. Yong, “Enhancing Imbalanced Sentiment Analysis: A GPT-3-Based Sentence-by-Sentence Generation Approach,” Applied Sciences (Switzerland), vol. 14, no. 2, pp. 1–24, 2024, doi: 10.3390/app14020622.

W. Chen, K. Yang, Z. Yu, Y. Shi, and C. L. P. Chen, A survey on imbalanced learning: latest research, applications and future directions, vol. 57, no. 6. 2024. doi: 10.1007/s10462-024-10759-6.

N. B. Yadav, “‘Harnessing Customer Feedback for Product Recommendations: An Aspect-Level Sentiment Analysis Framework,’” Human-Centric Intelligent Systems, vol. 3, no. 2, pp. 57–67, 2023, doi: 10.1007/s44230-023-00018-2.

F. Sağlam, DatRel: a noise-tolerant data relocation approach for effective synthetic data generation in imbalanced classifiers, vol. 114, no. 5. Springer US, 2025. doi: 10.1007/s10994-025-06755-8.

M. Wojciuk, Z. Swiderska-Chadaj, K. Siwek, and A. Gertych, “Improving classification accuracy of fine-tuned CNN models: Impact of hyperparameter optimization,” Heliyon, vol. 10, no. 5, pp. 1–11, 2024, doi: 10.1016/j.heliyon.2024.e26586.

J. Shin et al., “Exploring the Effectiveness of Machine Learning and Deep Learning Algorithms for Sentiment Analysis: A Systematic Literature Review,” Computers, Materials and Continua, vol. 84, no. 3, pp. 4105–4153, 2025, doi: 10.32604/cmc.2025.066910.

S. A. Amamra, “Random Forest-Based Machine Learning Model Design for 21,700/5 Ah Lithium Cell Health Prediction Using Experimental Data,” Physchem, vol. 5, no. 1, pp. 1–19, 2025, doi: 10.3390/physchem5010012.

J. A. Nasir, O. S. Khan, and I. Varlamis, “Fake news detection: A hybrid CNN-RNN based deep learning approach,” International Journal of Information Management Data Insights, vol. 1, no. 1, pp. 1–13, 2021, doi: 10.1016/j.jjimei.2020.100007.

N. Saraswathi, T. Sasi Rooba, and S. Chakaravarthi, “Improving the accuracy of sentiment analysis using a linguistic rule-based feature selection method in tourism reviews,” Measurement: Sensors, vol. 29, no. May, pp. 1–7, 2023, doi: 10.1016/j.measen.2023.100888.

E. Elhosary and O. Moselhi, “Evaluating Natural Language Processing Algorithms for Improved Hazard and Operability Analysis,” Geodata and AI, vol. 4, no. May, pp. 1–15, 2025, doi: 10.1016/j.geoai.2025.100026.

A. Alharthi, M. Alaryani, and S. Kaddoura, “A comparative study of machine learning and deep learning models in binary and multiclass classification for intrusion detection systems,” Array, vol. 26, no. May, pp. 1–15, 2025, doi: 10.1016/j.array.2025.100406.

Supriyono, A. P. Wibawa, Suyono, and F. Kurniawan, “Advancements in natural language processing: Implications, challenges, and future directions,” Telematics and Informatics Reports, vol. 16, no. November, pp. 1–17, 2024, doi: 10.1016/j.teler.2024.100173.

S. M. Al-Selwi et al., “RNN-LSTM: From applications to modeling techniques and beyond—Systematic review,” Journal of King Saud University - Computer and Information Sciences, vol. 36, no. 5, pp. 1–34, 2024, doi: 10.1016/j.jksuci.2024.102068.

A. O. Almagrabi, “A Deep CNN-LSTM-Based Feature Extraction for Cyber-Physical System Monitoring,” Computers, Materials and Continua, vol. 76, no. 2, pp. 2079–2093, 2023, doi: 10.32604/cmc.2023.039683.

Q. Gu, Z. Wang, H. Zhang, S. Sui, and R. Wang, “Aspect-Level Sentiment Analysis Based on Syntax-Aware and Graph Convolutional Networks,” Applied Sciences (Switzerland), vol. 14, no. 2, pp. 1–13, 2024, doi: 10.3390/app14020729.

B. Ihnaini, B. Abuhaija, E. A. Mills, and M. Mahmuddin, “Semantic similarity on multimodal data: A comprehensive survey with applications,” Journal of King Saud University - Computer and Information Sciences, vol. 36, no. 10, pp. 1–37, 2024, doi: 10.1016/j.jksuci.2024.102263.

M. J. Abbass, R. Lis, and W. Rebizant, “A Predictive Model Using Long Short-Time Memory (LSTM) Technique for Power System Voltage Stability,” Applied Sciences (Switzerland), vol. 14, no. 16, pp. 1–13, 2024, doi: 10.3390/app14167279.

D. G. da Silva and A. A. de M. Meneses, “Comparing Long Short-Term Memory (LSTM) and bidirectional LSTM deep neural networks for power consumption prediction,” Energy Reports, vol. 10, no. May, pp. 3315–3334, 2023, doi: 10.1016/j.egyr.2023.09.175.

J. Campino, “Unleashing the transformers: NLP models detect AI writing in education,” Journal of Computers in Education, vol. 12, no. 2, pp. 645–673, 2025, doi: 10.1007/s40692-024-00325-y.

P. W. Cahyo, U. S. Aesyi, W. A. Setianto, and T. Sulaiman, “A Novel Named Entity Recognition approach of Indonesian fake news using part of speech and BERT model on presidential election,” International Journal of Information Management Data Insights, vol. 5, no. 2, pp. 1–11, 2025, doi: 10.1016/j.jjimei.2025.100354.

A. Alamsyah and Y. Sagama, “Empowering Indonesian internet users: An approach to counter online toxicity and enhance digital well-being,” Intelligent Systems with Applications, vol. 22, no. May, pp. 1–16, 2024, doi: 10.1016/j.iswa.2024.200394.

A. F. Hidayatullah, R. A. Apong, D. T. C. Lai, and A. Qazi, “Pre-trained language model for code-mixed text in Indonesian, Javanese, and English using transformer,” Soc. Netw. Anal. Min., vol. 15, no. 1, pp. 1–17, 2025, doi: 10.1007/s13278-025-01444-9.

A. S. Ekakristi, A. F. Wicaksono, and R. Mahendra, “Intermediate-task transfer learning for Indonesian NLP tasks,” Natural Language Processing Journal, vol. 12, no. May, pp. 1–16, 2025, doi: 10.1016/j.nlp.2025.100161.

S. Tabinda Kokab, S. Asghar, and S. Naz, “Transformer-based deep learning models for the sentiment analysis of social media data,” Array, vol. 14, no. April, pp. 1–12, 2022, doi: 10.1016/j.array.2022.100157.

W. Khan, A. Daud, K. Khan, S. Muhammad, and R. Haq, “Exploring the frontiers of deep learning and natural language processing: A comprehensive overview of key challenges and emerging trends,” Natural Language Processing Journal, vol. 4, no. July, pp. 1–31, 2023, doi: 10.1016/j.nlp.2023.100026.

M. Alfian, U. L. Yuhana, D. Siahaan, H. Munazharoh, and E. Pardede, “Out-of-Vocabulary Handling in Part-of-Speech Tagging: A Semantic Web-Driven Systematic Review,” Int. J. Semant. Web Inf. Syst., vol. 21, no. 1, pp. 1–36, 2025, doi: 10.4018/IJSWIS.388421.

O. S. Ekundayo and A. E. Ezugwu, “Deep learning: Historical overview from inception to actualization, models, applications and future trends,” Appl. Soft Comput., vol. 181, no. May, pp. 1–31, 2025, doi: 10.1016/j.asoc.2025.113378.

J. A. Liem, D. T. D. Utomo, B. Ignasio, and T. W. Cenggoro, “Deep Learning And Machine Learning For Indonesian Online Gambling Website Promotion Detection,” Procedia Comput. Sci., vol. 269, pp. 979–992, 2025, doi: 10.1016/j.procs.2025.09.040.

R. Kusumaningrum, I. Z. Nisa, R. Jayanto, R. P. Nawangsari, and A. Wibowo, “Deep learning-based application for multilevel sentiment analysis of Indonesian hotel reviews,” Heliyon, vol. 9, no. 6, pp. 1–12, 2023, doi: 10.1016/j.heliyon.2023.e17147.

J. Fehle, U. Kruschwitz, N. C. Hellwig, and C. Wolff, “Leveraging fine-tuning of large language models for aspect-based sentiment analysis in resource-scarce environments,” Knowl. Based. Syst., vol. 336, no. January, pp. 1–23, 2026, doi: 10.1016/j.knosys.2026.115277.

A. G. Ghifari, G. Y. Ananada, and K. Purwandari, “A Comparative Sentiment Analysis of Public Opinion on Indonesia’s National Football Coach Using CRNN and SVM,” Procedia Comput. Sci., vol. 269, pp. 1485–1493, 2025, doi: 10.1016/j.procs.2025.09.090.


Bila bermanfaat silahkan share artikel ini

Berikan Komentar Anda terhadap artikel Random Forest, LSTM, and IndoBERT Comparison for TikTok App Sentiment Analysis

Dimensions Badge
Article History
Submitted: 2026-04-21
Published: 2026-06-30
Abstract View: 0 times
PDF Download: 0 times
How to Cite
Saputra, I., Mesran, M., & Rasli, R. (2026). Random Forest, LSTM, and IndoBERT Comparison for TikTok App Sentiment Analysis. Building of Informatics, Technology and Science (BITS), 8(1), 570-579. https://doi.org/10.47065/bits.v8i1.9714
Issue
Section
Articles

Most read articles by the same author(s)

1 2 > >>