Klasifikasi Popularitas Buku pada Platform Digital Menggunakan Random Forest Berbasis Fitur Perilaku Pembaca
Abstract
The assessment of book popularity on digital community platforms has long relied on single indicators such as average ratings, while popularity is a multidimensional construct involving broader community engagement. This study aims to: (1) analyze reader behavior patterns as a popularity indicator for digital books in the Book-Crossing dataset; (2) develop a multiclass popularity classification model using the Random Forest algorithm; and (3) generate strategic recommendations for authors, publishers, and digital platform managers. A quantitative approach was employed utilizing secondary data from the Book-Crossing dataset, which after preprocessing yielded 10,234 books eligible for analysis. A Random Forest Classifier was constructed using six reader behavior and book metadata features, with popularity labels (Popular, Moderately Popular, Less Popular) determined through quartile distribution analysis to address class imbalance. Results show the model achieved an accuracy of 91.2% with an F1-Score of 0.94 for the Popular category, while popularity distribution follows a power law pattern in which 10.0% of books fall into the Popular category and 58.2% reside in the Less Popular category. These findings demonstrate that combining multivariable reader behavior features can produce accurate and measurable popularity classification, while also confirming the popularity bias inherent in digital content ecosystems.
Downloads
References
M. Zhou, G. H. Chen, P. Ferreira, and M. D. Smith, “Consumer Behavior in the Online Classroom: Using Video Analytics and Machine Learning to Understand the Consumption of Video Courseware,” J. Mark. Res., vol. 58, no. 6, pp. 1079–1100, 2021, doi: 10.1177/00222437211042013.
S. Daniil, M. Cuper, C. C. S. Liem, J. van Ossenbruggen, and L. Hollink, Hidden Author Bias in Book Recommendation, vol. 1, no. 1. Association for Computing Machinery, 2022. [Online]. Available: http://arxiv.org/abs/2209.00371
S. Tripathi, A. Tiwari, K. Upreti, and G. V. Radhakrishnan, “A Machine Learning Approach to Consumer Behavior Analysis in Social Media-Influenced E-Book Markets,” J. Appl. Sci. Technol. Trends, vol. 6, no. 2, pp. 242–250, 2025, doi: 10.38094/jastt62318.
E. Abdu and A. Mamuye, “Book Recommendation Using Collaborative Filtering Algorithm,” Appl. Comput. Intell. Soft Comput., vol. 2023, pp. 1–12, Mar. 2023, doi: 10.1155/2023/1514801.
F. Wayesa, M. Leranso, G. Asefa, and A. Kedir, “Pattern-based hybrid book recommendation system using semantic relationships,” Sci. Rep., vol. 13, no. 1, pp. 1–12, 2023, doi: 10.1038/s41598-023-30987-0.
Y. Afoudi, M. Lazaar, and M. Al Achhab, “Hybrid recommendation system combined content-based filtering and collaborative prediction using artificial neural network,” Simul. Model. Pract. Theory, vol. 113, p. 102375, 2021, doi: https://doi.org/10.1016/j.simpat.2021.102375.
Ega Ranaldi Pebriansyah, Susanti, Rahmiati, and Triyani Arita Fitri, “Prediction of Library Book Borrowing Patterns Using The Random Forest Algorithm,” J. Ris. Inform., vol. 7, no. 4, pp. 327–335, 2025, doi: 10.34288/jri.v7i4.409.
M. Naghiaei, H. A. Rahmani, and Y. Deldjoo, “CPFair: Personalized Consumer and Producer Fairness Re-ranking for Recommender Systems,” in Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, in SIGIR ’22. New York, NY, USA: Association for Computing Machinery, 2022, pp. 770–779. doi: 10.1145/3477495.3531959.
P. K. Sahoo, V. R. S. Dhanish, and A. N. Kumar, “Book Recommendation Using Collaborative Filtering Algorithm,” Appl. Comput. Intell. Soft Comput., vol. 2023, no. 04, pp. 365–369, 2023, doi: 10.1155/2023/1514801.
Y. Du, L. Peng, S. Dou, X. Su, and X. Ren, “Research on Personalized Book Recommendation Based on Improved Similarity Calculation and Data Filling Collaborative Filtering Algorithm,” Comput. Intell. Neurosci., vol. 2022, pp. 1–11, Sep. 2022, doi: 10.1155/2022/1900209.
C. J. Arizmendi et al., “Predicting student outcomes using digital logs of learning behaviors: Review, current standards, and suggestions for future work,” Behav. Res. Methods, vol. 55, no. 6, pp. 3026–3054, 2023, doi: 10.3758/s13428-022-01939-9.
H. Wang and N. Li, “A Click-Through Rate Prediction Method Based on Cross-Importance of Multi-Order Features,” 2024, [Online]. Available: http://arxiv.org/abs/2405.08852
M. Cinelli, G. de Francisci Morales, A. Galeazzi, W. Quattrociocchi, and M. Starnini, “The echo chamber effect on social media,” Proc. Natl. Acad. Sci. U. S. A., vol. 118, no. 9, 2021, doi: 10.1073/pnas.2023301118.
J. Prent and M. Mansoury, “Correcting Popularity Bias in Recommender Systems via Item Loss Equalization,” 2024. [Online]. Available: https://api.semanticscholar.org/CorpusID:273186697
A. Klimashevskaia, D. Jannach, M. Elahi, and C. Trattner, “A survey on popularity bias in recommender systems,” User Model. User-adapt. Interact., vol. 34, no. 5, pp. 1777–1834, 2024, doi: 10.1007/s11257-024-09406-0.
M. Mansoury, F. Duijvestijn, and I. Mourabet, “Potential Factors Leading to Popularity Unfairness in Recommender Systems: A User-Centered Analysis,” pp. 1–16, 2023, [Online]. Available: http://arxiv.org/abs/2310.02961
Y. Zoralioglu and E. Yalcin, “Dynamic feedback loops in recommender systems: Analyzing fairness, popularity bias, and user group disparities,” J. Intell. Inf. Syst., 2026, doi: 10.1007/s10844-026-01025-y.
X. Yang et al., “Click-through rate prediction using transfer learning with fine-tuned parameters,” Inf. Sci. (Ny)., vol. 612, pp. 188–200, 2022, doi: https://doi.org/10.1016/j.ins.2022.08.009.
K. Xiang, Research on Deep Learning Recommendation System Based on Book Scoring, jurnal inf. Sci 2025. doi: 10.1145/3718751.3718838.
Pratama, “Analisis Sentimen BRImo dan BCA Mobile Menggunakan Support Vector Machine dan Lexicon Based,” Jurnal Ilmu Komputer dan Informatika., vol. 3, no. 2, pp. 1439-1450, 2023.
Damayanti et al., “Sentiment Analysis of Alfagift Application User Reviews Using Long Short-Term Memory ( LSTM ) and Support Vector Machine ( SVM ) Methods,” Jurnal DECODE, vol. 4, no. 2, pp. 509-521, 2024, doi: 10.51454/decode.v4i2.478.
Arumi,, Implementation of Naïve bayes Method for Predictor Prevalence Level for Malnutrition Toddlers in Magelang City. Jurnal RESTI.,2023, doi: 10.29207/resti.v7i2.4438
Jayaditya, I. K. A., & Kadyanan, I. G. A. G. A., Implementasi Random Forest pada Klasifikasi Penyakit Kardiovaskular dengan Hyperparameter Tuning Grid Search. Jurnal Nasional Teknologi Informasi dan Aplikasinya (JNATIA), 2023, doi: 10.24843/JNATIA.2023.v02.i01.p25.
Sholihah, N., & Hermawan, A., Implementation of Random Forest and SMOTE Methods for Economic Status Classification in Cirebon City. Jurnal Teknik Informatika (JUTIF), 2023, doi: 10.52436/1.jutif.2023.4.6.1135.
Bila bermanfaat silahkan share artikel ini
Berikan Komentar Anda terhadap artikel Klasifikasi Popularitas Buku pada Platform Digital Menggunakan Random Forest Berbasis Fitur Perilaku Pembaca
Pages: 347-356
Copyright (c) 2026 Al Imran, Sandi Firmasyah, Siti Intan Nurfadilah, Muhammad Alamsyah, Farizi Ilham

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under Creative Commons Attribution 4.0 International License that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (Refer to The Effect of Open Access).


