AI-Assisted Labeling for Indonesian Hadith Classification using Four Thematic Categories: Evaluating TF-IDF+SVM and IndoBERT
Abstract
Hadith is the second primary source of Islamic law after the Qur'an, and its large volume makes manual thematic classification time-consuming and inefficient. This study proposes an AI-assisted labeling approach to construct an Indonesian thematic hadith dataset and evaluates the performance of two text classification methods, namely TF-IDF + Support Vector Machine (SVM) and IndoBERT. The dataset consists of 6,600 Indonesian-translated Sahih Bukhari hadiths collected from the Hadith API and categorized into four thematic classes: aqidah, ibadah, akhlak, and muamalah. The annotation process employed Gemini 2.5 Flash with a structured prompt and JSON-based output format, followed by validation performed by a hadith researcher validation on a randomly selected 5% sample, achieving an overall agreement of 73.3%. The annotated data were divided into training and testing sets using an 80:20 stratified split. Model performance was evaluated using Accuracy, Macro F1-score, and Weighted F1-score. Experimental results show that TF-IDF + SVM achieved an Accuracy of 73.1%, a Macro F1-score of 70.1%, and a Weighted F1-score of 73.0%, while IndoBERT achieved an Accuracy of 72.3%, a Macro F1-score of 69.7%, and a Weighted F1-score of 72.2%. The results indicate that the conventional TF-IDF + SVM approach slightly outperformed the Transformer-based IndoBERT model on the proposed dataset. The main contributions of this study are the construction of an Indonesian thematic hadith dataset through AI-assisted labeling and a comparative evaluation of conventional and Transformer-based methods for Indonesian hadith classification.
Downloads
References
Adigunawan, & Fathoni, R. (2023). Machine Learning and Transformer-based Model for Sentiment Analysis of Indonesian E-Commerce Reviews. Indonesian Journal of Computer Science, 12(2), 284–301. http://ijcs.stmikindonesia.ac.id/ijcs/index.php/ijcs/article/view/3135
Bartlett, C. W., Bossenbroek, J., Ueyama, Y., McCallinhart, P., Peters, O. A., Santillan, D. A., Santillan, M. K., Trask, A. J., & Ray, W. C. (2023). Invasive or More Direct Measurements Can Provide an Objective Early-Stopping Ceiling for Training Deep Neural Networks on Non-invasive or Less-Direct Biomedical Data. SN Computer Science, 4(2), 1–12. https://doi.org/10.1007/s42979-022-01553-8
Cahyawijaya, S., Lovenia, H., Aji, A. F., Winata, G. I., Wilie, B., Koto, F., Mahendra, R., Wibisono, C., Romadhony, A., Vincentio, K., Santoso, J., Moeljadi, D., Wirawan, C., Hudi, F., Wicaksono, M. S., Parmonangan, I. H., Alfina, I., Putra, I. F., Rahmadani, S., … Purwarianti, A. (2023). NusaCrowd: Open Source Initiative for Indonesian NLP Resources. Proceedings of the Annual Meeting of the Association for Computational Linguistics, 13745–13818. https://doi.org/10.18653/v1/2023.findings-acl.868
Darfiansa, L. S., Fitriyani, & Larasati, S. S. A. (2025). Optimizing IndoBERT for Revised Bloom’s Taxonomy Question Classification Using Neural Network Classifier. Journal of Information Systems Engineering and Business Intelligence, 11(2), 226–237. https://doi.org/10.20473/jisebi.11.2.226-237
Geng, S., Cooper, H., Moskal, M., Jenkins, S., Berman, J., Ranchin, N., West, R., Horvitz, E., & Nori, H. (2025). JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language Models. http://arxiv.org/abs/2501.10868
Gilardi, F., Alizadeh, M., & Kubli, M. (2023). ChatGPT outperforms crowd workers for text-annotation tasks. Proceedings of the National Academy of Sciences of the United States of America, 120(30), 1–3. https://doi.org/10.1073/pnas.2305016120
Grandini, M., Bagli, E., & Visani, G. (2020). Metrics for Multi-Class Classification: an Overview. 1–17. http://arxiv.org/abs/2008.05756
Handayani, R. N., Najiyah, I., & Wisnuwardana, D. A. (2023). Classification of Bulughul Maraam Categories: Prohibitions, Recommendations, and Information Using Extreme Learning Machine and Fasttext. Jurnal Online Informatika, 8(2), 242–251. https://doi.org/10.15575/join.v8i2.1205
Irugalbandara, C. (2024). Meaning Typed Prompting: A Technique for Efficient, Reliable Structured Output Generation. http://arxiv.org/abs/2410.18146
Joseph, V. R. (2022). Optimal ratio for data splitting. Statistical Analysis and Data Mining, 15(4), 531–538. https://doi.org/10.1002/sam.11583
Kowsari, K., Meimandi, K. J., Heidarysafa, M., Mendu, S., Barnes, L., & Brown, D. (2019). Text classification algorithms: A survey. Information (Switzerland), 10(4), 1–68. https://doi.org/10.3390/info10040150
Lutfiah, D., Harahap, N. S., Yanto, F., & Cynthia, E. P. (2026). Klasifikasi Multi-Label Hadits Shahih Muslim Menggunakan Metode Support Vector Machine (SVM). Bulletin of Data Science, 5(3), 158–169. https://doi.org/10.47065/bulletinds.v5i3.10147
Minaee, S., Kalchbrenner, N., Cambria, E., Nikzad, N., Chenaghlu, M., & Gao, J. (2021). Deep Learning Based Text Classification: A Comprehensive Review. 1(1), 1–43. http://arxiv.org/abs/2004.03705
Nabiilah, G. Z., Al Faraby, S., & Dwifebri Purbolaksono, M. (2021). Classification of Hadith Topic of Indonesian Translation Using K-Nearest Neighbor and Chi-Square. International Journal on Information and Communication Technology (IJoICT), 7(2), 11–22. https://doi.org/10.21108/ijoict.v7i2.573
Ni’mah. SM, K., Arifi, A., & Abror, I. (2024). Hadith as a Source of Islamic Law: Its Role and Significance. Studi Multidisipliner: Jurnal Kajian Keislaman, 11(2), 193–204. https://doi.org/10.24952/multidisipliner.v11i2.13302
Rainio, O., Teuho, J., & Klén, R. (2024). Evaluation metrics and statistical tests for machine learning. Scientific Reports, 14(1), 1–14. https://doi.org/10.1038/s41598-024-56706-x
Ramadhan, T. I., Supriatman, A., & Kurniawan, T. R. (2024). Passage Retrieval untuk Question Answering Bahasa Indonesia Menggunakan BERT dan FAISS. Jurnal Algoritma, 21(2), 156–163. https://doi.org/10.33364/algoritma/v.21-2.2100
Rasenda, R., Fabrianto, L., & Faizah, N. M. (2024). Optimizing Hadith Classification With Neural Networks: a Study on Bukhari and Muslim Texts. JIKO (Jurnal Informatika Dan Komputer), 7(2), 145–149. https://doi.org/10.33387/jiko.v7i2.8732
Sahoo, P., Singh, A. K., Saha, S., Jain, V., Mondal, S., & Chadha, A. (2025). A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications. http://arxiv.org/abs/2402.07927
Sunendar, N. S., & Saputra, I. (2025). Comparative Performance of Transformer and Lstm Models for Indonesian Information Retrieval With Indobert. Jurnal Pilar Nusa Mandiri, 21(2), 228–233. https://doi.org/10.33480/pilar.v21i2.6920
Susanti, S., Najiyah, I., Ramdhani, Y., Herliana, A., Muckti, M. K., & Oktaviani, F. R. (2025). Searching Sahih Hadiths Based on Queries using Neural Models and FastText. Journal of Applied Data Sciences, 6(1), 272–285. https://doi.org/10.47738/jads.v6i1.467
Taha, K., Yoo, P. D., Yeun, C., Homouz, D., & Taha, A. (2024). A comprehensive survey of text classification techniques and their research applications: Observational and experimental insights. Computer Science Review, 54(August), 100664. https://doi.org/10.1016/j.cosrev.2024.100664
Taufiqurrahman, F., Faraby, S. Al, & Purbolaksono, M. D. (2021). Klasifikasi Teks Multi Label pada Hadis Terjemahan Bahasa Indonesia Menggunakan Chi Square dan SVM. E-Proceeding of Engineering, 8(5), 10650–10659.
Tsoumakas, G., & Katakis, I. (2009). Multi-Label Classification: An Overview. Database Technologies: Concepts, Methodologies, Tools, and Applications: Volumes 1-4, 1, 309–319. https://doi.org/10.4018/978-1-60566-058-5.ch021
Zakiah, R., Ahmad, N., Safaat Harahap, N., Agustian, S., Iskandar, I., Sanjaya, S., Studi, P., Informatika, T., Sains, F., & Teknologi, D. (2025). MALCOM: Indonesian Journal of Machine Learning and Computer Science Comparison of Random Forest and Long Short-Term Memory Performance in Multilabel Text Classification of Bukhari Hadith Translation . 5(July), 862–874.
Bila bermanfaat silahkan share artikel ini
Berikan Komentar Anda terhadap artikel AI-Assisted Labeling for Indonesian Hadith Classification using Four Thematic Categories: Evaluating TF-IDF+SVM and IndoBERT
Pages: 1182-1191
Copyright (c) 2026 Domi Sepri, Ahmad Fauzi, Firman Firman, Muhammad Nabil

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under Creative Commons Attribution 4.0 International License that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (Refer to The Effect of Open Access).













