AI-Assisted Labeling for Indonesian Hadith Classification using Four Thematic Categories: Evaluating TF-IDF+SVM and IndoBERT


  • Domi Sepri * Mail Universitas Islam Negeri Imam Bonjol Padang, Padang, Indonesia
  • Ahmad Fauzi Universitas Islam Negeri Imam Bonjol Padang, Padang, Indonesia
  • Firman Firman Universitas Islam Negeri Imam Bonjol Padang, Padang, Indonesia
  • Muhammad Nabil Universitas Islam Negeri Imam Bonjol Padang, Padang, Indonesia
  • (*) Corresponding Author
Keywords: AI-assisted Labeling; Hadith Classification; IndoBERT; Support Vector Machine; Natural Language Processing

Abstract

Hadith is the second primary source of Islamic law after the Qur'an, and its large volume makes manual thematic classification time-consuming and inefficient. This study proposes an AI-assisted labeling approach to construct an Indonesian thematic hadith dataset and evaluates the performance of two text classification methods, namely TF-IDF + Support Vector Machine (SVM) and IndoBERT. The dataset consists of 6,600 Indonesian-translated Sahih Bukhari hadiths collected from the Hadith API and categorized into four thematic classes: aqidah, ibadah, akhlak, and muamalah. The annotation process employed Gemini 2.5 Flash with a structured prompt and JSON-based output format, followed by validation performed by a hadith researcher validation on a randomly selected 5% sample, achieving an overall agreement of 73.3%. The annotated data were divided into training and testing sets using an 80:20 stratified split. Model performance was evaluated using Accuracy, Macro F1-score, and Weighted F1-score. Experimental results show that TF-IDF + SVM achieved an Accuracy of 73.1%, a Macro F1-score of 70.1%, and a Weighted F1-score of 73.0%, while IndoBERT achieved an Accuracy of 72.3%, a Macro F1-score of 69.7%, and a Weighted F1-score of 72.2%. The results indicate that the conventional TF-IDF + SVM approach slightly outperformed the Transformer-based IndoBERT model on the proposed dataset. The main contributions of this study are the construction of an Indonesian thematic hadith dataset through AI-assisted labeling and a comparative evaluation of conventional and Transformer-based methods for Indonesian hadith classification.

Downloads

Download data is not yet available.

References

Adigunawan, & Fathoni, R. (2023). Machine Learning and Transformer-based Model for Sentiment Analysis of Indonesian E-Commerce Reviews. Indonesian Journal of Computer Science, 12(2), 284–301. http://ijcs.stmikindonesia.ac.id/ijcs/index.php/ijcs/article/view/3135

Bartlett, C. W., Bossenbroek, J., Ueyama, Y., McCallinhart, P., Peters, O. A., Santillan, D. A., Santillan, M. K., Trask, A. J., & Ray, W. C. (2023). Invasive or More Direct Measurements Can Provide an Objective Early-Stopping Ceiling for Training Deep Neural Networks on Non-invasive or Less-Direct Biomedical Data. SN Computer Science, 4(2), 1–12. https://doi.org/10.1007/s42979-022-01553-8

Cahyawijaya, S., Lovenia, H., Aji, A. F., Winata, G. I., Wilie, B., Koto, F., Mahendra, R., Wibisono, C., Romadhony, A., Vincentio, K., Santoso, J., Moeljadi, D., Wirawan, C., Hudi, F., Wicaksono, M. S., Parmonangan, I. H., Alfina, I., Putra, I. F., Rahmadani, S., … Purwarianti, A. (2023). NusaCrowd: Open Source Initiative for Indonesian NLP Resources. Proceedings of the Annual Meeting of the Association for Computational Linguistics, 13745–13818. https://doi.org/10.18653/v1/2023.findings-acl.868

Darfiansa, L. S., Fitriyani, & Larasati, S. S. A. (2025). Optimizing IndoBERT for Revised Bloom’s Taxonomy Question Classification Using Neural Network Classifier. Journal of Information Systems Engineering and Business Intelligence, 11(2), 226–237. https://doi.org/10.20473/jisebi.11.2.226-237

Geng, S., Cooper, H., Moskal, M., Jenkins, S., Berman, J., Ranchin, N., West, R., Horvitz, E., & Nori, H. (2025). JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language Models. http://arxiv.org/abs/2501.10868

Gilardi, F., Alizadeh, M., & Kubli, M. (2023). ChatGPT outperforms crowd workers for text-annotation tasks. Proceedings of the National Academy of Sciences of the United States of America, 120(30), 1–3. https://doi.org/10.1073/pnas.2305016120

Grandini, M., Bagli, E., & Visani, G. (2020). Metrics for Multi-Class Classification: an Overview. 1–17. http://arxiv.org/abs/2008.05756

Handayani, R. N., Najiyah, I., & Wisnuwardana, D. A. (2023). Classification of Bulughul Maraam Categories: Prohibitions, Recommendations, and Information Using Extreme Learning Machine and Fasttext. Jurnal Online Informatika, 8(2), 242–251. https://doi.org/10.15575/join.v8i2.1205

Irugalbandara, C. (2024). Meaning Typed Prompting: A Technique for Efficient, Reliable Structured Output Generation. http://arxiv.org/abs/2410.18146

Joseph, V. R. (2022). Optimal ratio for data splitting. Statistical Analysis and Data Mining, 15(4), 531–538. https://doi.org/10.1002/sam.11583

Kowsari, K., Meimandi, K. J., Heidarysafa, M., Mendu, S., Barnes, L., & Brown, D. (2019). Text classification algorithms: A survey. Information (Switzerland), 10(4), 1–68. https://doi.org/10.3390/info10040150

Lutfiah, D., Harahap, N. S., Yanto, F., & Cynthia, E. P. (2026). Klasifikasi Multi-Label Hadits Shahih Muslim Menggunakan Metode Support Vector Machine (SVM). Bulletin of Data Science, 5(3), 158–169. https://doi.org/10.47065/bulletinds.v5i3.10147

Minaee, S., Kalchbrenner, N., Cambria, E., Nikzad, N., Chenaghlu, M., & Gao, J. (2021). Deep Learning Based Text Classification: A Comprehensive Review. 1(1), 1–43. http://arxiv.org/abs/2004.03705

Nabiilah, G. Z., Al Faraby, S., & Dwifebri Purbolaksono, M. (2021). Classification of Hadith Topic of Indonesian Translation Using K-Nearest Neighbor and Chi-Square. International Journal on Information and Communication Technology (IJoICT), 7(2), 11–22. https://doi.org/10.21108/ijoict.v7i2.573

Ni’mah. SM, K., Arifi, A., & Abror, I. (2024). Hadith as a Source of Islamic Law: Its Role and Significance. Studi Multidisipliner: Jurnal Kajian Keislaman, 11(2), 193–204. https://doi.org/10.24952/multidisipliner.v11i2.13302

Rainio, O., Teuho, J., & Klén, R. (2024). Evaluation metrics and statistical tests for machine learning. Scientific Reports, 14(1), 1–14. https://doi.org/10.1038/s41598-024-56706-x

Ramadhan, T. I., Supriatman, A., & Kurniawan, T. R. (2024). Passage Retrieval untuk Question Answering Bahasa Indonesia Menggunakan BERT dan FAISS. Jurnal Algoritma, 21(2), 156–163. https://doi.org/10.33364/algoritma/v.21-2.2100

Rasenda, R., Fabrianto, L., & Faizah, N. M. (2024). Optimizing Hadith Classification With Neural Networks: a Study on Bukhari and Muslim Texts. JIKO (Jurnal Informatika Dan Komputer), 7(2), 145–149. https://doi.org/10.33387/jiko.v7i2.8732

Sahoo, P., Singh, A. K., Saha, S., Jain, V., Mondal, S., & Chadha, A. (2025). A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications. http://arxiv.org/abs/2402.07927

Sunendar, N. S., & Saputra, I. (2025). Comparative Performance of Transformer and Lstm Models for Indonesian Information Retrieval With Indobert. Jurnal Pilar Nusa Mandiri, 21(2), 228–233. https://doi.org/10.33480/pilar.v21i2.6920

Susanti, S., Najiyah, I., Ramdhani, Y., Herliana, A., Muckti, M. K., & Oktaviani, F. R. (2025). Searching Sahih Hadiths Based on Queries using Neural Models and FastText. Journal of Applied Data Sciences, 6(1), 272–285. https://doi.org/10.47738/jads.v6i1.467

Taha, K., Yoo, P. D., Yeun, C., Homouz, D., & Taha, A. (2024). A comprehensive survey of text classification techniques and their research applications: Observational and experimental insights. Computer Science Review, 54(August), 100664. https://doi.org/10.1016/j.cosrev.2024.100664

Taufiqurrahman, F., Faraby, S. Al, & Purbolaksono, M. D. (2021). Klasifikasi Teks Multi Label pada Hadis Terjemahan Bahasa Indonesia Menggunakan Chi Square dan SVM. E-Proceeding of Engineering, 8(5), 10650–10659.

Tsoumakas, G., & Katakis, I. (2009). Multi-Label Classification: An Overview. Database Technologies: Concepts, Methodologies, Tools, and Applications: Volumes 1-4, 1, 309–319. https://doi.org/10.4018/978-1-60566-058-5.ch021

Zakiah, R., Ahmad, N., Safaat Harahap, N., Agustian, S., Iskandar, I., Sanjaya, S., Studi, P., Informatika, T., Sains, F., & Teknologi, D. (2025). MALCOM: Indonesian Journal of Machine Learning and Computer Science Comparison of Random Forest and Long Short-Term Memory Performance in Multilabel Text Classification of Bukhari Hadith Translation . 5(July), 862–874.


Bila bermanfaat silahkan share artikel ini

Berikan Komentar Anda terhadap artikel AI-Assisted Labeling for Indonesian Hadith Classification using Four Thematic Categories: Evaluating TF-IDF+SVM and IndoBERT

Dimensions Badge
Article History
Published: 2026-07-29
Abstract View: 27 times
PDF Download: 11 times
Issue
Section
Articles