Optimasi Model Retrieval-Augmented Generation Menggunakan Algoritma Indeks HNSW Lokal pada Small Language Model untuk Mitigasi Halusinasi Hadis Bukhari
Abstract
Small Language Models (SLMs) offer high local computational efficiency but possess a systemic vulnerability to information hallucination. This vulnerability becomes a critical risk when the model is applied to sensitive domains demanding absolute accuracy, such as Sharia law and Islamic sacred literature. The inaccurate citation of religious texts can lead to theological misguidance. To address this issue, this study proposes a low-memory, local Retrieval-Augmented Generation (RAG) architectural solution. This system is designed by integrating the open-source Llama-3.2-1B-Instruct model with the PostgreSQL pgvector vector database. Document retrieval optimization is performed based on the HNSW (Hierarchical Navigable Small World) indexing algorithm and a 16-bit precision quantization technique (halfvec) on 7,003 chunks of the Sahih al-Bukhari Hadith corpus. The primary objective of this research is to design a high-precision hallucination mitigation system that operates independently (on-premise), while making a tangible contribution to the development of a low-cost digital theological assistant that preserves privacy and data sovereignty. Mitigation efficacy was automatically evaluated using the Ragas framework against 50 theological test queries, while database efficiency was physically tested on consumer-grade local computer hardware. Experimental results indicate that the proposed architecture is capable of significantly improving the faithfulness metric by 117.4% (from 0.4120 to 0.8960) and answer relevance by 50.0% (from 0.6120 to 0.9180), while simultaneously accelerating inference response time by up to 50.7%. On the database side, halfvec quantization successfully reduced physical table storage space by 36.2% and accelerated index construction time by 16.5% with an absolute accuracy (recall) retention rate of 1.0000. This study proves that a high-precision religious virtual assistant is highly feasible to execute independently without relying on third-party cloud computing services.
Downloads
References
P. Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” in NIPS ’20: Proceedings of the 34th International Conference on Neural Information Processing Systems, Dec. 2020, pp. 9459–9474. doi: https://doi.org/10.48550/arXiv.2005.11401Focusto.
C. D. Manning, P. Raghavan, and H. Schütze, Introduction to Information Retrieval. Cambridge University Press, 2009.
Abhimanyu Dubey et al., “The Llama 3 Herd of Models,” arXiv preprint arXiv:2407.21783, 2024, doi: https://doi.org/10.48550/arXiv.2407.21783.
Y. A. Malkov and D. A. Yashunin, “Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 42, no. 4, pp. 824–836, Apr. 2020, doi: https://doi.org/10.1109/TPAMI.2018.2889473.
OpenAI, “New embedding models and API updates,” Jan. 2024.
Mohammed Amine Mouhoub, “Islamic Large Language Models: From Knowledge Acquisition to Trustworthy and Hallucination-Resistant AI,” arXiv preprint arXiv:2606.16629, Jun. 2026, doi: https://doi.org/10.48550/arXiv.2606.16629.
Zahra Khalila et al., “Investigating Retrieval-Augmented Generation in Quranic Studies: A Study of 13 Open-Source Large Language Models,” International Journal of Advanced Computer Science and Applications, vol. 16, no. 2, May 2025, doi: https://doi.org/10.48550/arXiv.2503.16581.
E. Elrefai, T. Khaled, and A. Soliman, “ThinkDrill at IslamicEval 2025 Shared Task: LLM Hybrid Approach for Qur’an and Hadith Question Answering,” in Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks, Stroudsburg, PA, USA: Association for Computational Linguistics, Nov. 2025, pp. 528–533. doi: 10.18653/v1/2025.arabicnlp-sharedtasks.72.
Lianmin Zheng et al., “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena,” in NIPS ’23: Proceedings of the 37th International Conference on Neural Information Processing Systems, Dec. 2023, pp. 46595–46626. doi: https://doi.org/10.52202/075280-2020.
Shahul Es, Jithin James, Luis Espinosa Anke, and Steven Schockaert, “RAGAS: Automated Evaluation of Retrieval Augmented Generation,” in Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, Association for Computational Linguistics, Mar. 2024, pp. 150–158. doi: https://doi.org/10.18653/v1/2024.eacl-demo.16.
Fei Fang, Yi Liu, and Chen Qian, “d-HNSW: A High-performance Vector Search Engine on Disaggregated Memory,” SIGMETRICS Abstracts ’26: Abstracts of the 2026 ACM SIGMETRICS International Conference on Measurement and Modeling of Computer Systems, pp. 229–231, Jun. 2026, doi: https://doi.org/10.1145/3801489.3806852.
Nils Reimers and Iryna Gurevych, “Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, Nov. 2019, pp. 3982–3992. doi: https://doi.org/10.48550/arXiv.1908.10084.
Michael Armbrust, Ali Ghodsi, Reynold Xin, and Matei Zaharia, “Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics,” in Proceedings of the 11th Annual Conference on Innovative Data Systems Research (CIDR), Jan. 2021. doi: http://cidrdb.org/cidr2021/papers/cidr2021_paper17.pdf.
Yunfan Gao et al., “Retrieval-Augmented Generation for Large Language Models: A Survey,” arXiv preprint arXiv:2312.10997, Mar. 2024, doi: http://doi.org/10.48550/arXiv.2312.10997.
Xiaohua Wang et al., “Searching for Best Practices in Retrieval-Augmented Generation,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Nov. 2024, pp. 17716–17736. doi: https://doi.org/10.48550/arXiv.2407.01219.
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer, “LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale,” in Advances in Neural Information Processing Systems, Dec. 2022, pp. 30381–30393. doi: https://doi.org/10.48550/arXiv.2208.07339.
Shahul Es, Jithin James, Luis Espinosa-Anke, and Steven Schockaert, “Ragas: Automated Evaluation of Retrieval Augmented Generation,” in arXiv preprint arXiv:2309.15217, Sep. 2023, pp. 1–8. doi: https://doi.org/10.48550/arXiv.2309.15217.
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa, “Large Language Models are Zero-Shot Reasoners,” in Advances in Neural Information Processing Systems, Jan. 2022, pp. 22199–22213. doi: https://doi.org/10.48550/arXiv.2205.11916.
James Jie Pan, Jianguo Wang, and Guoliang Li, “Survey of Vector Database Management Systems,” The VLDB Journal, vol. 33, pp. 1–25, Jul. 2024, doi: https://doi.org/10.1007/s00778-024-00864-x.
Lei Huang et al., “A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions,” arXiv preprint arXiv:2311.05232, vol. 43, Mar. 2025, doi: https://doi.org/10.1145/3703155.
Saqib Hakak et al., “Digital Hadith authentication: Recent advances, open challenges, and future directions,” Transactions on Emerging Telecommunications Technologies, vol. 33, no. 10, p. e4604, Jun. 2022, doi: https://doi.org/10.1002/ett.3977.
Chien Van Nguyen et al., “A Survey on Small Language Models,” arXiv preprint arXiv:2410.20011, 2024, doi: https://doi.org/10.48550/arXiv.2410.20011.
N. Sengupta et al., “Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models,” arXiv e-prints, p. arXiv:2308.16149, Aug. 2023, doi: 10.48550/arXiv.2308.16149.
Bila bermanfaat silahkan share artikel ini
Berikan Komentar Anda terhadap artikel Optimasi Model Retrieval-Augmented Generation Menggunakan Algoritma Indeks HNSW Lokal pada Small Language Model untuk Mitigasi Halusinasi Hadis Bukhari
Pages: 953-966
Copyright (c) 2026 Rizqi Ari Putra, Suhartono Suhartono, Muhammad Faisal

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under Creative Commons Attribution 4.0 International License that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (Refer to The Effect of Open Access).






















