Optimasi Model Retrieval-Augmented Generation Menggunakan Algoritma Indeks HNSW Lokal pada Small Language Model untuk Mitigasi Halusinasi Hadis Bukhari


  • Rizqi Ari Putra * Mail Universitas Islam Negeri Maulana Malik Ibrahim Malang, Malang, Indonesia
  • Suhartono Suhartono Universitas Islam Negeri Maulana Malik Ibrahim Malang, Malang, Indonesia
  • Muhammad Faisal Universitas Islam Negeri Maulana Malik Ibrahim Malang, Malang, Indonesia
  • (*) Corresponding Author
Keywords: Hadith Bukhari; Hallucination Mitigation; Retrieval-Augmented Generation; Small Language Model; Vector Quantization

Abstract

Small Language Models (SLMs) offer high local computational efficiency but possess a systemic vulnerability to information hallucination. This vulnerability becomes a critical risk when the model is applied to sensitive domains demanding absolute accuracy, such as Sharia law and Islamic sacred literature. The inaccurate citation of religious texts can lead to theological misguidance. To address this issue, this study proposes a low-memory, local Retrieval-Augmented Generation (RAG) architectural solution. This system is designed by integrating the open-source Llama-3.2-1B-Instruct model with the PostgreSQL pgvector vector database. Document retrieval optimization is performed based on the HNSW (Hierarchical Navigable Small World) indexing algorithm and a 16-bit precision quantization technique (halfvec) on 7,003 chunks of the Sahih al-Bukhari Hadith corpus. The primary objective of this research is to design a high-precision hallucination mitigation system that operates independently (on-premise), while making a tangible contribution to the development of a low-cost digital theological assistant that preserves privacy and data sovereignty. Mitigation efficacy was automatically evaluated using the Ragas framework against 50 theological test queries, while database efficiency was physically tested on consumer-grade local computer hardware. Experimental results indicate that the proposed architecture is capable of significantly improving the faithfulness metric by 117.4% (from 0.4120 to 0.8960) and answer relevance by 50.0% (from 0.6120 to 0.9180), while simultaneously accelerating inference response time by up to 50.7%. On the database side, halfvec quantization successfully reduced physical table storage space by 36.2% and accelerated index construction time by 16.5% with an absolute accuracy (recall) retention rate of 1.0000. This study proves that a high-precision religious virtual assistant is highly feasible to execute independently without relying on third-party cloud computing services.

Downloads

Download data is not yet available.

References

P. Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” in NIPS ’20: Proceedings of the 34th International Conference on Neural Information Processing Systems, Dec. 2020, pp. 9459–9474. doi: https://doi.org/10.48550/arXiv.2005.11401Focusto.

C. D. Manning, P. Raghavan, and H. Schütze, Introduction to Information Retrieval. Cambridge University Press, 2009.

Abhimanyu Dubey et al., “The Llama 3 Herd of Models,” arXiv preprint arXiv:2407.21783, 2024, doi: https://doi.org/10.48550/arXiv.2407.21783.

Y. A. Malkov and D. A. Yashunin, “Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 42, no. 4, pp. 824–836, Apr. 2020, doi: https://doi.org/10.1109/TPAMI.2018.2889473.

OpenAI, “New embedding models and API updates,” Jan. 2024.

Mohammed Amine Mouhoub, “Islamic Large Language Models: From Knowledge Acquisition to Trustworthy and Hallucination-Resistant AI,” arXiv preprint arXiv:2606.16629, Jun. 2026, doi: https://doi.org/10.48550/arXiv.2606.16629.

Zahra Khalila et al., “Investigating Retrieval-Augmented Generation in Quranic Studies: A Study of 13 Open-Source Large Language Models,” International Journal of Advanced Computer Science and Applications, vol. 16, no. 2, May 2025, doi: https://doi.org/10.48550/arXiv.2503.16581.

E. Elrefai, T. Khaled, and A. Soliman, “ThinkDrill at IslamicEval 2025 Shared Task: LLM Hybrid Approach for Qur’an and Hadith Question Answering,” in Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks, Stroudsburg, PA, USA: Association for Computational Linguistics, Nov. 2025, pp. 528–533. doi: 10.18653/v1/2025.arabicnlp-sharedtasks.72.

Lianmin Zheng et al., “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena,” in NIPS ’23: Proceedings of the 37th International Conference on Neural Information Processing Systems, Dec. 2023, pp. 46595–46626. doi: https://doi.org/10.52202/075280-2020.

Shahul Es, Jithin James, Luis Espinosa Anke, and Steven Schockaert, “RAGAS: Automated Evaluation of Retrieval Augmented Generation,” in Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, Association for Computational Linguistics, Mar. 2024, pp. 150–158. doi: https://doi.org/10.18653/v1/2024.eacl-demo.16.

Fei Fang, Yi Liu, and Chen Qian, “d-HNSW: A High-performance Vector Search Engine on Disaggregated Memory,” SIGMETRICS Abstracts ’26: Abstracts of the 2026 ACM SIGMETRICS International Conference on Measurement and Modeling of Computer Systems, pp. 229–231, Jun. 2026, doi: https://doi.org/10.1145/3801489.3806852.

Nils Reimers and Iryna Gurevych, “Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, Nov. 2019, pp. 3982–3992. doi: https://doi.org/10.48550/arXiv.1908.10084.

Michael Armbrust, Ali Ghodsi, Reynold Xin, and Matei Zaharia, “Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics,” in Proceedings of the 11th Annual Conference on Innovative Data Systems Research (CIDR), Jan. 2021. doi: http://cidrdb.org/cidr2021/papers/cidr2021_paper17.pdf.

Yunfan Gao et al., “Retrieval-Augmented Generation for Large Language Models: A Survey,” arXiv preprint arXiv:2312.10997, Mar. 2024, doi: http://doi.org/10.48550/arXiv.2312.10997.

Xiaohua Wang et al., “Searching for Best Practices in Retrieval-Augmented Generation,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Nov. 2024, pp. 17716–17736. doi: https://doi.org/10.48550/arXiv.2407.01219.

Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer, “LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale,” in Advances in Neural Information Processing Systems, Dec. 2022, pp. 30381–30393. doi: https://doi.org/10.48550/arXiv.2208.07339.

Shahul Es, Jithin James, Luis Espinosa-Anke, and Steven Schockaert, “Ragas: Automated Evaluation of Retrieval Augmented Generation,” in arXiv preprint arXiv:2309.15217, Sep. 2023, pp. 1–8. doi: https://doi.org/10.48550/arXiv.2309.15217.

Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa, “Large Language Models are Zero-Shot Reasoners,” in Advances in Neural Information Processing Systems, Jan. 2022, pp. 22199–22213. doi: https://doi.org/10.48550/arXiv.2205.11916.

James Jie Pan, Jianguo Wang, and Guoliang Li, “Survey of Vector Database Management Systems,” The VLDB Journal, vol. 33, pp. 1–25, Jul. 2024, doi: https://doi.org/10.1007/s00778-024-00864-x.

Lei Huang et al., “A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions,” arXiv preprint arXiv:2311.05232, vol. 43, Mar. 2025, doi: https://doi.org/10.1145/3703155.

Saqib Hakak et al., “Digital Hadith authentication: Recent advances, open challenges, and future directions,” Transactions on Emerging Telecommunications Technologies, vol. 33, no. 10, p. e4604, Jun. 2022, doi: https://doi.org/10.1002/ett.3977.

Chien Van Nguyen et al., “A Survey on Small Language Models,” arXiv preprint arXiv:2410.20011, 2024, doi: https://doi.org/10.48550/arXiv.2410.20011.

N. Sengupta et al., “Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models,” arXiv e-prints, p. arXiv:2308.16149, Aug. 2023, doi: 10.48550/arXiv.2308.16149.


Bila bermanfaat silahkan share artikel ini

Berikan Komentar Anda terhadap artikel Optimasi Model Retrieval-Augmented Generation Menggunakan Algoritma Indeks HNSW Lokal pada Small Language Model untuk Mitigasi Halusinasi Hadis Bukhari

Dimensions Badge
Article History
Submitted: 2026-06-26
Published: 2026-07-14
Abstract View: 41 times
PDF Download: 26 times
How to Cite
Putra, R., Suhartono, S., & Faisal, M. (2026). Optimasi Model Retrieval-Augmented Generation Menggunakan Algoritma Indeks HNSW Lokal pada Small Language Model untuk Mitigasi Halusinasi Hadis Bukhari. Journal of Information System Research (JOSH), 7(4), 953-966. https://doi.org/10.47065/josh.v7i4.10467
Issue
Section
Articles

Most read articles by the same author(s)