Comparative Evaluation of Local Llama and Mistral Models in RAG for Helpdesk Documentation


  • Septian Pratama * Mail Universitas Pamulang, Tangerang Selatan, Indonesia
  • Sajarwo Anggai Universitas Pamulang, Tangerang Selatan, Indonesia
  • Murni Handayani Universitas Pamulang, Tangerang Selatan, Indonesia
  • (*) Corresponding Author
Keywords: Retrieval-Augmented Generation; Local Large Language Model; Llama; Mistral; Helpdesk Knowledge Retrieval

Abstract

The increasing volume of technical documentation and repetitive support requests at PT XYZ has made it difficult for helpdesk personnel to retrieve accurate information efficiently and consistently. The 114 internal applications that PT XYZ oversees have different documentation, troubleshooting protocols, and frequently asked questions, which leads to dispersed knowledge sources and drawn-out problem solving. This study uses locally deployed Large Language Models (LLMs), specifically Llama and Mistral, to design and assess a Retrieval-Augmented Generation (RAG) based chatbot for internal helpdesk knowledge retrieval in order to solve this issue. The suggested solution combines generative language models with semantic document retrieval in PostgreSQL using PGVector. The study compares the effectiveness of local LLMs in RAG and non-RAG configurations, employing 1,068 internal question-answer pairs for knowledge retrieval and 214 evaluation questions for performance assessment. ROUGE, BLEU, and Cosine Similarity measures are used for evaluation. According to experimental findings, RAG greatly enhances both models' performance. While Mistral improved from 0.0984 to 0.8983, Llama ROUGE-1 score rose from 0.1710 to 0.7345. With a ROUGE-1 score of 0.8983, a BLEU-1 score of 0.8161, and a Cosine Similarity score of 0.8916, Mistral with RAG performed the best of all setups. These results show that integrating RAG with locally installed LLMs greatly improves contextual correctness and response relevance, enabling safe and scalable implementation for enterprise helpdesk knowledge management.

Downloads

Download data is not yet available.

References

D. A. Susanto, “Perancangan Aplikasi Sistem Informasi Helpdesk pada Kantor ABC,” Jurnal Tera, vol. 2, no. 2, pp. 26–33, Sep. 2022, doi: 10.59832/jt.v2i2.119.

W. M. H. Sasmita, S. Sumpeno, and R. F. Rachmadi, “Improving Government Helpdesk Service with an AI-Powered Chatbot Built on the Rasa Framework,” Jurnal RESTI, vol. 9, no. 2, pp. 393–403, Apr. 2025, doi: 10.29207/resti.v9i2.6293.

R. H. Nugroho, I. R. Kusumasari, V. Febrianto, M. A. Farhan N. H, and M. R. Mahardika, “Strategi Teknologi Artificial Intelligence (AI) dalam Pengambilan Keputusan Bisnis di Era Digital,” Jurnal Bisnis dan Komunikasi Digital, vol. 2, no. 2, p. 7, Dec. 2024, doi: 10.47134/jbkd.v2i2.3476.

G. D. Albert and A. Voutama, “Pengembangan Chatbot Berbasis PDF Menggunakan Local Retrieval-Augmented Generation (RAG) dan Ollama,” Jurnal Informatika dan Teknik Elektro Terapan, vol. 13, no. 2, Apr. 2025, doi: 10.23960/jitet.v13i2.6361.

M. S. Y. Lubis, “Implementasi Artificial Intelligence pada System Manufaktur Terpadu,” in Prosiding Seminar Nasional Teknik Universitas Islam Sumatera Utara (UISU), 2021, pp. 1–7.

M. D. A. Muhajir, N. Prastiti, and M. Koeshardianto, “Implementasi Chatbot Menggunakan Framework Langchain Berbasis LLM GPT (Studi Kasus : Panduan Akademik Universitas Trunojoyo),” Jurnal Mahasiswa Teknik Informatika, vol. 9, no. 2, 2025, doi: 10.36040/jati.v9i2.13003.

Y. Tribber, Kusnadi, and M. Asfi, “Implementasi Retrieval Augmented Generation untuk Layanan Informasi Kampus dengan Chatbot Berbasis AI,” in Prosiding Seminar Nasional Sistem Informasi dan Teknologi (SISFOTEK) ke-8, 2024, pp. 594–600.

Asmaidin and C. B. Santoso, “Evaluasi Metode Retrieval pada Chatbot Domain Khusus Berbasis Retrieval-Augmented Generation,” Journal Scientific and Applied Informatics, vol. 09, no. 1, pp. 105–111, Jan. 2026, doi: 10.36085/jsai.v9i1.9897.

Daerobby, Tukiyat, and A. Musyafa, “Analisis dan Evaluasi Kinerja Chatbot Penerimaan Mahasiswa Baru Berbasis LLM dengan Pendekatan RAG,” Jurnal Innovation and Future Technology, vol. 8, no. 1, pp. 53–61, Feb. 2026, doi: 10.47080/iftech.v8i1.4452.

F. R. Fatonah, D. S. Maylawati, and E. Nurlatifah, “Chatbot Edukasi Pra-Nikah Berbasis Telegram Menggunakan Bidirectional Encoder Representations from Transformers (BERT),” Jurnal Algoritma, vol. 21, no. 2, pp. 29–40, Nov. 2024, doi: 10.33364/algoritma/v.21-2.1657.

T. Helviansyah, N. S. Harahap, M. Irsyad, and B. S. Negara, “Sistem Tanya Jawab Berbasis Chatbot Website Menggunakan Gemini AI pada Data Fiqih Kontemporer,” Journal of Information System Management, vol. 7, no. 1, pp. 38–47, Jun. 2025, doi: 10.24076/joism.2025v7i1.2082.

N. N. Qoniah and D. Ramadhani, “Optimization of Feminacare Chatbot Application Using SeaLLM Model,” Indonesian Journal of Informatic Research and Software Engineering, vol. 5, no. 1, pp. 68–78, Mar. 2025, doi: 10.57152/ijirse.v5i1.2043.

A. G and V. K, “RAG Based Chatbot Using LLMs,” Interantional Journal of Scientific Research in Engineering and Management, vol. 8, no. 6, pp. 1–4, Jun. 2024, doi: 10.55041/IJSREM35600.

A. Bora, “Comparative Analysis of Retrieval-Augmented Generation (RAG) Based Large Language Models (LLM) for Medical Chatbot Applications,” Master’s thesis, University of Lincoln, Lincoln, 2024. doi: 10.13140/RG.2.2.33650.93124.

T. Muhammad, R. Rahardiansyah, R. Setya Perdana, and T. N. Fatyanosa, “Analisis Teknik Embedding Model NV-Embed pada Large Language Models Berbasis Retrieval Augmented Generation,” Jurnal Pengembangan Teknologi Informasi dan Ilmu Komputer, vol. 9, no. 2, pp. 2548–964, Feb. 2025, [Online]. Available: http://j-ptiik.ub.ac.id

M. Klesel and H. F. Wittmann, “Retrieval-Augmented Generation (RAG),” Business and Information Systems Engineering, vol. 67, pp. 551–561, Jun. 2025, doi: 10.1007/s12599-025-00945-3.

K. Fathoni et al., “EMasjid, Islamic Chatbot Application With RAG (Retrieval Augmented Generation),” Informatics, Electrical and Electronics Engineering (Infotron), vol. 5, no. 1, pp. 49–58, May 2025, doi: 10.33474/infotron.v3i2.22894.

Y. Gao et al., “Retrieval-Augmented Generation for Large Language Models: A Survey,” arXiv preprint arXiv:2312.10997, pp. 1–21, Dec. 2023, doi: 10.48550/arXiv.2312.10997.

M. Barbella and G. Tortora, “ROUGE Metric Evaluation for Text Summarization Techniques,” SSRN Electronic Journal, pp. 1–31, May 2022, doi: 10.2139/ssrn.4120317.

F. F. Kartamanah, A. R. Atmadja, and I. Budiman, “Analyzing PEGASUS Model Performance with ROUGE on Indonesian News Summarization,” Jurnal dan Penelitian Teknik Informatika (Sinkron), vol. 9, no. 1, pp. 31–42, Jan. 2025, doi: 10.33395/sinkron.v9i1.14278.

I. Haromain, S. Munir, and A. Rahmah, “Analisa Prompt Engineering pada Large Language Model dengan Retrieval-Augmented Generation untuk Informasi Obat dan Vitamin,” Indonesian Journal Computer Science, vol. 4, no. 2, Oct. 2025, doi: 10.31294/ijcs.v4i2.10005.

M. Yang, “Application of Multimodal Generation Model in Short Video Content Personalized Generation,” Informatica, vol. 49, no. 21, Dec. 2025, doi: 10.31449/inf.v49i21.9838.

J. Ilmiah, T. Harapan, A. Bunga Permata, T. Listyorini, and E. Supriyati, “Pengembangan Mesin Penerjemah Bahasa Indonesia ke Bahasa Daerah Kudus Menggunakan NMT Berbasis RNN-GRU,” Jurnal Ilmiah Teknologi Harapan, vol. 14, no. 1, pp. 15–27, Mar. 2026, doi: 10.35447/jitekh.v14i1.1314.

A. Sanjaya, A. Bagus Setiawan, U. Mahdiyah, I. Nur Farida, A. Risky Prasetyo, and U. Nusantara PGRI Kediri, “Measurement Of Meaning Similarity Using Cosine Similarity and Word Synonyms Database,” vol. 10, no. 4, 2023, doi: 10.25126/jtiik.2023106864.

M. Qamal, M. Iqbal, and Y. Afrillia, “Utilizing Text Mining For Encryption Algorithm Recommendation Using Content-Based Filtering<b></b>,” JITK (Jurnal Ilmu Pengetahuan dan Teknologi Komputer), vol. 11, no. 4, pp. 1272–1280, May 2026, doi: 10.33480/jitk.v11i4.7663.


Bila bermanfaat silahkan share artikel ini

Berikan Komentar Anda terhadap artikel Comparative Evaluation of Local Llama and Mistral Models in RAG for Helpdesk Documentation

Dimensions Badge
Article History
Submitted: 2025-12-09
Published: 2026-06-30
Abstract View: 0 times
PDF Download: 0 times
How to Cite
Pratama, S., Anggai, S., & Handayani, M. (2026). Comparative Evaluation of Local Llama and Mistral Models in RAG for Helpdesk Documentation. Building of Informatics, Technology and Science (BITS), 8(1), 450-458. https://doi.org/10.47065/bits.v8i1.8895
Issue
Section
Articles