https://ejurnal.seminar-id.com/index.php/bulletinds/issue/feedBulletin of Data Science2026-07-13T20:42:51+07:00Support Journalseminar.id2020@gmail.comOpen Journal Systems<p data-start="0" data-end="433">The <strong><em data-start="4" data-end="30">Bulletin of Data Science</em></strong> is a journal managed and published by the <a href="https://fkpt.org/"><strong><span class="hover:entity-accent entity-underline inline cursor-pointer align-baseline"><span class="whitespace-normal">Forum Kerjasama Pendidikan Tinggi</span></span> (FKPT)</strong></a>, located in <span class="hover:entity-accent entity-underline inline cursor-pointer align-baseline"><span class="whitespace-normal">Medan</span></span>, <span class="hover:entity-accent entity-underline inline cursor-pointer align-baseline"><span class="whitespace-normal">North Sumatra</span></span>. The journal publishes research in the field of <strong>Informatics</strong>, particularly those related to <strong>Data Science</strong>. It holds the online ISSN <a href="https://issn.perpusnas.go.id/terbit/detail/20210917311475138">2807-9493 (online)</a>, based on Decree No. 0005.28079493/K.4/SK.ISSN/2021.09 issued on September 20, 2021.<br>The <strong><em data-start="439" data-end="465">Bulletin of Data Science</em></strong> is published three times a year, namely in October (<strong>Issue 1</strong>), February (<strong>Issue 2</strong>), and June (<strong>Issue 3</strong>). It is indexed in <strong> <a href="https://scholar.google.com/citations?hl=id&user=c9HhReoAAAAJ">Google Scholar</a> |</strong><strong> <a href="https://garuda.kemdikbud.go.id/journal/view/24551">Portal Garuda </a>| <a href="https://portal.issn.org/resource/ISSN/2807-9493#">ROAD</a> | <a href="https://app.dimensions.ai/discover/publication?and_facet_source_title=jour.1492637">Dimensions</a> | Science and Technology Index (SINTA 4) </strong></p> <p> </p>https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/9871Deteksi Kelelahan Wajah Mahasiswa Akibat Begadang Berbasis Convolutional Neural Network (CNN)2026-06-18T02:59:34+07:00Doni Ekrian Saputradoniekriansaputra25@gmail.comYovi Apridiansyahyoviapridiansyah@umb.ac.idAgung Kharisma Hidayahkharisma@umb.ac.idArdi Wijayaardiwijaya@umb.ac.id<p>This research is entitled Detecting Physical Facial Changes Due to Staying Up Late in College Students Using Convolutional Neural Network (CNN). Staying up late or lack of sleep is a common habit among students due to academic pressure, social activities, and prolonged use of digital devices. This condition can cause changes in facial physical characteristics such as red eyes, dark circles under the eyes, and a dull appearance. Identification of these changes is generally still done subjectively, so a system is needed that can detect them automatically by utilizing image processing technology and artificial intelligence. This research aims to design and implement a Convolutional Neural Network (CNN) model to detect changes in facial physical characteristics due to staying up late in students. The dataset used is 330 facial images obtained by taking photos using a cellphone camera in JPEG format. The dataset consists of four categories: normal faces, dull faces, red eyes, and dark circles under the eyes. The data then goes through a pre-processing stage, dividing the dataset into 300 training data and 30 test data, and the model training process using the MobileNetV2 architecture with the Python programming language. Test results show that the CNN model is capable of classifying changes in facial physical characteristics with good performance. Based on an evaluation using a Confusion Matrix, the model achieved a precision of 89%, a recall of 87%, an F1-score of 86%, and an accuracy of 87%. The CNN method is considered effective in automatically detecting changes in facial physical characteristics due to staying up late in students.</p>2026-06-17T00:00:00+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/9930Sistem Pakar Diagnosis Penyakit Lambung Menggunakan Naïve Bayes dengan Prior Klinis untuk Klasifikasi Gastritis, GERD, dan Dispepsia2026-06-30T20:29:05+07:00Fatma Sarifatmasari2482@gmail.comAgus Sidiq Purnomosidiq@mercubuana-yogya.ac.id<p>Penyakit lambung seperti gastritis, gastroesophageal reflux disease (GERD), dan dispepsia merupakan gangguan yang sering dialami masyarakat dengan karakteristik gejala yang saling mirip, sehingga menyulitkan proses diagnosis awal secara tepat. Penelitian ini bertujuan untuk mengembangkan sistem pakar yang dapat membantu mengidentifikasi jenis penyakit lambung berdasarkan gejala yang dialami pengguna. Sistem dibangun dengan memanfaatkan metode Naïve Bayes yang disesuaikan melalui pemberian penekanan pada tingkat kepentingan gejala, sehingga setiap gejala tidak dianggap memiliki pengaruh yang sama dalam proses penentuan hasil. Basis pengetahuan disusun dari 25 gejala yang diperoleh melalui kajian literatur dan validasi pakar, serta didukung oleh data rekam medis untuk membentuk dasar perhitungan. Pengujian dilakukan menggunakan 30 data uji yang telah diverifikasi oleh pakar. Hasil pengujian menunjukkan bahwa sistem mampu mengklasifikasikan 27 data dengan benar, sehingga diperoleh tingkat akurasi sebesar 90%. Ketidaksesuaian hasil terjadi pada beberapa kasus dengan kombinasi gejala yang sangat mirip antar penyakit. Secara umum, sistem mampu memberikan hasil diagnosis yang cukup mendekati dengan pertimbangan pakar serta menyajikan tingkat keyakinan untuk setiap kemungkinan penyakit. Dengan demikian, sistem ini dapat membantu mengenali kemungkinan penyakit lambung sejak awal berdasarkan gejala yang dialami.</p>2026-06-17T16:57:06+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/10071Optimasi Hiperparameter Gradient Boosting Regressor Untuk Prediksi Curah Hujan Berdasarkan Data Cuaca Harian BMKG2026-06-30T20:25:47+07:00Mohamad Arif Abdul Syukurmarifabdulsyukur@gmail.comSuhartono Suhartonosuhartono@ti.uin-malang.ac.idM. Imamudinimamudin@ti.uin-malang.ac.id<p><strong>Abstrak</strong><strong>−</strong>Prediksi curah hujan harian memiliki peran krusial dalam mitigasi bencana hidrometeorologi dan perencanaan sektor pertanian. Namun, karakteristik data curah hujan yang bersifat <em>zero-inflated</em> dan non-linear menjadi tantangan utama dalam menghasilkan prediksi yang akurat. Penelitian ini bertujuan untuk membangun model prediksi curah hujan menggunakan algoritma <em>Gradient Boosting Regressor</em> dengan membandingkan dua skenario: model <em>default</em> dan model yang telah dioptimasi. Dataset yang digunakan adalah data cuaca harian dari Badan Meteorologi, Klimatologi, dan Geofisika (BMKG) sebanyak 942 observasi, yang mencakup variabel suhu minimum (TN), suhu maksimum (TX), suhu rata-rata (TAVG), kelembaban rata-rata (RH_AVG), durasi penyinaran matahari (SS), serta kecepatan angin rata-rata (FF_AVG). Proses pra-pemrosesan data dilakukan melalui penanganan nilai kosong menggunakan <em>K-Nearest Neighbors Imputer</em> dan standarisasi fitur dengan <em>StandardScaler</em>. Optimasi <em>hyperparameter</em> dilakukan menggunakan <em>RandomizedSearchCV</em> dengan validasi silang deret waktu (<em>TimeSeriesSplit</em>). Hasil penelitian menunjukkan bahwa model teroptimasi mengungguli model <em>default</em> dengan mencapai nilai <em>Root Mean Squared Error</em> (RMSE) sebesar 12,207 dan <em>Mean Absolute Error</em> (MAE) sebesar 7,704. Temuan ini membuktikan bahwa integrasi pra-pemrosesan yang sistematis dan optimasi parameter secara empiris secara signifikan meningkatkan kemampuan model dalam mempelajari pola variabilitas curah hujan harian.</p>2026-06-18T00:00:00+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/10079Klasifikasi Serangan Uniform Resource Locator Phishing Menerapkan Metode Backpropagation Neural Network2026-06-30T20:19:06+07:00Novi Yantinovi_yanti@uin-suska.ac.idM. Rival Kurniawan12050116070@students.uin-suska.ac.idRahmad Abdillahrahmad.abdillah@uin-suska.ac.idBenny Sukma Negarabsnegara@uin-suska.ac.id<p><em>Phishing</em> merupakan bentuk kejahatan siber yang bertujuan mencuri informasi sensitif pengguna melalui metode penipuan, seperti manipulasi tautan link <em>Uniform Resource Locator</em> (URL) pada situs web. Melihat serangan siber yang makin pesat, deteksi dini terhadap URL berbahaya menjadi sangat krusial. Namun, model deteksi terdahulu masih memiliki tantangan karena batasan jumlah dataset dan efisiensi kinerja komputasi. Pada penelitian ini dilakukan klasifikasi serangan URL <em>phishing</em> menerapkan metode <em>Backpropagation Neural Network</em> (BPNN) dengan integrasi <em>Knowledge Discovery in Database</em> (KDD). Data yang digunakan bersumber dari dataset publik Kaggle (PhiUSIIL-2024) berjumlah 235.795 data URL dengan 56 fitur. Tahapan KDD dimulai dari seleksi data dengan mereduksi 5 atribut tidak relevan, hasil pembersihan data menjadi 234.611 data valid, dan transformasi menggunakan <em>Min-Max Scaler</em>. Hasil penelitian menunjukkan bahwa model BPNN memiliki kinerja paling optimal menggunakan arsitektur 50 neuron <em>input</em>, 75 neuron <em>hidden</em>, dan 1 neuron <em>output</em> (50-75-1). Hasil pengujian dengan skenario pembagian rasio data latih dan uji 90:10, <em>learning rate</em> 0.1, nilai 100 <em>epoch</em>, model memberikan nilai akurasi sebesar 99.982%, presisi 99.970%, <em>recall</em> 100%, dan <em>F1-Score</em> 99.985%. Pendekatan ini menghasilkan sistem deteksi URL <em>phishing</em> yang sangat adaptif, stabil, dan berakurasi tinggi dalam menghadapi variasi serangan siber modern.</p>2026-06-19T17:05:16+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/10061Implementation of the Certainty Factor Method in Determining the Type of Treatment for Dry Facial Skin2026-06-30T19:58:30+07:00Mesran Mesranmesran.skom.mkom@gmail.comRima Marhamahrima.marhamah144@gmail.comMahendra Maulia Lubismahendralubiss@gmail.com<p>Facial skin is one of the most sensitive parts of the body and is vulnerable to various disorders, one of which is dry skin. Dry facial skin is characterized by a dull appearance, a tendency to experience irritation, and the emergence of fine lines or wrinkles. This condition not only affects physical appearance but can also reduce the skin’s protective function against infections and exposure to free radicals. This study aims to identify dry facial skin conditions based on symptoms experienced by users by applying the Certainty Factor (CF) method. The CF method is selected because it is capable of handling uncertainty in decision-making processes, particularly in Expert System applications. In its implementation, the system calculates the certainty value from a combination of symptoms entered by the user to determine the probability level of dry facial skin conditions. The results of the study indicate that the system produces a certainty value of 22.32%, suggesting that the user has a 22.32% probability of experiencing dry facial skin. These findings demonstrate that a CF-based expert system can be utilized as an initial supporting tool for identifying skin conditions, although further validation by experts in the field of Dermatology is still required</p>2026-06-19T00:00:00+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/9960Sistem Pendukung Keputusan Dalam Pemilihan Mahasiswa Berprestasi Menerapkan Metode MOORA dan MOOSRA menggunakan Pembobotan Rank Order Centroid (ROC)2026-06-30T19:38:38+07:00Dwi Setyasihdwi_set@staff.gunadarma.ac.idOctarina Budi Lestariocta_bl@staff.gunadarma.ac.idSyarifah Azharina Syafrudins_azharina@staff.gunadarma.ac.idIndah Wahyuniiwahyuni@staff.gunadarma.ac.id<p>Each university must conduct student selection so that the tertiary institution can propose students who excel in participating in the competition activities carried out by the Higher Education. A student is someone who is undergoing a college level to achieve the desired goals and a great responsibility to be a student at each university. Higher education has the meaning of a place that is used by students and lecturers as a place for students to study in gaining knowledge given by lecturers and teaching students in providing knowledge. In the event, students who excel are used by students so that they can be dubbed students. A student is a student who has expertise that exceeds a certain field and has achieved achievements. Students who excel must receive an award from the university for their achievements. In giving awards to students, there must be a logistical selection process and the results must be real. So that students can excel, there are several criteria that have been made by the university such as achievement index, scientific writing, organizational activity, mastery of English. From several criteria that have been made, students who want to take part in the selection must meet the criteria above, which are the criteria as a measurement in the selection of outstanding students. A decision support system (Decision Support System) has the meaning of a system that can perform data processing using computer assistance. This system aims to assist in overcoming problems in processing data on the selection of outstanding students which will produce real data for students in applying the MOORA and MOOSRA methods. The MOORA and MOOSRA methods aim to assist the support system in making decisions from the data above to obtain absolute and accurate results. The results of the selection of outstanding students using the MOORA and MOOSRA methods were compared, the results of the MOORA method which was ranked 1 on the A7 alternative named Azmi had a value of 8.1103 while the MOOSRA method was A4 worth 39.1545 on behalf of Waldi.</p>2026-06-19T00:00:00+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/10147Klasifikasi Multi-Label Hadits Shahih Muslim Menggunakan Metode Support Vector Machine (SVM)2026-06-30T19:33:36+07:00Dinda Lutfiah12250125077@students.uin-suska.ac.idNazruddin Safaat Harahapnazruddin.safaat@uin-suska.ac.idFebi Yantofebiyanto@uin-suska.ac.idEka Pandu Cynthiaeka.pandu.cynthia@uin-suska.ac.id<p><strong>Abstract</strong><strong>−</strong>In Islamic studies, hadith is a form of text used as a reference in understanding religious teachings and practices. The problem of this research is the complexity of hadith classification because a single text can contain several categories of meaning at once, such as recommendations, prohibitions, and information, so that a single-label approach is not sufficient to represent the contents of the hadith as a whole. This study aims to apply Support Vector Machine (SVM) for multi-label classification of the Indonesian translation of Sahih Muslim Hadith. The dataset consists of 5,362 hadith texts, namely 4,557 training data and 805 test data. The research stages include text preprocessing, TF-IDF feature extraction, One-Vs-Rest classification, and evaluation using precision, recall, F1-score, hamming loss, and subset accuracy. The baseline SVM obtained micro precision of 0.8990, micro recall of 0.8080, micro F1-score of 0.8511, macro precision of 0.6404, macro recall of 0.4446, macro F1-score of 0.4798, hamming loss of 0.1226, and subset accuracy of 0.6696. After hyperparameter tuning using GridSearchCV, the model obtained micro precision of 0.7919, micro recall of 0.8395, micro F1-score of 0.8150, macro precision of 0.5482, macro recall of 0.6341, macro F1-score of 0.5809, hamming loss of 0.1652, and subset accuracy of 0.5764. The results showed that tuning increased macro recall and macro F1-score, but decreased micro F1-score, hamming loss, and subset accuracy. The main contribution lies in the performance evaluation of label imbalance using macro and micro metrics as well as the analysis of dominant features in each category of hadith meaning. Thus, SVM can be used for multi-label classification of Sahih Muslim Hadith, although label imbalance still affects the model performance.</p>2026-06-20T00:09:23+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/10142Klasifikasi Multi-Label Hadis Terjemahan Shahih Bukhari Berdasarkan Kategori Anjuran, Larangan, dan Informasi Menggunakan Metode Extreme Gradient Boosting (XGBoost)2026-06-30T19:21:33+07:00Tasya Dwi Yanti12250123081@students.uin-suska.ac.idFitra Kurniafitra.k@uin-suska.ac.idNazruddin Safaat Harahapnazruddin.safaat@uin-suska.ac.idIwan Iskandariwan.iskandar@uin-suska.ac.idReski Mai Candrareski.candra@uin-suska.ac.id<p>Hadith is one of the sources of Islamic teachings that contains meanings such as recommendations, prohibitions, and<br>information. The large number of hadiths makes manual classifications time consuming, therefore an automatic classification method<br>is needed. This study aims to perform multi-label classifications on translated Shahih Bukhari hadiths using the XGBoost method with<br>the Binary Relevence approach and hyperparameter tuning using GridSearchCV. the result showed that the implementation of<br>hyperparameter tuning improved model performance, especially in recall and F1-score values. In addition, the hamming loss value<br>decraesed from 9.19% to 9.10%, indicating that hyperparameter tuning was able to reduce prediction errors in the model.</p>2026-06-21T00:39:56+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/9966Sistem Pendukung Keputusan Pemilihan Kaprodi Terbaik Dengan Penerapan Metode MOORA Menggunakan Pembobotan ROC2026-06-21T00:52:42+07:00Erma Sovaerma_sova@staff.gunadarma.ac.idMakmun Makmunmakmun@staff.gunadarma.ac.idDini Dwi Ermawatidinidwi@staff.gunadarma.ac.id<p>Di dalam dunia sekolah tinggi ataupun universitas tentunya banyak mahasiswa-mahasiswi yang telah mengetahui istilah kata kaprodi atau yang lebih dikenal ketua program studi.Kaprodi sebagai elemen dalam manajemen sekolah tinggi atau universitas untuk merealisasikan visi, misi dan tujuan program studi yang relevan dengan visi, misi dan tujuan lembaga secara keseluruhan. Dalam pemilihan kaprodi terbaik biasanya pihak sekolah tinggi maupun universitas hanya melihat pada kriteria Umpan Balik Mahasiswa saja, padahal ada beberapa kriteria yang dapat dijadikan penilaian dalam pemilihan kaprodi terbaik seperti, Pendidikan, Jumlah Kegiatan PKM, Jumlah Score Sinta, Lama Mengajar dan Umpan Balik Mahasiswa.Pada pemilihan kaprodi terbaik dibutuhkan sebuah sistem pendukung keputusan. Metode yang digunakan dalam penelitian ini ialah Metode MOORA. Metode ini dipilih karena mampu menghasilkan alternatif dari sejumlah alternatif. Adapun yang menjadi hasil perhitungan MOORA alternatif terbaik dari penelitian ini yaitu altenatif altenatif <strong>A11</strong> atas nama <strong>“Sulaiman, M.Kom “</strong> dengan nilai <strong>Yi = 0,4736</strong>, sebagai kaprodi terbaik.</p>2026-06-21T00:52:42+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/9781Analisis Komparatif Algoritma Ridge, LASSO, MLP, dan Gaussian Process Regressor pada Prediksi Efisiensi Kendaraan Listrik2026-06-21T23:54:53+07:00Annida Purnamawatiannida.npr@bsi.ac.idMonikka Nur Winnartomonikka.mnt@bsi.ac.idMely Mailasarimely.mly@bsi.ac.id<p>Penelitian ini bertujuan untuk melakukan analisis komparatif terhadap beberapa algoritma regresi dalam memprediksi efisiensi kendaraan listrik. efisiensi kendaraan listrik menjadi permasalahan penting dalam mendukung pengembangan transportasi berkelanjutan. Penelitian ini dilakukan untuk membandingkan kinerja beberapa algoritma regresi dalam menghasilkan prediksi dengan tingkat kesalahan rendah dan akurasi tinggi. Metode yang digunakan adalah Cross-Industry Standard Process for Data Mining (CRISP-DM) yang meliputi tahapan pemahaman bisnis, pemahaman data, persiapan data, pemodelan, evaluasi, dan penerapan. Data yang digunakan berupa dataset kendaraan listrik yang memuat atribut terkait efisiensi energi. Model yang dianalisis terdiri dari Ridge, LASSO, MLP, dan Gaussian Process Regressor. Kebaruan penelitian ini terletak pada analisis komparatif yang terstruktur menggunakan pendekatan CRISP-DM serta evaluasi kinerja model regresi pada konteks prediksi efisiensi kendaraan listrik yang masih terbatas dikaji. Hasil pengujian menunjukkan bahwa model LASSO memberikan kinerja terbaik dengan nilai Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), dan R-Squared (R²). Hasil penelitian menunjukkan bahwa algoritma LASSO memiliki performa terbaik dengan nilai MAE sebesar 2.13, RMSE sebesar 16.10, dan R² sebesar 0.974. Sementara itu, MLP menunjukkan performa terendah dibandingkan model lainnya. Dengan demikian, pemilihan algoritma yang tepat terbukti berpengaruh signifikan terhadap akurasi prediksi.</p>2026-06-21T00:00:00+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/10145Klasifikasi Ujaran Kebencian Menggunakan 5-Fold Ensemble Weighted Probability Averaging berbasis Arsitektur Twitter-RoBERTa2026-06-30T19:14:44+07:00Ibra Sahrian Alsa12150114195@students.uin-suska.ac.idSurya Agustiansurya.agustian@uin-suska.ac.idFitra Kurniafitra.k@uin-suska.ac.idPizaini Pizainipizaini@uin-suska.ac.idSiska Kurnia Gustisiskakurniagusti@uin-suska.ac.id<p class="western" align="justify"><span style="font-size: small;"><em>Social media platforms have become a critical medium for hate speech propagation at unprecedented scale, with over 66.8 million user reports regarding hateful conduct recorded on platform X during the first half of 2024 alone. This study proposes an end-to-end NLP pipeline for automated hate speech classification using the domain-adapted Twitter-RoBERTa architecture, evaluated on the HASOC (Hate Speech and Offensive Content Identification) English datasets from 2020 and 2021. The core challenge addressed is Transformer fine-tuning instability on relatively small annotated corpora caused by extreme sensitivity to random seed initialization and suboptimal hyperparameter configurations. Three methodological innovations are synergistically integrated: (1) Bayesian Optimization via the Optuna framework for automated adaptive hyperparameter search with 15 trials; (2) Stratified 5-Fold Cross-Validation for robust, reproducible data partitioning; and (3) Weighted Probability Averaging (WPA) as the ensemble aggregation strategy. Results demonstrate that the proposed architecture achieves a Macro F1-Score of 80.99% on Subtask 1A and 64.70% on Subtask 1B, positioning it competitively against 65 international research teams on the official HASOC 2021 leaderboard.</em></span></p>2026-06-22T00:00:00+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/10262Prediksi Harga Saham BBCA Menggunakan Metode CNN, LSTM dan CNN-LSTM Berbasis Variasi Timestep dan Rasio Split pada Data OHLCV Time Series2026-06-30T19:08:34+07:00Rifki Faiz Azzurananda12250113451@students.uin-suska.ac.idSiska Kurnia Gustisiskakurniagusti@uin-suska.ac.idElin Haeranielin.haerani@uin-suska.ac.idEka Pandu Cynthiaeka.pandu.cynthia@uin-suska.ac.id<p>Predicting stock prices is a real challenge because financial data moves unpredictably and its patterns are hard to figure out. This study implements and compares three deep learning models, namely Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM), and hybrid CNN-LSTM, to predict the stock price of PT Bank Central Asia Tbk (BBCA) using OHLCV (Open, High, Low, Close, Volume) data from 2015 to 2025. The focus of this study is to compare the performance of each model under different experimental settings, rather than to propose a new architecture. Experiments were carried out using timesteps of 30, 60, and 90, combined with data split ratios of 70:30, 80:20, and 90:10. Model performance was measured using Root Mean Square Error (RMSE), Mean Absolute Percentage Error (MAPE), and the coefficient of determination (R²). The results show that CNN gave the most stable performance across all tested scenarios, while LSTM was able to pick up temporal patterns but still showed lagging in some conditions. The hybrid CNN-LSTM model came out on top at timestep 90 with an 80:20 data split, reaching an RMSE of 130.5432, MAPE of 1.0669%, and R² of 0.9581. This shows that the CNN-LSTM hybrid works better at producing accurate and responsive predictions compared to single models.</p>2026-06-23T15:54:30+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/10129Analisis Sentimen Komentar YouTube Terkait Indonesia Tidak Lolos Kualifikasi Piala Dunia Menggunakan Support Vector Machine2026-06-30T19:03:46+07:00Idris Syahrudinidrissyahrudin15552@gmail.comSufajar Butsiantosufajar@pelitabangsa.ac.idAndini Putri Riandaniandiniriandani@pelitabangsa.ac.id<p>Penelitian ini bertujuan untuk menganalisis sentimen opini publik pada media sosial YouTube terkait Tim Nasional Indonesia tidak lolos dalam kualifikasi Piala Dunia. Fokus utama penelitian ini adalah mengklasifikasikan komentar pengguna ke dalam tiga kategori, yaitu positif, negatif, dan netral, untuk memahami pola reaksi masyarakat terhadap performa tim nasional yang di mana komentar-komentar yang muncul dapat dianalisis untuk memahami kecenderungan sentimen dan pola interaksi pengguna. Metodologi yang digunakan mengikuti kerangka Knowledge Discovery in Databases. Data sebanyak 15.169 komentar dikumpulkan melalui proses crawling menggunakan YouTube Data API v3. Data tersebut kemudian melalui tahap pra-pemrosesan yang intensif, meliputi case folding, cleaning, tokenisasi, normalisasi kata, dan stopword removal, sehingga menghasilkan 7.584 data yang siap untuk diklasifikasikan. Algoritma Support Vector Machine dengan kernel linear diterapkan pada dataset yang dibagi dengan rasio 80:20 antara data latih dan data uji. Hasil penelitian menunjukkan bahwa mayoritas opini masyarakat didominasi oleh sentimen negatif (44,03%) dibandingkan sentimen netral (31,34%) dan positif (24,62%). Pengujian model SVM menghasilkan tingkat akurasi sebesar 80,43% dengan nilai weighted F1-score sebesar 0,81. Hal ini menunjukkan bahwa metode SVM sangat efektif dalam menangani data teks yang tidak terstruktur dan noisy komentar media sosial YouTube. Penelitian ini membuktikan keandalan algoritma SVM dalam analisis sentimen bersekala besar.</p>2026-06-23T17:55:17+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/10179Evaluasi Seleksi Fitur Algoritma Genetika pada Klasifikasi Penyakit Stroke Menggunakan K-Nearest Neighbor pada Dataset Overlapping2026-06-30T19:00:41+07:00Aqmal Syarif Fadilah12250113564@students.uin-suska.ac.idSiska Kurnia Gustisiskakurniagusti@uin-suska.ac.idIwan Iskandariwan.iskandar@uin-suska.ac.idIis Afriantyiis.afrianty@uin-suska.ac.id<p>Stroke is a disease that disrupts brain function due to obstructed blood flow and requires early detection to reduce the risk of disability and death. One of the classification methods widely used in stroke prediction is K-Nearest Neighbor (KNN), but this algorithm is sensitive to irrelevant features and complex data distribution. This study aims to evaluate the effectiveness of feature selection using Genetic Algorithm (GA) on the performance of KNN classification on stroke prediction datasets characterized by high overlap between classes. The dataset used consists of 15,000 patient data with 21 initial attributes, which after the preprocessing stage changed to 38 attributes. The evaluation model was carried out using 10-Fold Stratified Cross Validation at three data sharing ratios, namely 90:10, 80:20, and 70:30, with parameter K = 13. The results showed that GA was effective in reducing the data dimensionality significantly to 15–17 selected features and increasing the average internal validation fitness value by around 3%. However, evaluation of the test data shows that the performance of both the KNN and GA-KNN models remains stagnant at an accuracy range of 48–52%, with ROC-AUC values approaching 0.500. This condition indicates that GA tends to find local optimal solutions that are sensitive to the division of the training data, making it difficult for the model to generalize when processing new data. The extreme level of overlap in this dataset is further demonstrated through comparative experiments using feature engineering and the Random Forest algorithm, which both stalled at around 50% accuracy. These results indicate that GA feature selection successfully addresses data complexity, but is not automatically capable of solving class criteria problems if overlap is an inherent characteristic of the data itself.</p>2026-06-23T22:52:11+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/10306Analisis Kinerja Random Forest dan XGBoost serta Ensemble Soft Voting untuk Klasifikasi Stroke pada Dataset yang Tidak Seimbang2026-06-30T18:57:27+07:00Bagus Fadhil Husain12250115278@students.uin-suska.ac.idSiska Kurnia Gustisiskakurniagusti@uin-suska.ac.idMuhammad Fikrymuhammad.fikry@uin-suska.ac.idReski Mai Candrareski.candra@uin-suska.ac.id<p>Stroke is one of the leading causes of disability and death worldwide, highlighting the need for predictive methods that can support early stroke risk detection. This study aims to implement an Ensemble Soft Voting method based on the combination of Random Forest and XGBoost algorithms and compare its performance with individual models for stroke classification using the Framingham Heart Study dataset, which exhibits a high degree of class imbalance. The initial dataset consisted of 11.627 records with 39 attributes and underwent several preprocessing stages, including attribute selection based on medical literature, data deduplication, handling of missing value using MICE, outlier treatment, feature selection using ANOVA F-Test, and data balancing using SMOTE-ENN. Model validation was performed using Stratified K-Fold Cross Validation with k values ranging from 3 to 20, while hyperparameter optimization was conducted using GridSearchCV. The results showed that Random Forest achieved the best performance in terms of F1-Score and Recall, with values of 36.52% and 67.74%, respectively, using a 90:10 train-test split and k=5. Meanwhile, XGBoost achieved the highest Accuracy and Precision of 82.98% and 25.71%, respectively, whereas Ensemble Soft Voting obtained the highest ROC-AUC value of 81.29%. PCA and t-SNE visualizations revealed a considerable degree of overlap between the Stroke and No Stroke classes, making the classification task more challenging. The findings indicate that Random Forest provides the best balance in identifying stroke cases within the dataset, although data characteristics remain a major factor influencing classification performance.</p>2026-06-27T22:44:08+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/10323Analisis Sentimen Multikelas Program Makan Bergizi Gratis di Media Sosial X Menggunakan VaderSentiment dan SVM Berbasis Oversampling SMOTE2026-06-30T18:53:39+07:00Mohamad Hegar Sukmana Wibowohegarsw@gmail.comMuhammad Fatchanfatchan@pelitabangsa.ac.idAsep Supriyantoasep.supriyanto@pelitabangsa.ac.id<p>Media sosial X (Twitter) menjadi ruang diskusi publik yang menghasilkan data opini masyarakat secara real-time terhadap berbagai kebijakan pemerintah, termasuk Program Makan Bergizi Gratis (MBG). Analisis sentimen terhadap kebijakan publik umumnya masih terbatas pada klasifikasi biner serta belum optimal dalam menangani ketidakseimbangan distribusi data. Penelitian ini bertujuan membangun model analisis sentimen multikelas terhadap opini masyarakat mengenai Program MBG dengan mengintegrasikan VaderSentiment sebagai metode ekstraksi fitur, Support Vector Machine (SVM) sebagai algoritma klasifikasi, dan Synthetic Minority Oversampling Technique (SMOTE) untuk mengatasi imbalanced dataset. Penelitian menggunakan pendekatan kuantitatif dengan metode eksperimental. Data berupa tweet berbahasa Indonesia yang relevan dengan MBG dikumpulkan melalui teknik crawling. Tahapan pengolahan data meliputi preprocessing seperti cleaning, case folding, tokenizing, dan stopword removal. Skor sentimen yang dihasilkan oleh VaderSentiment digunakan sebagai fitur numerik dalam proses klasifikasi SVM. Selanjutnya, teknik SMOTE diterapkan pada data latih untuk menyeimbangkan distribusi kelas sentimen positif, netral, dan negatif. Evaluasi model dilakukan menggunakan confusion matrix serta metrik akurasi, precision, recall, dan F1-score. Penelitian ini diharapkan menunjukkan bahwa kombinasi VaderSentiment, SVM, dan SMOTE mampu meningkatkan kinerja klasifikasi sentimen multikelas dibandingkan kondisi sebelum penyeimbangan data. Secara metodologis, penelitian ini berkontribusi pada pengembangan analisis sentimen berbasis machine learning, serta secara praktis mendukung evaluasi kebijakan publik berbasis data melalui pemanfaatan opini masyarakat di media sosial.</p>2026-06-27T00:00:00+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/10246Pengelompokkan Tingkat Kesejahteraan dan Tekanan Psikologis Mahasiswa Menggunakan Mental Health Inventory-38 dengan Algoritma K-Medoids2026-06-30T18:47:50+07:00Muhammad Abdan Syakura12250115210@students.uin-suska.ac.idElvia Budianitaelvia.budianita@uin-suska.ac.idOkfalisa Oktafalisaokfalisa@uin-suska.ac.idFitri Insanifitri.insani@uin-suska.ac.idYuli Widiningsihyuli.widiningsih@uin-suska.ac.id<p>Mental health among university students is an increasingly critical issue, particularly in academic environments characterized by high levels of pressure and demands. This study aims to cluster mental health patterns of students at the Faculty of Science and Technology, Universitas Islam Negeri Sultan Syarif Kasim Riau, from the 2022–2025 cohort, using the <em>K-Medoids</em> algorithm based on the Mental Health Inventory-38 (MHI-38) instrument, which has been modified and validated. Data were collected through an online questionnaire distributed to students, yielding 522 valid records from 559 total respondents following the data cleaning process. Each response was transformed into numerical values based on Favorable and Unfavorable categories using a five-point Likert scale. The <em>K-Medoids</em> algorithm was applied to form clusters, with the number of clusters tested ranging from K=2 to K=10. Clustering quality was evaluated using two metrics, namely the Silhouette Coefficient and the Davies-Bouldin Index (DBI). The results indicate that the optimal number of clusters is K=2, with a Silhouette Coefficient value of 0.2308 and a DBI value of 1.9926. Cluster 0 comprises 280 students categorized as having psychological well-being, while Cluster 1 comprises 242 students categorized as experiencing psychological distress. The findings of this study are expected to serve as a reference for the university in designing more targeted mental health support programs for students.</p>2026-06-27T23:20:43+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/10220Evaluasi Kinerja Arsitektur Ringan YOLOv11n dalam Deteksi Kecacatan Fisik Cangkang Telur untuk Pendukung Sistem Penjaminan Mutu2026-06-30T18:39:58+07:00Caesar Gian Indrarizkycaesargian@student.telkomuniversity.ac.idFigo Mandalafigomandala@student.telkomuniversity.ac.idAisy Hafidzah Fadlillahaisyfadlillah@student.telkomuniversity.ac.id<p>Kecacatan fisik pada cangkang telur dapat menurunkan kualitas produk dan meningkatkan risiko kontaminasi selama proses distribusi maupun konsumsi. Proses inspeksi yang masih dilakukan secara manual memiliki keterbatasan dalam hal konsistensi dan efisiensi, terutama pada skala produksi yang besar. Penelitian ini bertujuan untuk mengem-bangkan model deteksi kecacatan cangkang telur menggunakan arsitektur YOLOv11 guna mendukung proses <em>quality assurance </em>secara otomatis. Dataset yang digunakan terdiri atas 968 citra telur dengan dua kelas, yaitu <em>Damaged </em>dan <em>Normal</em>, serta 1.294 anotasi objek yang telah dibagi ke dalam data pelatihan, validasi, dan pengujian. Model dilatih menggunakan skema <em>5-Fold Cross Validation </em>dan dievaluasi menggunakan metrik <em>precision</em>, <em>recall</em>, <em>F1-Score</em>, <em>mAP50</em>, dan <em>mAP50-95</em>. Hasil penelitian menunjukkan bahwa model memperoleh rata-rata <em>precision </em>sebesar 0,95, <em>recall </em>sebesar 0,93, <em>F1-Score </em>sebesar 0,94, <em>mAP50 </em>sebesar 0,96, dan <em>mAP50-95 </em>sebesar 0,92. Selain itu, model mampu mendeteksi objek telur normal maupun telur cacat dengan performa yang tinggi pada berbagai kondisi citra. Hasil tersebut menunjukkan bahwa YOLOv11 memiliki potensi yang baik untuk diterapkan pada sistem inspeksi kualitas telur secara otomatis.</p>2026-06-27T00:00:00+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/10251Pemanfaatan Nlp Untuk Ekstrasi Profil Kompetensi Dari Portofolio Digital Sebagai Dasar Sistem Rekomendasi2026-06-30T18:42:47+07:00Moch Fahmimuhammadimhaf22@gmail.comDanang Arbian Sulistyodanangarbian@gmail.com<p>Seleksi penerima beasiswa merupakan proses penting dalam mendukung pemerataan akses pendidikan bagi mahasiswa yang memiliki potensi akademik maupun nonakademik. Namun, proses seleksi yang masih dilakukan secara konvensional sering menghadapi kendala berupa banyaknya dokumen yang harus dianalisis serta keterbatasan dalam mengidentifikasi profil kompetensi mahasiswa secara komprehensif. Penelitian ini mengusulkan pemanfaatan <em>Natural Language Processing</em> (NLP) untuk mengekstraksi profil kompetensi dari portofolio digital mahasiswa sebagai dasar pengembangan sistem rekomendasi beasiswa. Data yang digunakan berupa 1.042 dokumen biodata dan portofolio mahasiswa dalam format PDF yang memuat informasi akademik, pengalaman organisasi, aktivitas kemahasiswaan, serta kondisi sosial ekonomi keluarga. Tahapan penelitian meliputi ekstraksi teks dari dokumen, pembersihan dan normalisasi data, pembentukan representasi fitur menggunakan metode <em>Term Frequency-Inverse Document Frequency</em> (TF-IDF), serta integrasi fitur tekstual, numerik, dan kategorikal untuk menghasilkan profil kompetensi yang lebih representatif. Selanjutnya, algoritma <em>Random Forest</em> diterapkan untuk mengklasifikasikan kelayakan penerima beasiswa dan menghasilkan skor rekomendasi bagi setiap kandidat. Evaluasi model dilakukan menggunakan skema pembagian data sebesar 75% untuk pelatihan dan 25% untuk pengujian. Hasil penelitian menunjukkan bahwa model mampu mencapai tingkat akurasi sebesar 74%, nilai <em>F1-Score</em> sebesar 0,46 pada kelas penerima beasiswa, dan nilai <em>Average Precision</em> sebesar 0,526. Temuan tersebut mengindikasikan bahwa pendekatan NLP efektif dalam mengidentifikasi dan mengolah informasi kompetensi yang tersebar pada portofolio digital sehingga dapat mendukung proses seleksi beasiswa yang lebih objektif, cepat, dan terukur. Selain itu, sistem yang dikembangkan mampu menghasilkan pemeringkatan kandidat berdasarkan tingkat kelayakan yang diperoleh dari kombinasi aspek akademik, pengalaman organisasi, serta kondisi sosial ekonomi, sehingga dapat menjadi alternatif solusi dalam meningkatkan kualitas pengambilan keputusan pada program beasiswa berbasis data.</p>2026-06-28T16:21:27+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/10235Pengelompokan Tingkat Kesejahteraan dan Tekanan Psikologis Mahasiswa Menggunakan Mental Health Inventory-38 dengan Algoritma K-Means2026-06-29T21:41:10+07:00Yadullah Asy-syakiri12250115109@students.uin-suska.ac.idElvia Budianitaelvia.budianita@uin-suska.ac.idIwan Iskandariwan.iskandar@uin-suska.ac.idFadhilah Syafriafadhilah.syafria@uin-suska.ac.idYuli Widiningsihyuli.widiningsih@uin-suska.ac.id<p>Student mental health is an urgent issue, particularly in STEM environments with high academic demands. This study aims to categorize the mental health patterns of students in the Faculty of Science and Technology at UIN Sultan Syarif Kasim Riau from the 2022–2025 cohorts using the K-Means Clustering algorithm. Data were collected via the MHI-38 questionnaire, which was adapted into Indonesian and validated by a clinical psychologist. Of the 559 data points, 522 valid data points were used after the data cleaning stage. Clustering evaluation utilized the Davies-Bouldin Index (DBI) and the Silhouette Coefficient with two distance calculation methods Euclidean and Manhattan across the range of k=2 to 10. The best configuration was obtained at k=2; Euclidean Distance yielded a DBI of 2.03 and a Silhouette of 0.18, while Manhattan Distance yielded a DBI of 2.19 and a Silhouette of 0.20. The clustering formed two clusters: Cluster 0 consisted of 334 students (64%) categorized as having psychological well-being with an average score of 3.82, and Cluster 1 consisted of 188 students (36%) categorized as experiencing psychological stress with an average score of 2.98. Cross-cohort analysis showed that the 2022 cohort was the most stable (72.4% in Cluster 0), while the 2023 and 2024 cohorts had a proportion of Cluster 1 of around 40%. By major, Computer Science was dominated by Cluster 1 (61.8%), while Mathematics and Electrical Engineering were dominated by Cluster 0 (83.3% and 82%, respectively). These results are expected to provide information to program administrators so they can evaluate their courses.</p>2026-06-28T19:24:07+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/10443A Meta-Synthesis of Factual Accuracy and Citation Hallucination in LLM Academic Assistants2026-06-30T18:30:31+07:00Rizki Anantamarizkianantama@unibamadura.ac.idMohammad Iqbal Bachtiariqbalbachtiar@unibamadura.ac.idZeinor Rahmanzeinorrahman@unibamadura.ac.idMohammad Ilham Bahriilhambahri@unibamadura.ac.id<p>The integration of Large Language Models (LLMs) in higher education presents a paradox between learning efficiency and the risk of misinformation due to the hallucination phenomenon. This study aims to comprehensively evaluate the factual accuracy and referential integrity of LLMs when acting as academic assistants. This research employs a comparative quantitative design through secondary data synthesis from three main empirical studies extracted from global databases. Independent variables include LLM model type, academic discipline, and prompt complexity, while dependent variables encompass concordance rate, citation fabrication rate, and Levenshtein distance deviation on Digital Object Identifiers (DOI). The results indicate that LLMs achieve factual accuracy above 90% on structured analytical tasks but show fatal vulnerability in referential integrity, with citation fabrication rates reaching 55% in GPT-3.5 and DOI hallucination reaching 89.4% in the humanities domain. These findings prove that students' trust in LLM outputs must not be absolute. The novelty of this research lies in the formulation of the "Dual-Layer Evaluation Framework" which separates conceptual validity from referential validity, providing an empirical foundation for educational institutions to formulate stricter digital literacy policies and the development of retrieval-augmented generation-based mitigation systems.</p>2026-06-28T00:00:00+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/10446Performance Evaluation of a Real-Time Teacher Attendance System Based on Haar Cascade and LBPH2026-06-30T18:34:48+07:00Zeinor Rahmanzeinorrahman@unibamadura.ac.idAli Fikrialifikri.student@unibamadura.ac.idRizki Anantamarizkianantama@unibamadura.ac.idMohammad Iqbal Bachtiariqbalbachtiar@unibamadura.ac.id<p>The teacher attendance process at Madrul Muttaqien Madrasah Diniyah is still conducted conventionally, creating the potential for recording errors and attendance data manipulation. This study develops an automated attendance application based on image processing to detect and recognize faces. The face detection process utilizes the Haar Cascade Classifier method, while face recognition is performed using the Local Binary Patterns Histograms (LBPH) method. System performance was evaluated through two testing scenarios: (1) a recall test involving 15 facial samples under various conditions, including shadow-covered faces, low-light environments, and tilted head positions, which achieved a recall value of 100%; and (2) a distance test conducted at distances of 40.27 cm, 57.69 cm, 63.16 cm, and 70 cm. The system successfully detected and recognized faces with a recall value of 100% at distances up to 63.16 cm but failed to recognize faces at a distance of 70 cm. These results indicate that the application is highly reliable under challenging lighting and orientation conditions and remains effective at distances of up to approximately 63 cm, beyond which its detection capability decreases significantly</p>2026-06-29T01:01:28+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/10294Perbandingan Teorema Bayes dan Dempster Shafer untuk Identifikasi Penyakit Kulit pada Kucing2026-06-30T20:52:26+07:00Ila Yati Betiilayb@unived.ac.idAchmad Fikri Sallabyfikrisallaby@unived.ac.idDeti Karmanitadetikarmanita@unived.ac.idSaika Damayantisaikadamayanti@unived.ac.id<p><strong>Abstrak</strong><strong>−</strong>Penyakit kulit merupakan salah satu masalah kesehatan yang paling sering terjadi pada kucing domestik, namun diagnosis klinisnya sering kali dihadapkan pada ketidakpastian akibat kemiripan gejala antar-penyakit. Penelitian ini bertujuan untuk membandingkan efektivitas dan karakteristik performa antara metode Teorema Bayes dan Data-Driven Dempster-Shafer dalam mengidentifikasi penyakit kulit pada kucing berbasis pemrograman Python. Data penelitian yang digunakan bersumber dari 120 rekam medis klinis yang mencakup diagnosis Flea Allergy (68 kasus), Ringworm (21 kasus), Dermatitis (4 kasus), Scabies (4 kasus), dan kondisi Healthy atau sehat (23 kasus). Nilai densitas awal (belief matrix) pada metode Dempster-Shafer dibentuk secara objektif menggunakan probabilitas kemunculan gejala pada 97 populasi kucing sakit (data-driven approach). Simulasi pengujian dilakukan secara komparatif dengan menginput kombinasi gejala klinis aktif G1, G2, dan G3. Hasil analisis menunjukkan bahwa kedua metode secara konsisten menempatkan Flea Allergy sebagai diagnosis dengan nilai kepastian tertinggi. Metode Teorema Bayes menghasilkan nilai probabilitas posterior individual yang mutlak sebesar 89,02%, sekaligus tetap mendeteksi probabilitas status Healthy sebesar 3,50%. Sementara itu, metode Dempster-Shafer mampu mereduksi ambiguitas informasi (uncertainty) dengan mengonvergensikan hasil secara spesifik pada subset kelompok penyakit {Flea Allergy, Ringworm, Scabies} dengan nilai belief sebesar 69,07%, serta langsung mengeliminasi opsi Healthy menjadi 0,00% karena tidak relevan dengan gejala aktif. Penelitian ini membuktikan bahwa integrasi kedua metode berbasis Python menghasilkan kalkulasi yang 100% konsisten antara perhitungan manual dan sistem, sehingga layak diimplementasikan sebagai sistem penunjang keputusan klinis yang objektif bagi praktisi medis hewan.</p> <p> </p>2026-06-30T20:40:16+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/10449Klasifikasi Popularitas Buku pada Platform Digital Menggunakan Random Forest Berbasis Fitur Perilaku Pembaca2026-07-01T23:44:23+07:00Al Imranimranalimran160703@gmail.comSandi Firmasyahsandifirmansyah17507@gmail.comSiti Intan Nurfadilahoaku4544@gmail.comMuhammad Alamsyahmalamsyah149@gmail.comFarizi Ilhamdosen02954@unpam.ac.id<p>The assessment of book popularity on digital community platforms has long relied on single indicators such as average ratings, while popularity is a multidimensional construct involving broader community engagement. This study aims to: (1) analyze reader behavior patterns as a popularity indicator for digital books in the Book-Crossing dataset; (2) develop a multiclass popularity classification model using the Random Forest algorithm; and (3) generate strategic recommendations for authors, publishers, and digital platform managers. A quantitative approach was employed utilizing secondary data from the Book-Crossing dataset, which after preprocessing yielded 10,234 books eligible for analysis. A Random Forest Classifier was constructed using six reader behavior and book metadata features, with popularity labels (Popular, Moderately Popular, Less Popular) determined through quartile distribution analysis to address class imbalance. Results show the model achieved an accuracy of 91.2% with an F1-Score of 0.94 for the Popular category, while popularity distribution follows a power law pattern in which 10.0% of books fall into the Popular category and 58.2% reside in the Less Popular category. These findings demonstrate that combining multivariable reader behavior features can produce accurate and measurable popularity classification, while also confirming the popularity bias inherent in digital content ecosystems.</p>2026-06-30T00:00:00+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/10391Hybrid Weighted K-Nearest Neighbor dan Naive Bayes untuk Klasifikasi Diabetes Melitus2026-07-03T00:52:29+07:00Aklima Laduna Ramadyaladunaaklima67@gmail.comSolikhun Solikhunsolikhun@amiktunasbangsa.ac.idTimbo Faritcan P. Siallagantimbofaritcansiallagan@gmail.com<p><strong>Abstract</strong><strong>−</strong>Diabetes mellitus is a chronic metabolic disease with a continuously increasing prevalence and the potential to cause various serious complications if not detected early. The utilization of machine learning has become one of the effective approaches to support a faster and more accurate diagnostic process. This study proposes a Hybrid Weighted K-Nearest Neighbor (WKNN)–Naive Bayes model that combines a distance-based approach with a probabilistic approach for the classification of diabetes mellitus. The dataset used is the Pima Indians Diabetes Dataset, consisting of 768 patient records with eight predictor attributes and one target variable. The preprocessing stage included missing value handling using imputation, data normalization using RobustScaler, feature engineering, and class balancing using the Synthetic Minority Over-sampling Technique (SMOTE). Model evaluation was performed using Stratified 10-Fold Cross Validation and compared against K-Nearest Neighbor (KNN), Weighted K-Nearest Neighbor (WKNN), and Naive Bayes models. The results show that the Hybrid WKNN–Naive Bayes model achieved an Accuracy of 88.50%, Precision of 85.32%, Recall of 93.00%, F1-Score of 88.99%, and AUC-ROC of 91.09%, demonstrating good classification performance for diabetes mellitus diagnosis. The Weighted K-Nearest Neighbor (WKNN) model also achieved an Accuracy of 88.50%, with a Recall of 94.00%, F1-Score of 89.10%, and AUC-ROC of 91.73%, while the Naive Bayes model produced an Accuracy of 71.00%. These results indicate that the Hybrid WKNN–Naive Bayes approach is able to deliver competitive classification performance through the combination of distance-based and probabilistic methods, making it a promising candidate for use as a decision support system in the early detection of diabetes mellitus.</p>2026-06-30T00:00:00+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bulletinds/article/view/10534Integrasi BWM–MARCOS untuk Seleksi Pemasok Pangan yang Konsisten dan Robust pada Kafetaria Universitas2026-07-13T20:42:51+07:00Erienika Lompoliuerienika.lompoliu@unklab.ac.idRegi Fernando Najoanreginajoan@unklab.ac.idWilsen Grivin Mokodaserwilsenm@unklab.ac.idGeorge M. W. Tangkagmwtangka@gamil.com<p>The cafeteria of XYZ University serves as an essential service hub for students, particularly dormitory residents. Therefore, the selection of food suppliers must systematically consider quality, food safety, delivery reliability, service, and cost. This study integrates two Multi-Criteria Decision-Making (MCDM) methods—Best Worst Method (BWM) and Measurement of Alternatives and Ranking according to COmpromise Solution (MARCOS)—to produce a more objective, consistent, and transparent supplier selection model. BWM determines criterion weights through linear optimization with a consistency ratio of CR = 0.08 ≤ 0.25, confirming valid weights. These weights are then applied in MARCOS to generate the final supplier ranking. Five evaluation criteria are used: quality & freshness (C1 = 0.35), delivery reliability (C2 = 0.25), food safety compliance (C3 = 0.20), service responsiveness (C4 = 0.15), and price (C5 = 0.05), applied to four supplier alternatives (Alpha, Beta, Gamma, Delta). Results show the final ranking Alpha ≻ Gamma ≻ Beta ≻ Delta with utility values of 1.2534, 1.1336, 0.9921, and 0.8224 respectively. Sensitivity analysis with ±10% weight variation confirms ranking stability, demonstrating the model's robustness. This model provides an auditable numerical framework, easy to communicate to non-technical stakeholders, and serves as a basis for procurement policy and periodic supplier performance evaluation.</p>2026-07-13T00:00:00+07:00##submission.copyrightStatement##