https://ejurnal.seminar-id.com/index.php/bits/issue/feedBuilding of Informatics, Technology and Science (BITS)2026-09-20T12:44:12+07:00Support Journalseminar.id2020@gmail.comOpen Journal Systems<p style="text-align: justify;">Building of Informatics, Technology and Science (BITS) is an open-access media in publishing scientific articles that contain the results of research in information technology and computers. Paper that enters this journal will be checked for plagiarism and peer-review first to maintain its quality. This journal is managed by Forum Kerjasama Pendidikan Tinggi (FKPT) published 4 times a year in <strong>June (No 1), September (No 2), December (No 3), </strong>and <strong>March (No 4) </strong>with ISSN <a href="https://issn.brin.go.id/terbit/detail/1557033587" target="_blank" rel="noopener">2684-8910 (Print)</a> and <a href="https://issn.brin.go.id/terbit/detail/1557037175" target="_blank" rel="noopener">2685-3310 (Online)</a>. The existence of this journal is expected to develop research and make a real contribution to improving research resources in the field of information technology and computers. BITS Journal, indexed by : <a href="https://scholar.google.com/citations?user=oy-dtP8AAAAJ&hl=id&citsig=AMD79orr29I2On4MNhRIxcFHJxCpCrUMQA">Google Scholar</a> | <a href="https://garuda.kemdikbud.go.id/journal/view/15844">Portal Garuda </a>| <a href="https://app.dimensions.ai/discover/publication?search_mode=content&search_text=10.47065&search_type=kws&search_field=full_search&and_facet_source_title=jour.1407312">Dimensions</a> | <a href="https://onesearch.id/Search/Results?lookfor=Building+of+Informatics%2C+Technology+and+Science+%28BITS%29&type=AllFields&limit=20&sort=relevance">Indonesia One Search</a> | <a href="https://moraref.kemenag.go.id/archives/journal/98984515036262163">Moraref</a> | <a href="https://index.pkp.sfu.ca/index.php/browse/index/10161">PKP Index</a> | <a href="https://www.scilit.net/journal/6109244">SCILIT</a> | <a href="https://explore.openaire.eu/search/dataprovider?datasourceId=issn___print::8c94c96cf14c5cea949a4b30da0dcea5">OpenAire</a> | <a href="https://portal.issn.org/resource/ISSN/2685-3310">ROAD</a> | <a href="https://search.crossref.org/?q=Building+of+Informatics%2C+Technology+and+Science+%28BITS%29&from_ui=yes">Crossref</a> | <a href="https://sinta.kemdikbud.go.id/journals/profile/7790">Science and Technology Index (Peringkat SINTA 3)</a> | <a href="https://www.base-search.net/Search/Results?type=all&lookfor=2685-3310&ling=1&oaboost=1&name=&thes=&refid=dcresen&newsearch=1">BASE</a> | <a href="https://www.worldcat.org/search?q=2685-3310&qt=results_page">Worldcut.Org.</a><br><strong>Building of Informatics, Technology and Science (BITS)</strong>, has been reaccredited with a <strong>SINTA rating of 3</strong> through the Decree of the Director General of Strengthening Research and Development of the Ministry of Research, Technology and Higher Education based on number <a href="https://drive.google.com/file/d/1Lq3pCoZZmZwoZMSVsAuCM-0seprhkwee/view?usp=sharing">72/E/KPT/2024</a>, dated April 1, 2024 regarding the results Electronic Scientific Periodic Accreditation Period I 2024 from <strong>Volume 5 No 1 (2023)</strong> to <strong>Volume 9 No 4 (2028)</strong>.</p>https://ejurnal.seminar-id.com/index.php/bits/article/view/10794Comparative Customer Segmentation Pipelines for E-Commerce Using K-Means-KNN and UMAP-K-Means-XGBoost2026-09-08T16:23:21+07:00Dzidan Aditya Gumilangdzdnaditya@gmail.comEndang Lestari Ruskanendanglestari@unsri.ac.idArdina Arianiardinaariani@unsri.ac.idKen Dhita Taniakenya.tania@gmail.comAhmad Rifairifai.bae@gmail.com<p>The rapid expansion of e-commerce has generated massive volumes of customer data that remain underutilized for supporting Customer Relationship Management (CRM) strategies. Conventional customer segmentation approaches commonly employ a pipeline consisting of K-Means clustering followed by K-Nearest Neighbors (KNN) classification. However, this approach exhibits limitations in handling high-dimensional data and maintaining classification performance on large-scale datasets. This study presents a comparative analysis of two customer segmentation pipelines: the conventional K-Means-KNN pipeline and the proposed Uniform Manifold Approximation and Projection (UMAP)-K-Means-XGBoost pipeline. The experiments were conducted using the E-Commerce Shopper Behavior & Lifestyle dataset, comprising approximately one million customer records and eight selected features representing transactional, psychographic, and financial behavioral characteristics. Clustering performance was evaluated using the Silhouette Score, Davies-Bouldin Index, and Calinski-Harabasz Index, while classification performance was assessed using accuracy, precision, recall, and F1-score. Experimental results demonstrate that incorporating UMAP improves cluster separability by preserving the intrinsic structure of high dimensional data, whereas XGBoost consistently outperforms KNN in downstream classification, achieving an accuracy exceeding 99%. These findings indicate that the UMAP-K-Means-XGBoost pipeline provides a more robust, scalable, and interpretable framework for customer segmentation, thereby offering more reliable decision support for data-driven CRM strategies in e-commerce environments.</p>2026-09-08T16:19:16+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bits/article/view/10936Decision Support System for Determining Signature Menus Using the Integration of BMW, MOORA, and Copeland Scores2026-09-09T10:32:12+07:00Meta Ardi Setiawanmetaflash55@gmail.comAnteng Widodoanteng.widodo@umk.ac.idNoor Latifahnoor.latifah@umk.ac.id<p>Menu optimisation is a crucial yet challenging challenge for business sustainability because the speciality coffee sector is marked by dynamic consumer tastes and a high degree of product diversity. This study tackles a basic operational issue at Inti Coffee, where a subjective, intuition-based method is currently used to identify "signature offerings"the important menu items that ought to be given priority for promotion, inventory, and resource allocation. This current process, which mainly depends on the owner, operational manager, and lead barista's intuition, is intrinsically vulnerable to individual cognitive biases, personal taste preferences, and inconsistent evaluation, which frequently results in less than ideal menu performance and lost revenue opportunities. The main goal of this research is to develop and deploy an online Decision Support System (DSS) that offers a transparent, data-driven, and organised framework for objectively ranking menu items in order to get around these restrictions. A hybrid multi-criteria decision-making (MCDM) technique is included into the suggested system to guarantee group unanimity and robustness. In order to minimise pairwise comparison inconsistencies and capture the knowledge of the three primary decision-makers, the Best-Worst Method (BWM) is first used to systematically extract the relative relevance weights of six different evaluation criteria. Second, a thorough dataset of 500 real sales transactions is used to assess and rank 51 menu alternatives using the MOORA (Multi-Objective Optimisation on the basis of Ratio Analysis) method. Both financial parameters (total items sold, HPP or cost of goods sold, and profit margin) and operational characteristics (uniqueness of taste score, preparation time, and ingredient lifetime) are included in the evaluation criteria. In order to successfully resolve any potential conflicts between decision-makers, the Copeland Score is finally used to combine the individual preference rankings into a single, collective group score. This study makes three main contributions: first, it develops a novel integrated DSS framework that integrates BWM, MOORA, and Copeland Score into a single unified workflow specifically designed for signature menu selection; second, it involves three different stakeholders (owner, operational manager, and head barista) in the group decision-making process, ensuring that the final recommendation reflects a balanced consensus rather than individual bias; and third, it creates a fully functional web-based system with statistical validation tools that empirically verify the accuracy and dependability of the recommendations using Spearman correlation, RMSE, and overlap ratio against actual customer preferences. The creation of a useful, web-based DSS that turns menu curating from an art to a science and offers a reproducible model for other food and beverage businesses is the main contribution of this research. With a MOORA score of 0.249180 and a Copeland Score of 50, the interim results show that the system is able to identify "Kopi Susu Aren" as the best-performing signature item. A poll involving 130 participants was carried out to verify the system's output against human judgement in the actual world. In addition to a low Root Mean Square Error (RMSE) of 5.5734 and an overlap ratio of 66.7%, the results show a strong positive correlation (Spearman's rho = 0.9238) between the system's ranks and participant feedback, proving the system's high accuracy and practical applicability. In order to provide scalability and usability for continuous operational choices, the DSS is implemented using a combination of PHP Native, Python, MySQL, and a testing dashboard based on Streamlit.</p>2026-09-08T00:00:00+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bits/article/view/10680Perbandingan XGBoost dan Random Forest Untuk Prediksi Pembatalan Pesanan Marketplace Berbasis SHAP2026-09-09T10:45:22+07:00Dian Pertiwi111202214602@mhs.dinus.ac.idUsman Sudibyousman.sudibyo@dsn.dinus.ac.id<p>The high rate of order cancellations in marketplaces is a challenge that can impact operational efficiency, inventory management, and customer service quality. The ability to identify orders potentially canceled early is crucial so businesses can take more appropriate mitigation measures. The purpose of this study is to compare the performance of the Random Forest, and XGBoost algorithms in the marketplace order cancellations prediction process and to interpret the factors that influence predictions using SHAP (SHapley Additive exPlanations)-based explainable machine learning. The study used a marketplace transaction dataset that had undergone data cleaning, feature transformations, and class through the application of the Synthetic Minority Oversampling Technique (SMOTE) method. Next, the model performance is evaluated using the Accuracy, Precision, Recall, and F1-Score metrics. Four model scenarios were tested: Random Forest + SMOTE, Random Forest Tuned + SMOTE, XGBoost + SMOTE, and XGBoost Tuned + SMOTE. Based on the research results, it is indicated that the XGBoost Tuned + SMOTE model produces the best performance with accuracy values reaching 86.62%, precision reaching 51.43%, recall reaching 34.95%, and F1-Score reaching 41.62%. Meanwhile, the XGBoost + SMOTE model produced the highest recall of 37.09% with an F1-Score of 41.48%. Considering that the research objective is to detect potentially canceled orders in imbalance data, the XGBoost + SMOTE model was chosen as the best model because it produced the highest recall of 37,09% with a competitive F1-Score of 41,48%, making it more capable of detecting minority classes compared to other models. SHAP analysis showed that several transaction features contribute more dominantly to the prediction results produced by the model. These findings can be used as a basis for decision-making to reduce the risk of order cancellations and improve the effectiveness of marketplace operations.</p>2026-09-08T00:00:00+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bits/article/view/10357Performance Evaluation of Naive Bayes, SVM, LSTM, and IndoBERT for Fashion Review Sentiment Classification2026-09-09T11:13:29+07:00Muhammad Iqbal Ismaildosen03423@unpam.ac.idDenies Susantodosen02890@unpam.ac.idIndar Riyantodosen10118@unpam.ac.idMelinne Maldini Rosadydosen03352@unpam.ac.id<p>The rapid growth of e-commerce has increased the volume of customer reviews, making it difficult for fashion businesses to manually identify customer sentiment, product concerns, and satisfaction drivers. This study aims to evaluate the performance of Naive Bayes, Support Vector Machine (SVM), Long Short-Term Memory (LSTM), and IndoBERT in classifying sentiment polarity in Indonesian fashion product reviews, while also identifying topic-level sentiment patterns using Latent Dirichlet Allocation (LDA). The dataset consisted of 23,487 Shopee Indonesia fashion reviews collected during the 2025 observation period. After data cleaning, duplicate removal, text normalization, tokenization, stopword removal, and lemmatization, 22,961 valid reviews were retained and categorized into positive, neutral, and negative classes using rating-based labeling. To reduce majority-class bias, stratified data splitting and class-weighted learning were applied, while five-fold cross-validation was used to evaluate model stability. The results show that IndoBERT achieved the highest performance with an accuracy of 93.41%, precision of 92.87%, recall of 92.15%, and F1-score of 92.51%, outperforming LSTM with 90.80% accuracy, SVM with 88.90%, and Naive Bayes with 84.60%. The findings indicate that transformer-based contextual representation is more effective in handling noisy Indonesian fashion review text, including informal expressions, abbreviations, and context-dependent sentiment. Topic-based analysis further revealed that positive sentiment was mainly associated with product quality, material comfort, design, and value, while negative sentiment was driven by sizing mismatch, product inconsistency, delivery delay, and customer service issues. This study contributes to sentiment classification research by providing a comparative evaluation of machine learning, deep learning, and transformer-based models, and offers practical insights for improving fashion e-commerce product communication, customer satisfaction, and data-driven digital marketing decisions.</p>2026-09-08T00:00:00+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bits/article/view/10507Komparasi Algoritma Machine Learning pada Prediksi Kemenangan Tim VALORANT Berbasis Delta Analysis2026-09-09T11:23:31+07:00Bintang Ilham Kurniawanbintangilhamkurniawan@gmail.comSri Winarnosri.winarno@dsn.dinus.ac.id<p>The development of the Valorant esports ecosystem through the VALORANT Champions Tour (VCT) creates a need for objective data analysis. Current analysis is often trapped in subjective bias, risking inaccurate predictions. To address this problem, this research compares various team win prediction methods in VCT 2025. The selection of the VCT 2025 dataset provides novelty as it represents a franchise league structure with much more evenly distributed game meta and tactical complexity than previous seasons. The compared algorithms include Naïve Bayes, Decision Tree, XGBoost, CatBoost, and Stacking Classifier. The main contribution of this research is the application of a differential feature extraction technique (Delta Analysis) to objectively measure the performance gap between teams. Comparative experimental results on 1,507 match observations show that the Stacking Classifier achieved the most optimal accuracy (75.72%) compared to XGBoost (75.19%), Naïve Bayes (74.26%), CatBoost (73.59%), and Decision Tree (72.86%). These findings prove that Delta Analysis-based meta-learning synergy can provide more stable analytical generalization in handling modern esports data complexity.</p>2026-09-08T00:00:00+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bits/article/view/10945Superioritas Optimizer AdaGrad pada Arsitektur Spasio-Temporal PolyFace-LSTM untuk Prediksi Kepribadian Berbasis Video2026-09-09T11:48:51+07:00Muhammad Zidan Alif Oktavianm.zidanalif123@gmail.comSalamun Rohman Nudinsalamunrohman@unesa.ac.id<p>Conventional personality assessment generally relies on psychometric questionnaires, such as the Big Five Inventory (BFI). However, this method is susceptible to self-report bias and requires a relatively long evaluation time. On the other hand, the development of automated video-based personality analysis systems faces high computational challenges in simultaneously integrating facial spatial features and the temporal dynamics of micro-expressions. This study aims to propose a solution in the form of a hybrid spatio-temporal neural network architecture utilizing transfer learning techniques to predict personality more objectively. The proposed system integrates the Multi-Task Cascaded Convolutional Networks (MTCNN) algorithm for automatic face detection, the PolyFace model as a spatial feature extractor, and the Long Short-Term Memory (LSTM) algorithm to model temporal relationships between frames. Experiments were conducted by comparing three optimization algorithms, namely Adam, AdaGrad, and SGD, using the ChaLearn LAP 2017 dataset. The results show that the AdaGrad optimizer achieved the best generalization and prediction performance during the validation phase, with an average accuracy of 89.13%, outperforming SGD (88.34%) and Adam (88.33%). Theoretically and mathematically, the superiority of AdaGrad in this architecture is attributed to its ability to dynamically adjust the learning rate based on the accumulation of historical gradients. This mechanism makes AdaGrad considerably more stable in extracting complex spatio-temporal micro-expression features without becoming trapped in local minima. The highest performance achieved by AdaGrad was observed in the Agreeableness dimension, reaching an accuracy of 90.14%. The main contribution of this study is the development of an efficient (low-latency), precise spatio-temporal personality prediction model that avoids local minima, making it suitable for implementation in real-world automated psychological detection systems.</p>2026-09-08T00:00:00+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bits/article/view/10602Implementasi Teknik SMOTE Menggunakan Random Forest dan XGBoost pada Klasifikasi Tingkat Kualitas Udara2026-09-09T12:13:55+07:00Nanda Putri Karizkinanda_putri_karizki@teknokrat.ac.idHeni Sulistianihenisulistiani@teknokrat.ac.id<p>Air pollution is a global environmental issue with significant impacts on public health and environmental sustainability. The problem is compounded by the generally imbalanced class distribution in air quality data; consequently, the highest-risk category "Hazardous" often constitutes the minority class, making it the most difficult to detect using conventional classification models. This study implements a machine learning approach using Random Forest and XGBoost algorithms combined with the Synthetic Minority Oversampling Technique (SMOTE) to classify air quality levels. The dataset comprises 5,000 samples featuring nine environmental and demographic variables. Air quality is categorized into four classes Good, Moderate, Poor, and Hazardous with an initial imbalanced distribution (Good: 40%, Moderate: 30%, Poor: 20%, Hazardous: 10%). SMOTE was applied exclusively to the training data to balance the class distribution. Results indicate that XGBoost combined with SMOTE achieved the best performance, yielding an accuracy of 0.952, an F1-Score of 0.952, and an average cross-validation score of 0.969. This represents a 0.022 improvement over a previous study that utilized a Decision Tree model without SMOTE (achieving 0.930 accuracy). For the "Hazardous" minority class, the XGBoost-SMOTE combination improved recall from 0.80 (Random Forest without SMOTE) to 0.87, while also achieving more balanced precision and F1-Score values (0.87). CO levels and proximity to industrial areas emerged as the dominant features, contrasting with PM2.5 in the earlier study. These findings confirm that combining ensemble methods with SMOTE effectively addresses class imbalance. Evaluation was conducted using precision, recall, and F1-Score metrics, alongside 5-fold stratified cross-validation to ensure model stability across all classes including the "Hazardous" minority class, which is most critical for public health. These results suggest that combining ensemble algorithms with data balancing techniques can serve as a practical reference for developing machine learning-based air quality monitoring systems.</p>2026-09-08T00:00:00+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bits/article/view/10560Optimasi Hyperparameter Grid Search untuk Multi-Algoritma Klasifikasi Tingkat Hidrasi pada Daily Water Intake2026-09-09T12:25:04+07:00Julfiana Rizkiah Futrijulfiana_rizkiah_futri@teknokrat.ac.idHeni Sulistianihenisulistiani@teknokrat.ac.id<p>Water is a vital element for sustaining human physiological functions, ranging from metabolic processes to cognitive and physical performance. Insufficient daily water intake can trigger various health problems that significantly reduce individual well being. Machine learning provides a data driven approach to classify hydration levels more objectively based on individual and environmental attributes. This study evaluates and compares the performance of three classification algorithms Random Forest, Naive Bayes, and K-Nearest Neighbor in predicting hydration status (Good or Poor) using the Daily Water Intake public dataset comprising 30,000 samples. Class imbalance in the dataset (79.7% Good, 20.3% Poor) was addressed through SMOTE, while Stratified K-Fold Cross Validation with K=10 was employed for model evaluation. Hyperparameter optimization was performed using Grid Search, and model performance was assessed through accuracy, precision, recall, F1-Score, and AUC-ROC metrics. Results demonstrate that Random Forest achieved the highest accuracy of 99.50% and AUC-ROC of 0.9999 after optimization thanks to its ensemble mechanism, which captures non-linear relationships among features, outperforming Naive Bayes, which is constrained by its feature-independence assumption. KNN showed the most notable improvement post optimization with an accuracy gain of 0.62% (from 97.62% to 98.24%), while Naive Bayes remained unchanged at 83.53% as its optimal parameter matched the default value. These findings offer evidence based guidance for algorithm selection in hydration classification systems.</p>2026-09-08T00:00:00+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bits/article/view/10763Comparative Evaluation of Machine Learning Models with Explainability and Fairness for Diabetes Risk Prediction2026-09-20T12:44:12+07:00Vanya Dwi Nabilavanyadwinabila@gmail.comBayu Wijaya Putrabayuwisata@gmail.comM Rudi Sanjayam.rudi.sjy@ilkom.unsri.ac.idDwi Rosa Indahindah812@unsri.ac.id<p>Diabetes mellitus remains a major global health challenge, making early risk prediction essential for timely intervention and prevention. This study proposes a trustworthy machine learning framework for diabetes risk prediction by integrating predictive performance, explainability, and fairness evaluation. Six classification algorithms Logistic Regression, Decision Tree, Random Forest, XGBoost, LightGBM, and CatBoost were comparatively evaluated, with LightGBM selected as the best baseline model. Hyperparameter optimization was subsequently performed using Optuna with macro F1-score as the optimization objective to better address the imbalanced multiclass nature of the dataset. Model performance was assessed using Accuracy, Balanced Accuracy, Precision, Recall, F1-score, and Receiver Operating Characteristic–Area Under the Curve (ROC-AUC). Model explainability was analyzed using SHapley Additive exPlanations (SHAP), while fairness was evaluated using Fairlearn based on gender and race through Demographic Parity Difference and Equalized Odds Difference.Experimental results show that hyperparameter optimization increased Balanced Accuracy from 0.3755 to 0.4891 and macro F1-score from 0.3797 to 0.4333, indicating improved recognition of minority classes. Although overall Accuracy decreased from 0.8357 to 0.7023, this trade-off reflects a more balanced classification across diabetes categories, which is preferable for imbalanced clinical datasets where identifying minority cases is essential for early risk detection. SHAP analysis identified Body Mass Index (BMI), Age, Physical Health Days, Mental Health Days, and Income Level as the most influential predictors. Fairness evaluation further demonstrated low demographic disparities across gender and race. These findings demonstrate that integrating predictive performance, explainability, and fairness enables the development of a more transparent, equitable, and clinically reliable diabetes risk prediction framework.</p>2026-09-08T00:00:00+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bits/article/view/9490Klasifikasi Status Gizi Bayi sebagai Upaya Deteksi Dini Masalah Gizi Menggunakan Algoritma Naive Bayes2026-09-09T13:05:14+07:00Muhammad Rafif ARrafif180507@gmail.comFathoni Fathonifathoni@unsri.ac.idKen Ditha Taniakenditha@ilkom.unsri.ac.idAllsela Meirizaallsela_meiriza@yahoo.co.idAli Ibrahimaliibrahim210784@gmail.com<p>The classification of infant nutritional status remains a significant challenge in the utilization of health data, particularly in developing predictive models that are accurate and reliable for supporting decision-making at the primary healthcare level. This study aims to implement the Naïve Bayes algorithm to classify infant nutritional status based on anthropometric indicators, namely weight-for-age (W/A) and height-for-age (H/A). The data used in this study were obtained from Posyandu activity reports at Citra Medika Public Health Center covering the period from 2023 to 2025, with a total of 10,654 infant records. The class distribution in the dataset includes normal, well-nourished, undernourished, severely undernourished, overnourished, and at risk of overnutrition categories. Since the class proportions are relatively balanced, no oversampling technique was applied. The research process involved several stages, including data cleaning, category normalization, class label transformation, and dataset splitting using stratified sampling with a composition of 80% training data and 20% testing data. The evaluation results indicate that the Naïve Bayes model achieved an accuracy of 82%. The precision values were 0.92 for the normal class, 0.39 for the malnutrition class, and 0.14 for the overweight class. The recall values were 0.89 for the normal class, 0.63 for the malnutrition class, and 0.08 for the overweight class. Meanwhile, the F1-scores were 0.83 for the normal class, 0.48 for the malnutrition class, and 0.11 for the overweight class. These findings suggest that the model demonstrates fairly good performance in classifying infant nutritional status based on anthropometric data. Therefore, the Naïve Bayes algorithm can be considered effective for classifying infant nutritional conditions using W/A and H/A indicators. This study is expected to contribute to the development of data-driven systems to enhance analytical quality and support decision-making in primary healthcare services.</p>2026-09-08T00:00:00+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bits/article/view/9667Analisis Disparitas Data Kebencanaan Indonesia sebagai Dasar Mitigasi Berbasis K-Means2026-09-09T13:18:33+07:00Vanisa Amalia Putrihello.vanisamalia@gmail.comFathoni Fathonifathoni@unsri.ac.idYadi Utamayadiutama@unsri.ac.id<p>Indonesia experiences a high frequency of disasters; however, differences in characteristics between national and global datasets create challenges for consistent disaster analysis. Previous studies generally rely on a single data source, limiting the ability to capture disparities in data representation comprehensively. This study aims to analyze disparities in disaster data representation and identify regional disaster patterns using a data mining approach. A Disaster Severity Index was constructed using Min-Max normalization, and the K-Means algorithm was applied to cluster regions based on disaster index and event frequency. The optimal number of clusters was determined using the Elbow Method and validated using the Silhouette Score, while a global dataset was used for comparison. The results indicate that the optimal model is achieved at K=4 with a Silhouette Score of 0.610883, indicating good cluster separation. Most regions exhibit moderate disaster characteristics, while a small number show extreme patterns with disproportionate relationships between frequency and severity. Differences between national and global datasets suggest variations in reporting mechanisms and data coverage. These findings demonstrate that relying on a single data source may lead to biased interpretations, highlighting the importance of multi-source data integration to improve the accuracy of disaster analysis.</p>2026-09-08T00:00:00+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bits/article/view/10960Analisis Perbandingan Random Forest dan XGBoost dalam Prediksi Pembatalan Pesanan Shopee Menggunakan Interpretasi SHAP2026-09-09T13:33:18+07:00Muhammad Bayu Samudra09031182328008@student.unsri.ac.idMira Afrinamira_afrina@unsri.ac.idAllsela Meirizaallsela@unsri.ac.idRizka Dhini Kurniarizkadhini@gmail.comArdina Arianiardinaariani@unsri.ac.id<p>Order cancellation is a significant issue for <em>e-commerce</em> platforms because it can result in revenue loss, increased operational costs, and reduced transaction management efficiency. This study aims to compare the performance of <em>Random Forest</em> and <em>Extreme Gradient Boosting</em> (XGBoost) in predicting customer order cancellations on Shopee and to interpret the factors contributing to the prediction outcomes using an <em>Explainable Artificial Intelligence</em> approach based on <em>SHapley Additive exPlanations</em> (SHAP). The study employed the <em>Shopee Consumer Behaviour Cancellation Order Analysis</em> dataset obtained from Kaggle. The research process consisted of data <em>preprocessing</em>, dataset partitioning using the <em>Stratified Train-Test Split</em> method with an 80:20 ratio, implementation of both algorithms, model evaluation using <em>Accuracy</em>, <em>Precision</em>, <em>Recall</em>, <em>F1-Score</em>, and <em>Receiver Operating Characteristic–Area Under Curve</em> (ROC-AUC), followed by model interpretation using SHAP. The evaluation results indicate that <em>Random Forest</em> achieved better performance across most evaluation metrics, while XGBoost obtained a slightly higher <em>ROC-AUC</em> value. SHAP analysis identified Total Discount, Buyer-Paid Shipping Cost, Estimated Shipping Cost, and Estimated Shipping Fee Deduction as the variables contributing most to order cancellation predictions.These findings indicate that combining predictive algorithms with an <em>Explainable Artificial Intelligence</em> approach not only supports the classification of order cancellations but also provides insights into the factors influencing prediction outcomes, which can serve as a consideration for reducing order cancellations on <em>e-commerce</em> platforms.</p>2026-09-08T00:00:00+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bits/article/view/10768Comparative IndoBERT, Logistic Regression, SVM, Random Forest, and Naïve Bayes for TikTok Deepfake Misinformation2026-09-09T18:45:18+07:00Tri Pidiantotriipiidiiantoo@gmail.comIwan Setiawan Wibisonoiwansetiawan@unw.ac.id<p>Generative AI has made deepfake content, especially face-swap videos and cloned voices, easier to produce and spread, and short-video platforms such as TikTok have become a major channel for this material. Most existing work still focuses on detecting the manipulated audio or video itself, leaving audience reactions in the comment section relatively unexplored. The problem this study addresses is that, without reading these comments, moderation systems have no way to know whether audiences are being misled by AI-generated content; the objective is to determine whether a fine-tuned IndoBERT model can flag such misinformation-leaning comments more reliably than conventional machine-learning classifiers. This paper takes a comparative approach to that gap, testing whether a fine-tuned IndoBERT model can flag potential misinformation in Indonesian-language TikTok comments more reliably than four TF-IDF-based classifiers: Logistic Regression, Random Forest, Multinomial Naive Bayes, and Linear SVM. Starting from 1,345 comments scraped across 15 TikTok videos, a multi-stage cleaning process left 1,071 usable comments, manually sorted into three labels: Misinformation, Skeptical, and Other. IndoBERT was fine-tuned with a weighted cross-entropy loss to offset class imbalance and evaluated through a train-validation-test split and Stratified 5-Fold Cross Validation. The fine-tuned model reached 82.41% accuracy and 81.58% F1-Macro, beating every machine-learning baseline by at least 13.84 points in F1-Macro, with five-fold results holding steady at an average F1-Macro of 80.17%. These results suggest IndoBERT is a stronger option than conventional machine learning for flagging misinformation-leaning comments on AI-generated content, offering a practical foundation for automated moderation on social platforms.</p>2026-09-08T00:00:00+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bits/article/view/10891Perbandingan Random Forest dan SVM dalam Analisis Sentimen Opini Masyarakat Terhadap Kenaikan Harga Pertamax2026-09-09T18:55:12+07:00Fahmy Wira Oktavierfahmy_wira_oktavier@teknokrat.ac.idAuliya Rahman Isnainauliyarahman@teknokrat.ac.id<p>The increase in Pertamax prices triggered various responses from the public, many of which were expressed through social media X. Due to the large number of opinions that emerged, sentiment analysis proved to be an appropriate way to automatically recognize public opinion trends. This study was conducted with the aim of determining how effective the "Random Forest" and "Support Vector Machine" (SVM) algorithms are in the process of grouping public sentiment regarding the increase in Pertamax prices. The research data was obtained by collecting information from platform X using keywords that match the topic of Pertamax. After filtering and removing duplicates, 5,006 records were obtained which were used as the dataset in this study. The steps of this study include data pre-processing, sentiment classification process, feature extraction using the Term Frequency Inverse Document Frequency (TF-IDF) technique, Using the SMOTE method to balance the number of classes and divide the dataset into categories, then classifying with two algorithms, namely Random Forest and Support Vector Machine (SVM). Confusion matrix is used to assess model performance with various metrics such as accuracy, precision, recall, and F1 score. This study shows that the data is dominated by negative sentiment, amounting to 51.64%, then dominated by positive sentiment at 35.60%, and neutral sentiment reached 12.76%. Before using SMOTE, Random Forest had an accuracy of 73% while SVM reached 75%. After using SMOTE, the accuracy of the Random Forest model increased to 85% and the accuracy of the SVM increased to 89%. This research contributes to improving the quality of sentiment analysis and serves as a reference for future studies; furthermore, it provides stakeholders with an understanding of public opinion trends regarding the price hike of Pertamax, thereby assisting them in decision-making.</p>2026-09-08T00:00:00+07:00##submission.copyrightStatement##https://ejurnal.seminar-id.com/index.php/bits/article/view/10916An Interval-Informed Hybrid CNN-LSTM for Ten-Category ECG Beat and Event Classification2026-09-09T19:11:01+07:00I Putu Bagus Erix Wijayaputu.bagus@bku.ac.idArda Ardiyansyahardaardiansyah24@gmail.comSatria Mandalamandalasatria@gmail.com<p>Automated electrocardiogram (ECG) classification remains challenging because waveform-based deep-learning models capture rich morphology, whereas clinically relevant conduction and timing intervals are not always represented explicitly. This study investigates whether a compact interval-informed representation can retain useful discriminative information for ten ECG beat and event categories. MIT-BIH Arrhythmia Database recordings were filtered and divided into seven-second segments. RR, PR, and QT intervals and QRS width were summarized using minimum, maximum, mean, median, skewness, and kurtosis, yielding 24 features. To prevent information leakage, the original samples were partitioned before oversampling; feature scaling and SMOTE were applied only to training data, while validation and held-out test data retained their original distributions. Leakage-controlled five-fold cross-validation yielded 92.15% ± 0.43% accuracy and 86.78% ± 0.68% macro-F1. On the untouched 826-sample test set, the model achieved 92.62% accuracy, 87.49% macro-F1, and 92.68% weighted F1. The most frequent bidirectional errors occurred between Normal and Premature Ventricular Contraction beats. These results support the use of interval-statistical features as a compact representation for multicategory ECG analysis, while showing that rare-class performance and fiducial reliability remain limiting factors for broader clinical use.</p>2026-09-08T00:00:00+07:00##submission.copyrightStatement##