Symptom-Based Classification of Migraine Severity Using Random Forest and LassoNet


  • Syehan Fariz Gustomo Telkom University, Bandung, Indonesia
  • Kemas Muslim Lhaksmana * Mail Telkom University, Bandung, Indonesia
  • (*) Corresponding Author
Keywords: Balanced Accuracy; LassoNet; Migraine Intensity; Multiclass Classification; Random Forest

Abstract

This study aims to develop and compare classification models for predicting symptom-based migraine intensity using a public dataset from Kaggle. The research was conducted on multiclass data with an imbalanced class distribution. Therefore, target creation and result interpretation must be carefully designed so that the resulting evaluation remains consistent with the characteristics of the data used. In this study, the Intensity variable was recoded into three operational classes: low, moderate, and high. The features used included symptoms, characteristics of migraine episodes, and `symptom_count`, which represents the number of symptoms in each sample. The two models compared were Random Forest and LassoNet, both of which were tested using Stratified 5-Fold Cross-Validation. Model performance was assessed using the Macro F1-score as the primary metric, Balanced Accuracy as the main supplementary metric, and Accuracy as a complementary metric. The test results showed that Random Forest performed better, with a Macro F1-score of 0.6366, Balanced Accuracy of 0.6181, and Accuracy of 0.6225. Meanwhile, LassoNet achieved a Macro F1-score of 0.2474, a Balanced Accuracy of 0.3333, and an Accuracy of 0.5900. These results indicate that symptom patterns in the dataset can still be utilized to distinguish migraine intensity within a computational classification framework, although the separation between closely related classes is not yet fully robust.

Downloads

Download data is not yet available.

References

S. D. Silberstein, “Migraine,” MSD Manual Professional Edition. Accessed: Jun. 13, 2026. [Online]. Available: https://www.msdmanuals.com/professional/neurologic-disorders/headache/migraine

Y. W. Woldeamanuel and R. P. Cowan, “Computerized migraine diagnostic tools: a systematic review,” Jan. 24, 2022doi: 10.1177/20406223211065235.

W. Wallace et al., “The diagnostic and triage accuracy of digital and online symptom checker tools: a systematic review,” Dec. 01, 2022, Nature Research. doi: 10.1038/s41746-022-00667-w.

Y. You, R. Ma, and X. Gui, “User Experience of Symptom Checkers: A Systematic Review,” Aug. 2022. Accessed: Jun. 30, 2026. [Online]. Available: https://arxiv.org/abs/2208.09100

A. Reddy and A. Reddy, “Migraine triggers, phases, and classification using machine learning models,” Front. Neurol., vol. 16, 2025, doi: 10.3389/fneur.2025.1555215.

S. J. Kim, H. J. Lee, S. H. Lee, S. Cho, K. M. Kim, and M. K. Chu, “Most bothersome symptom in migraine and probable migraine: A population-based study,” PLoS One, vol. 18, no. 11, Nov. 2023, doi: 10.1371/journal.pone.0289729.

A. K. Eigenbrodt et al., “Diagnosis and management of migraine in ten steps,” Nat. Rev. Neurol., vol. 17, no. 8, pp. 501–514, Aug. 2021, doi: 10.1038/s41582-021-00509-5.

N. Imai and Y. Matsumori, “Different effects of migraine associated features on headache impact, pain intensity, and psychiatric conditions in patients with migraine,” Sci. Rep., vol. 14, no. 1, Dec. 2024, doi: 10.1038/s41598-024-74253-3.

L. Sasse et al., “Overview of leakage scenarios in supervised machine learning,” J. Big Data, vol. 12, no. 1, Dec. 2025, doi: 10.1186/s40537-025-01193-8.

J. Yang, R. El-Bouri, O. O’Donoghue, A. S. Lachapelle, A. A. S. Soltan, and D. A. Clifton, “Deep Reinforcement Learning for Multi-class Imbalanced Training,” May 2022. [Online]. Available: http://arxiv.org/abs/2205.12070

M. Owusu-Adjei, J. Ben Hayfron-Acquah, T. Frimpong, and G. Abdul-Salaam, “Imbalanced class distribution and performance evaluation metrics: A systematic review of prediction accuracy for determining model performance in healthcare systems,” PLOS Digital Health, vol. 2, no. 11 November, Nov. 2023, doi: 10.1371/journal.pdig.0000290.

S. Han, B. D. Williamson, and Y. Fong, “Improving random forest predictions in small datasets from two-phase sampling designs,” BMC Med. Inform. Decis. Mak., vol. 21, no. 1, Dec. 2021, doi: 10.1186/s12911-021-01688-3.

Aurélien. Géron, Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow, 3rd ed. Sebastopol, CA: O’Reilly Media, 2022. Accessed: Jun. 13, 2026. [Online]. Available: https://www.oreilly.com/library/view/hands-on-machine-learning/9781098125967/

I. Lemhadri, F. Ruan, L. Abraham, and R. Tibshirani, “LassoNet: A Neural Network with Feature Sparsity,” Journal of Machine Learning Research, vol. 22, no. 127, pp. 1–29, 2021, [Online]. Available: http://jmlr.org/papers/v22/20-848.html.

V. Borisov, T. Leemann, K. Seßler, J. Haug, M. Pawelczyk, and G. Kasneci, “Deep Neural Networks and Tabular Data: A Survey,” IEEE Trans. Neural Netw. Learn. Syst., vol. 35, no. 6, pp. 7499–7519, Jun. 2024, doi: 10.1109/TNNLS.2022.3229161.

M. Salmi, D. Atif, D. Oliva, A. Abraham, and S. Ventura, “Handling imbalanced medical datasets: review of a decade of research,” Artif. Intell. Rev., vol. 57, no. 10, Oct. 2024, doi: 10.1007/s10462-024-10884-2.

A. C. Muller and S. Guido, Introduction to Machine Learning with Python, 1st ed. Sebastopol, CA: O’Reilly Media, 2016. Accessed: Jun. 13, 2026. [Online]. Available: https://www.oreilly.com/library/view/introduction-to-machine/9781449369880/

Sebastian. Raschka and Vahid. Mirjalili, Python Machine Learning: Machine Learning and Deep Learning with Python, scikit-learn, and TensorFlow 2, 3rd ed. Packt Publishing, Limited, 2019. Accessed: Jun. 19, 2026. [Online]. Available: https://www.packtpub.com/en-us/product/python-machine-learning-third-edition-9781789955750

P. Bruce, Andrew. Bruce, and Peter. Gedeck, Practical statistics for data scientists : 50+ essential concepts using R and Python, 2nd ed. O’Reilly Media, 2020.

J. Pagán, J. L. Risco-Martín, J. M. Moya, and J. L. Ayala, “Modeling methodology for the accurate and prompt prediction of symptomatic events in chronic diseases,” J. Biomed. Inform., vol. 62, pp. 136–147, Feb. 2024, doi: 10.1016/j.jbi.2016.05.008.

K. Ghosh, C. Bellinger, R. Corizzo, P. Branco, B. Krawczyk, and N. Japkowicz, “The class imbalance problem in deep learning,” Mach. Learn., vol. 113, no. 7, pp. 4845–4901, Jul. 2024, doi: 10.1007/s10994-022-06268-8.

G. Holste et al., “Long-Tailed Classification of Thorax Diseases on Chest X-Ray: A New Benchmark Study,” in Data Augmentation, Labelling, and Imperfections, Cham: Springer Nature Switzerland, 2022, pp. 22–32. doi: 10.1007/978-3-031-17027-0_3.

V. M. Wagner et al., “Real-World Benchmarking and Validation of Foundation Model Transformers for Endometrial Cancer Subtyping from Histopathology,” Nature, Apr., 2026. doi: 10.1101/2025.10.10.25337691.


Bila bermanfaat silahkan share artikel ini

Berikan Komentar Anda terhadap artikel Symptom-Based Classification of Migraine Severity Using Random Forest and LassoNet

Dimensions Badge
Article History
Submitted: 2026-05-21
Published: 2026-06-30
Abstract View: 0 times
PDF Download: 0 times
How to Cite
Gustomo, S. F., & Lhaksmana, K. (2026). Symptom-Based Classification of Migraine Severity Using Random Forest and LassoNet. Building of Informatics, Technology and Science (BITS), 8(1), 635-642. https://doi.org/10.47065/bits.v8i1.9900
Issue
Section
Articles

Most read articles by the same author(s)

1 2 > >>