Symptom-Based Classification of Migraine Severity Using Random Forest and LassoNet
Abstract
This study aims to develop and compare classification models for predicting symptom-based migraine intensity using a public dataset from Kaggle. The research was conducted on multiclass data with an imbalanced class distribution. Therefore, target creation and result interpretation must be carefully designed so that the resulting evaluation remains consistent with the characteristics of the data used. In this study, the Intensity variable was recoded into three operational classes: low, moderate, and high. The features used included symptoms, characteristics of migraine episodes, and `symptom_count`, which represents the number of symptoms in each sample. The two models compared were Random Forest and LassoNet, both of which were tested using Stratified 5-Fold Cross-Validation. Model performance was assessed using the Macro F1-score as the primary metric, Balanced Accuracy as the main supplementary metric, and Accuracy as a complementary metric. The test results showed that Random Forest performed better, with a Macro F1-score of 0.6366, Balanced Accuracy of 0.6181, and Accuracy of 0.6225. Meanwhile, LassoNet achieved a Macro F1-score of 0.2474, a Balanced Accuracy of 0.3333, and an Accuracy of 0.5900. These results indicate that symptom patterns in the dataset can still be utilized to distinguish migraine intensity within a computational classification framework, although the separation between closely related classes is not yet fully robust.
Downloads
References
S. D. Silberstein, “Migraine,” MSD Manual Professional Edition. Accessed: Jun. 13, 2026. [Online]. Available: https://www.msdmanuals.com/professional/neurologic-disorders/headache/migraine
Y. W. Woldeamanuel and R. P. Cowan, “Computerized migraine diagnostic tools: a systematic review,” Jan. 24, 2022doi: 10.1177/20406223211065235.
W. Wallace et al., “The diagnostic and triage accuracy of digital and online symptom checker tools: a systematic review,” Dec. 01, 2022, Nature Research. doi: 10.1038/s41746-022-00667-w.
Y. You, R. Ma, and X. Gui, “User Experience of Symptom Checkers: A Systematic Review,” Aug. 2022. Accessed: Jun. 30, 2026. [Online]. Available: https://arxiv.org/abs/2208.09100
A. Reddy and A. Reddy, “Migraine triggers, phases, and classification using machine learning models,” Front. Neurol., vol. 16, 2025, doi: 10.3389/fneur.2025.1555215.
S. J. Kim, H. J. Lee, S. H. Lee, S. Cho, K. M. Kim, and M. K. Chu, “Most bothersome symptom in migraine and probable migraine: A population-based study,” PLoS One, vol. 18, no. 11, Nov. 2023, doi: 10.1371/journal.pone.0289729.
A. K. Eigenbrodt et al., “Diagnosis and management of migraine in ten steps,” Nat. Rev. Neurol., vol. 17, no. 8, pp. 501–514, Aug. 2021, doi: 10.1038/s41582-021-00509-5.
N. Imai and Y. Matsumori, “Different effects of migraine associated features on headache impact, pain intensity, and psychiatric conditions in patients with migraine,” Sci. Rep., vol. 14, no. 1, Dec. 2024, doi: 10.1038/s41598-024-74253-3.
L. Sasse et al., “Overview of leakage scenarios in supervised machine learning,” J. Big Data, vol. 12, no. 1, Dec. 2025, doi: 10.1186/s40537-025-01193-8.
J. Yang, R. El-Bouri, O. O’Donoghue, A. S. Lachapelle, A. A. S. Soltan, and D. A. Clifton, “Deep Reinforcement Learning for Multi-class Imbalanced Training,” May 2022. [Online]. Available: http://arxiv.org/abs/2205.12070
M. Owusu-Adjei, J. Ben Hayfron-Acquah, T. Frimpong, and G. Abdul-Salaam, “Imbalanced class distribution and performance evaluation metrics: A systematic review of prediction accuracy for determining model performance in healthcare systems,” PLOS Digital Health, vol. 2, no. 11 November, Nov. 2023, doi: 10.1371/journal.pdig.0000290.
S. Han, B. D. Williamson, and Y. Fong, “Improving random forest predictions in small datasets from two-phase sampling designs,” BMC Med. Inform. Decis. Mak., vol. 21, no. 1, Dec. 2021, doi: 10.1186/s12911-021-01688-3.
Aurélien. Géron, Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow, 3rd ed. Sebastopol, CA: O’Reilly Media, 2022. Accessed: Jun. 13, 2026. [Online]. Available: https://www.oreilly.com/library/view/hands-on-machine-learning/9781098125967/
I. Lemhadri, F. Ruan, L. Abraham, and R. Tibshirani, “LassoNet: A Neural Network with Feature Sparsity,” Journal of Machine Learning Research, vol. 22, no. 127, pp. 1–29, 2021, [Online]. Available: http://jmlr.org/papers/v22/20-848.html.
V. Borisov, T. Leemann, K. Seßler, J. Haug, M. Pawelczyk, and G. Kasneci, “Deep Neural Networks and Tabular Data: A Survey,” IEEE Trans. Neural Netw. Learn. Syst., vol. 35, no. 6, pp. 7499–7519, Jun. 2024, doi: 10.1109/TNNLS.2022.3229161.
M. Salmi, D. Atif, D. Oliva, A. Abraham, and S. Ventura, “Handling imbalanced medical datasets: review of a decade of research,” Artif. Intell. Rev., vol. 57, no. 10, Oct. 2024, doi: 10.1007/s10462-024-10884-2.
A. C. Muller and S. Guido, Introduction to Machine Learning with Python, 1st ed. Sebastopol, CA: O’Reilly Media, 2016. Accessed: Jun. 13, 2026. [Online]. Available: https://www.oreilly.com/library/view/introduction-to-machine/9781449369880/
Sebastian. Raschka and Vahid. Mirjalili, Python Machine Learning: Machine Learning and Deep Learning with Python, scikit-learn, and TensorFlow 2, 3rd ed. Packt Publishing, Limited, 2019. Accessed: Jun. 19, 2026. [Online]. Available: https://www.packtpub.com/en-us/product/python-machine-learning-third-edition-9781789955750
P. Bruce, Andrew. Bruce, and Peter. Gedeck, Practical statistics for data scientists : 50+ essential concepts using R and Python, 2nd ed. O’Reilly Media, 2020.
J. Pagán, J. L. Risco-Martín, J. M. Moya, and J. L. Ayala, “Modeling methodology for the accurate and prompt prediction of symptomatic events in chronic diseases,” J. Biomed. Inform., vol. 62, pp. 136–147, Feb. 2024, doi: 10.1016/j.jbi.2016.05.008.
K. Ghosh, C. Bellinger, R. Corizzo, P. Branco, B. Krawczyk, and N. Japkowicz, “The class imbalance problem in deep learning,” Mach. Learn., vol. 113, no. 7, pp. 4845–4901, Jul. 2024, doi: 10.1007/s10994-022-06268-8.
G. Holste et al., “Long-Tailed Classification of Thorax Diseases on Chest X-Ray: A New Benchmark Study,” in Data Augmentation, Labelling, and Imperfections, Cham: Springer Nature Switzerland, 2022, pp. 22–32. doi: 10.1007/978-3-031-17027-0_3.
V. M. Wagner et al., “Real-World Benchmarking and Validation of Foundation Model Transformers for Endometrial Cancer Subtyping from Histopathology,” Nature, Apr., 2026. doi: 10.1101/2025.10.10.25337691.
Bila bermanfaat silahkan share artikel ini
Berikan Komentar Anda terhadap artikel Symptom-Based Classification of Migraine Severity Using Random Forest and LassoNet
Pages: 635-642
Copyright (c) 2026 Syehan Fariz Gustomo, Kemas Muslim Lhaksmana

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under Creative Commons Attribution 4.0 International License that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (Refer to The Effect of Open Access).





















