Diabetes Risk Prediction under Severe Class Imbalance: Comparing Firth Logistic Regression and SMOTEN-Based Logistic Regression
Abstract
Diabetes mellitus is a major noncommunicable disease, and reliable early-detection prediction can support Sustainable Development Goal (SDG) Target 3.4. This study compares standard logistic regression, Firth's Penalized Logistic Regression, and SMOTEN-based logistic regression for diabetes risk prediction under severe class imbalance using Wave 5 of the Indonesia Family Life Survey (IFLS5). The analysis included 115 respondents, comprising 8 diabetes cases (6.96%), with all predictors represented as categorical variables. Model performance was evaluated using 50 repeated stratified 75/25 train-test splits, with SMOTEN applied only to training data to prevent leakage. Standard logistic regression failed to converge across all partitions. Firth's Penalized Logistic Regression identified waist circumference as significantly associated with diabetes status (odds ratio = 10.69, 95% CI: 1.56–73.14) and achieved mean accuracy of 89.17%, sensitivity of 8.00%, specificity of 95.19%, and AUC of 0.692. SMOTEN-based logistic regression achieved accuracy of 86.90%, sensitivity of 12.00%, specificity of 92.44%, and AUC of 0.632. Brier scores were 0.089 and 0.098, respectively, compared with 0.064 for the prevalence benchmark, indicating poorer overall probabilistic prediction performance. These findings illustrate the challenges of rare-event prediction under severe class imbalance and highlight the need for larger samples and further validation before clinical application.
Keywords
Full Text:
PDFReferences
[1] J. Tanoey and H. Becher. “Diabetes prevalence and risk factors of early-onset adult diabetes: results from the Indonesian family life survey”. Global Health Action 14.1 (2021), p. 2001144. DOI: https://doi.org/10.1080/16549716.2021.2001144.
[2] F. R. Muharram, J. B. Swannjo, R. R. Melbiarta, and S. Martini. “Trends of diabetes and pre-diabetes in Indonesia 2013–2023: a serial analysis of national health surveys”. BMJ Open 15.9 (2025), e098575. DOI: https://doi.org/10.1136/bmjopen-2024-098575.
[3] M. Wahidin, A. Achadi, B. Besral, et al. “Projection of diabetes morbidity and mortality till 2045 in Indonesia based on risk factors and NCD prevention and control programs”. Scientific Reports 14 (2024), p. 5424. DOI: https://doi.org/10.1038/s41598-024-54563-2.
[4] United Nations. Goal 3: Ensure Healthy Lives and Promote Well-Being for All at All Ages. United Nations Sustainable Development Goals. 2026. Accessed August 19, 2026. URL: https://sdgs.un.org/goals/goal3.
[5] D. Ariyanto, A. Sofro, A. N. Hanifah, J. B. Prihanto, D. A. Maulana, and R. W. Romadhonia. “Logistic and probit regression modeling to predict the opportunities of diabetes in prospective athletes”. BAREKENG: Jurnal Ilmu Matematika dan Terapan 18.3 (2024), pp. 1391–1402. DOI: https://doi.org/10.30598/barekengvol18iss3pp1391-1402.
[6] T. Tamayo, C. Herder, and W. Rathmann. “Impact of early psychosocial factors (childhood socioeconomic factors and adversities) on future risk of type 2 diabetes, metabolic disturbances and obesity: a systematic review”. BMC Public Health 10 (2010), p. 525. DOI: https://doi.org/10.1186/1471-2458-10-525.
[7] H. He and E. A. Garcia. “Learning from imbalanced data”. IEEE Transactions on Knowledge and Data Engineering 21.9 (2009), pp. 1263–1284. DOI: https://doi.org/10.1109/TKDE.2008.239.
[8] D. Firth. “Bias reduction of maximum likelihood estimates”. Biometrika 80.1 (1993), pp. 27–38. DOI: https://doi.org/10.1093/biomet/80.1.27.
[9] S. Uno, H. Noma, and M. Gosho. “Firth-type penalized methods of the modified Poisson and least-squares regression analyses for binary outcomes”. Biometrical Journal 66.7 (2024), e202400004. DOI: https://doi.org/10.1002/bimj.202400004.
[10] S. Suhas, N. Manjunatha, C. N. Kumar, V. Benegal, G. N. Rao, M. Varghese, and G. Gururaj. “Firth’s penalized logistic regression: A superior approach for analysis of data from India’s National Mental Health Survey, 2016”. Indian Journal of Psychiatry 65.12 (2023), pp. 1208–1213. DOI: https://doi.org/10.4103/indianjpsychiatry.indianjpsychiatry_827_23.
[11] R. Puhr, G. Heinze, M. Nold, L. Lusa, and A. Geroldinger. “Firth’s logistic regression with rare events: accurate effect estimates and predictions?” Statistics in Medicine 36 (2017), pp. 2302–2317. DOI: https://doi.org/10.1002/sim.7273.
[12] S. Sadeghi, D. Khalili, A. Ramezankhani, M. A. Mansournia, and M. Parsaeian. “Diabetes mellitus risk prediction in the presence of class imbalance using flexible machine learning methods”. BMC Medical Informatics and Decision Making 22.1 (2022), p. 36. DOI: https://doi.org/10.1186/s12911-022-01775-z.
[13] M. Talebi Moghaddam, Y. Jahani, Z. Arefzadeh, A. Dehghan, M. Khaleghi, M. Sharafi, and G. Nikfar. “Predicting diabetes in adults: identifying important features in unbalanced data over a 5-year cohort study using machine learning algorithm”. BMC Medical Research Methodology 24 (2024), p. 220. DOI: https://doi.org/10.1186/s12874-024-02341-z.
[14] P. Sampath, G. Elangovan, K. Ravichandran, V. Shanmuganathan, S. Pasupathi, T. Chakrabarti, P. Chakrabarti, and M. Margala. “Robust diabetic prediction using ensemble machine learning models with synthetic minority over-sampling technique”. Scientific Reports 14 (2024), p. 28984. DOI: https://doi.org/10.1038/s41598-024-78519-8.
[15] Y. Jang. “Feature-based ensemble modeling for addressing diabetes data imbalance using the SMOTE, RUS, and random forest methods: a prediction study”. Ewha Medical Journal 48.2 (2025), e32. DOI: https://doi.org/10.12771/emj.2025.00353.
[16] R. van den Goorbergh, M. van Smeden, D. Timmerman, and B. Van Calster. “The harm of class imbalance corrections for risk prediction models: illustration and simulation using logistic regression”. Journal of the American Medical Informatics Association 29.9 (2022), pp. 1525–1534. DOI: https://doi.org/10.1093/jamia/ocac093.
[17] J. Strauss, F. Witoelar, and B. Sikoki. The Fifth Wave of the Indonesia Family Life Survey (IFLS5): Overview and Field Report. Tech. rep. WR-1143/1-NIA/NICHD. RAND Corporation, 2016. DOI: https://doi.org/10.7249/WR1143.1.
[18] N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer. “SMOTE: Synthetic Minority Over-sampling Technique”. Journal of Artificial Intelligence Research 16 (2002), pp. 321–357. DOI: https://doi.org/10.1613/jair.953.
[19] G. Heinze and M. Schemper. “A solution to the problem of separation in logistic regression”. Statistics in Medicine 21.16 (2002), pp. 2409–2419. DOI: https://doi.org/10.1002/sim.1047.
[20] P. Peduzzi, J. Concato, E. Kemper, T. R. Holford, and A. R. Feinstein. “A simulation study of the number of events per variable in logistic regression analysis”. Journal of Clinical Epidemiology 49.12 (1996), pp. 1373–1379. DOI: https://doi.org/10.1016/S0895-4356(96)00236-3.
[21] J. I. Ramírez-Manent, A. Martínez Jover, C. Silveira Martinez, P. Tomás-Gil, P. Martí-Lliteras, and A. A. López-González. “Waist circumference is an essential factor in predicting insulin resistance and early detection of metabolic syndrome in adults”. Nutrients 15.2 (2023), p. 257. DOI: https://doi.org/10.3390/nu15020257.
[22] Gary S. Collins et al. “TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods”. BMJ 385 (2024), e078378. DOI: https://doi.org/10.1136/bmj-2023-078378.
DOI: https://doi.org/10.18860/cauchy.v11i2.45736
Refbacks
- There are currently no refbacks.
Copyright (c) 2026 Lintang Dwi Laga Pertiwi, A'yunin Sofro

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Editorial Office
Mathematics Department,
Maulana Malik Ibrahim State Islamic University of Malang
Gajayana Street 50 Malang, East Java, Indonesia 65144
e-mail: cauchy@uin-malang.ac.id

CAUCHY: Jurnal Matematika Murni dan Aplikasi is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.







