Classification of Educational Vulnerability Status Using Boosting Methods on Imbalanced Socioeconomic Data

Andi Illa Erviani Nensi, Meavi Cintani, Mahda Al Maida, A. Qeis Tenridapi, Bagus Sartono, Aulia Rizki Firdawanti, Budi Susetyo, Gerry Alfa Dito

Abstract


Extreme class imbalance in educational data causes classification models to favor the majority class and fail to identify vulnerable students who are not attending school, despite this group requiring the greatest intervention. This study aims to develop a classification model for identifying educational vulnerability status among school-age children. The positive class corresponds to out-of-school children (minority class), while the negative class represents children who are still attending school. The dataset exhibits an extreme imbalance ratio of approximately 1:48, making conventional classification approaches ineffective in detecting minority-class observations.To address this challenge, four imbalance-aware boosting methods, namely AdaBoost-M2, SMOTEBoost, RBBoost, and RUSBoost, were compared. Model selection was conducted using stratified five-fold cross-validation on the training set, followed by final evaluation on an independent test set. Model performance was assessed using sensitivity, specificity, balanced accuracy, and Area Under the Curve (AUC), which are more appropriate than overall accuracy for highly imbalanced data. The results show that SMOTEBoost L50 with a threshold of 0.50 achieved the highest Balanced Accuracy (0.7474), Sensitivity (0.8824), and AUC (0.7745). However, its low minority-class precision (0.0452) and F1-score (0.0860) indicate a substantial false-positive trade-off. These findings demonstrate that integrating minority-class balancing strategies with boosting mechanisms can substantially improve the detection of vulnerable students in highly imbalanced socioeconomic data. Feature importance analysis revealed that education level or class was the most dominant predictor, followed by household size and frequency of internet use. Overall, this study provides empirical evidence regarding the effectiveness of imbalance-aware boosting methods for educational vulnerability detection and highlights their potential application in supporting more targeted educational interventions.

Keywords


boosting; class imbalance; educational vulnerability; SMOTEBoost; out-of-school student detection

Full Text:

PDF

References


[1] Republic of Indonesia, Law of the Republic of Indonesia Number 20 of 2003 concerning the National Education System. Jakarta, 2003. URL: https://peraturan.bpk.go.id/Home/Details/43920/uu-no-20-tahun. Accessed: 9 Aug. 2026.

[2] Badan Pusat Statistik, Statistik Pendidikan 2023. Jakarta, 2023. URL: https://www.bps.go.id/id/publication/2023/11/24/54557f7c1bd32f187f3cdab5/statistik-pendidikan-2023.html. Accessed: 9 Aug. 2026.

[3] Suharti, “Educational inequality and household socioeconomic status in Indonesia,” Journal of Indonesian Economy and Business, vol. 36, no. 2, pp. 123–140, 2021.

[4] N. F. Azzahra, Addressing Distance Learning Barriers in Indonesia amid the COVID-19 Pandemic. Jakarta, 2020. DOI: https://doi.org/10.35497/309162. URL: https://hdl.handle.net/10419/249436. Accessed: 9 Aug. 2026.

[5] United Nations Development Programme, Human Development Report 2019: Beyond Income, Beyond Averages, Beyond Today. New York: United Nations Development Programme, 2019. URL: https://hdr.undp.org/content/human-development-report-2019. Accessed: 9 Aug. 2026.

[6] A. I. E. Nensi, D. Gustiara, S. Shafa, and B. Susetyo, “Analysis of the relationship between literacy, numeracy and school accreditation rankings in Sulawesi using ordinal logistic regression and K-nearest neighbors,” Journal of Mathematics, Computations and Statistics, vol. 9, no. 2, pp. 280–293, June 2026. DOI: https://doi.org/10.35580/Jmathcos11260.

[7] H. He and E. A. Garcia, “Learning from imbalanced data,” IEEE Transactions on Knowledge and Data Engineering, vol. 21, no. 9, pp. 1263–1284, 2009. DOI: https://doi.org/10.1109/TKDE.2008.239.

[8] B. Krawczyk, “Learning from imbalanced data: Open challenges and future directions,” Progress in Artificial Intelligence, vol. 5, no. 4, pp. 221–232, 2016. DOI: https://doi.org/10.1007/s13748-016-0094-0.

[9] D. Micci-Barreca, “A preprocessing scheme for high-cardinality categorical attributes in classification and prediction problems,” ACM SIGKDD Explorations Newsletter, vol. 3, no. 1, pp. 27–32, 2001. DOI: https://doi.org/10.1145/507533.507538.

[10] M. Kuhn and K. Johnson, Applied Predictive Modeling. New York: Springer, 2013. DOI: https://doi.org/10.1007/978-1-4614-6849-3.

[11] C. Ambroise and G. J. McLachlan, “Selection bias in gene extraction on the basis of microarray gene-expression data,” Proceedings of the National Academy of Sciences, vol. 99, no. 10, pp. 6562–6566, 2002. DOI: https://doi.org/10.1073/pnas.102102699.

[12] G. C. Cawley and N. L. C. Talbot, “On over-fitting in model selection and subsequent selection bias in performance evaluation,” Journal of Machine Learning Research, vol. 11, no. 70, pp. 2079–2107, 2010. URL: https://www.jmlr.org/papers/v11/cawley10a.html.

[13] D. Krstajic, L. J. Buturovic, D. E. Leahy, and S. Thomas, “Cross-validation pitfalls when selecting and assessing regression and classification models,” Journal of Cheminformatics, vol. 6, no. 1, p. 10, 2014. DOI: https://doi.org/10.1186/1758-2946-6-10.

[14] N. V. Chawla, A. Lazarevic, L. O. Hall, and K. W. Bowyer, “SMOTEBoost: Improving prediction of the minority class in boosting,” in Proceedings of the 7th European Conference on Principles and Practice of Knowledge Discovery in Databases (PKDD), Cavtat-Dubrovnik, Croatia, 2003, pp. 107–119. DOI: https://doi.org/10.1007/978-3-540-39804-2_12.

[15] Y. Freund and R. E. Schapire, “A decision-theoretic generalization of on-line learning and an application to boosting,” Journal of Computer and System Sciences, vol. 55, no. 1, pp. 119–139, 1997. DOI: https://doi.org/10.1006/jcss.1997.1504.

[16] Y. Sun, M. S. Kamel, A. K. C. Wong, and Y. Wang, “Cost-sensitive boosting for classification of imbalanced data,” Pattern Recognition, vol. 40, no. 12, pp. 3358–3378, 2007. DOI: https://doi.org/10.1016/j.patcog.2007.04.009.

[17] C. Seiffert, T. M. Khoshgoftaar, J. Van Hulse, and A. Napolitano, “RUSBoost: A hybrid approach to alleviating class imbalance,” IEEE Transactions on Systems, Man, and Cybernetics – Part A: Systems and Humans, vol. 40, no. 1, pp. 185–197, 2010. DOI: https://doi.org/10.1109/TSMCA.2009.2029559.

[18] R. Gao and Z. Liu, “An improved AdaBoost algorithm for hyperparameter optimization,” Journal of Physics: Conference Series, vol. 1631, no. 1, p. 012048, 2020. DOI: https://doi.org/10.1088/1742-6596/1631/1/012048.

[19] T. R. Mahesh, V. V. Kumar, V. D. Kumar, O. Geman, M. Margala, and M. Guduri, “The stratified K-folds cross-validation and class-balancing methods with high-performance ensemble classifiers for breast cancer classification,” Healthcare Analytics, vol. 4, p. 100247, 2023. DOI: https://doi.org/10.1016/j.health.2023.100247.

[20] T. Fawcett, “An introduction to ROC analysis,” Pattern Recognition Letters, vol. 27, pp. 861–874, 2006. DOI: https://doi.org/10.1016/j.patrec.2005.10.010.




DOI: https://doi.org/10.18860/cauchy.v11i2.39670

Refbacks

  • There are currently no refbacks.


Copyright (c) 2026 Andi Illa Erviani Nensi

Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

Editorial Office
Mathematics Department,
Maulana Malik Ibrahim State Islamic University of Malang
Gajayana Street 50 Malang, East Java, Indonesia 65144
e-mail: cauchy@uin-malang.ac.id

Creative Commons License
CAUCHY: Jurnal Matematika Murni dan Aplikasi is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.