Clustering and Mixture Distribution Analysis of Average Years of Schooling in Papua (2010–2023)

Alvian Sroyer, Henderina Morin, Felix Reba, Jonathan Wororomi, Agustinus Languwuyo

Abstract


The purpose of this research is to analyze the distribution of the Human Development Index (HDI) in Papua based on the average years of schooling during the 2010–2023 period using the Gaussian Mixture-based Clustering approach. Data from 28 districts are grouped into five clusters according to their distributional characteristics. Each cluster is modeled using one of the four probability distributions: Inverse Gaussian, Rician, Weibull, or Nakagami. Parameter estimation was performed using the Maximum Likelihood Estimation (MLE) method, and the best distribution for each cluster was selected based on several information criteria (AIC, BIC, AICc, CAIC, and HQC) and validated through Kolmogorov-Smirnov (KS) and Anderson-Darling (AD) tests. The analysis results show that the Inverse Gaussian distribution fits Cluster 1 and Cluster 3, which represent districts with lower HDI schooling patterns. Cluster 2 is best described by the Rician distribution, indicating moderate HDI variability. The Weibull distribution fits Cluster 4, representing areas with moderately improving education. Cluster 5, with the highest and most stable HDI levels, is best modeled using the Nakagami distribution. The resulting mixture model, combining these four distributions, accurately reflects the HDI distribution patterns across Papua. Policy implications from this study include the development of cluster-based educational strategies tailored to regional characteristics to improve educational equity and human development across the province.


Keywords


Gaussian Mixture Model;Goodness of Fit; HDI Papua; Mixture Model; Probability Distribution;

Full Text:

PDF

References


[1] R. Karagiannis and G. Karagiannis. “Constructing composite indicators with Shannon entropy: The case of Human Development Index”. Socioeconomic Planning Sciences (2020). DOI: https://doi.org/10.1016/j.seps.2019.03.007.

[2] C. Türe and Y. Türe. “A model for the sustainability assessment based on the human development index in districts of Megacity Istanbul (Turkey)”. Environment, Development and Sustainability (2021). DOI: https://doi.org/10.1007/s10668-020-00735-9.

[3] Y. Jiang and C. Shi. “Estimating sustainability and regional inequalities using an enhanced sustainable development index in China”. Sustainable Cities and Society (2023). DOI: https://doi.org/10.1016/j.scs.2023.104555.

[4] S. Kwatra, A. Kumar, and P. Sharma. “A critical review of studies related to construction and computation of Sustainable Development Indices”. Ecological Indicators (2020). DOI: https://doi.org/10.1016/j.ecolind.2019.106061.

[5] W. Afalia, I. Hamda, S. A. Adriana, A. F. Alamsyah, and N. L. Wafiroh. “Determinants Of Human Development Index In Papua Province 2012-2021”. Wiga: Jurnal Penelitian Ilmu Ekonomi (2023). DOI: https://doi.org/10.30741/wiga.v13i2.1076.

[6] D. P. Rahmawati, I. N. Budiantara, D. D. Prastyo, and M. A. D. Octavanny. “Modeling of Human Development Index in Papua Province Using Spline Smoothing Estimator in Nonparametric Regression”. Journal of Physics: Conference Series (2021). DOI: https://doi.org/10.1088/1742-6596/1752/1/012018.

[7] D. A. N. Sirodj, I. M. Sumertajaya, and A. Kurnia. “Analisis Clustering Time Series untuk Pengelompokan Provinsi di Indonesia Berdasarkan Indeks Pembangunan Manusia Jenis Kelamin Perempuan”. Statistika: Journal of Theoretical Statistics and Its Applications (2023). DOI: https://doi.org/10.29313/statistika.v23i1.2181.

[8] A. M. Sikana and A. W. Wijayanto. “Analisis Perbandingan Pengelompokan Indeks Pembangunan Manusia Indonesia Tahun 2019 dengan Metode Partitioning dan Hierarchical Clustering”. Jurnal Ilmu Komputer (2021). DOI: https://doi.org/10.24843/jik.2021.v14.i02.p01.

[9] E. Luthfi and A. W. Wijayanto. “Analisis perbandingan metode hirearchical, k-means, dan k-medoids clustering dalam pengelompokkan indeks pembangunan manusia Indonesia”. INOVASI (2021). DOI: https://doi.org/10.30872/jinv.v17i4.10106.

[10] A.- Akramunnisa and F. Fajriani. “K-Means Clustering Analysis pada PersebaranTingkat Pengangguran Kabupaten/Kota di Sulawesi Selatan”. Jurnal Varian (2020). DOI: https://doi.org/10.30812/varian.v3i2.652.

[11] A. Budi and Samuel. “Klasterisasi Indeks Pembangunan Manusia (IPM) Per Kabupaten Di Indonesia Dengan Menggunakan Algoritma K-Means”. Jurnal Informatika dan Bisnis (2016).

[12] A. O. Lima, G. B. Lyra, M. C. Abreu, J. F. Oliveira-Júnior, M. Zeri, and G. Cunha-Zeri. “Extreme rainfall events over Rio de Janeiro State, Brazil: Characterization using probability distribution functions and clustering analysis”. Atmospheric Research 247.August 2020 (2021), p. 105221. DOI: https://doi.org/10.1016/j.atmosres.2020.105221.

[13] M. D. Doi, A. Rusgiyono, and T. Wuryandari. “Analisis k-Medoids dengan Validasi Indeks pada IPM Daerah 3T di Indonesia”. Jurnal Gaussian (2023). DOI: https://doi.org/10.14710/j.gauss.12.2.178-188.

[14] A. E. Ezugwu et al. “A comprehensive survey of clustering algorithms: State-of-the-art machine learning applications, taxonomy, challenges, and future research prospects”. Engineering Applications of Artificial Intelligence (2022). DOI: https://doi.org/10.1016/j.engappai.2022.104743.

[15] A. Rajabi, M. Eskandari, M. J. Ghadi, L. Li, J. Zhang, and P. Siano. “A comparative study of clustering techniques for electrical load pattern segmentation”. Renewable and Sustainable Energy Reviews (2020). DOI: https://doi.org/10.1016/j.rser.2019.109628.

[16] J. Li and J. Liu. “A modified extreme value perspective on best-performance life expectancy”. Journal of Population Research (2020). DOI: https://doi.org/10.1007/s12546-020-09248-8.

[17] M. S. Dhanoa et al. “A strategy for modelling heavy-tailed greenhouse gases (GHG) data using the generalised extreme value distribution: Are we overestimating GHG flux using the sample mean?” Atmospheric Environment (2020). DOI: https://doi.org/10.1016/j.atmosenv.2020.117500.

[18] H. Omrani, A. Alizadeh, and M. Amini. “A new approach based on BWM and MULTIMOORA methods for calculating semi-human development index: An application for provinces of Iran”. Socioeconomic Planning Sciences (2020). DOI: https://doi.org/10.1016/j.seps.2019.02.004.

[19] B. Giles-Corti, M. Lowe, and J. Arundel. “Achieving the SDGs: Evaluating indicators to be used to benchmark and monitor progress towards creating healthy and sustainable cities”. Health Policy (2020). DOI: https://doi.org/10.1016/j.healthpol.2019.03.001.

[20] S. A. Takyi, O. Amponsah, M. O. Asibey, and R. A. Ayambire. “An overview of Ghana’s educational system and its implication for educational equity”. International Journal of Leadership in Education (2021). DOI: https://doi.org/10.1080/13603124.2019.1613565.

[21] A. Sinha, D. Balsalobre-Lorente, M. W. Zafar, and M. M. Saleem. “Analyzing global inequality in access to energy: Developing policy framework by inequality decomposition”. Journal of Environmental Management (2022). DOI: https://doi.org/10.1016/j.jenvman.2021.114299.

[22] T. Ladi, A. Mahmoudpour, and A. Sharifi. “Assessing impacts of the water poverty index components on the human development index in Iran”. Habitat International (2021). DOI: https://doi.org/10.1016/j.habitatint.2021.102375.

[23] G. Resce. “Wealth-adjusted Human Development Index”. Journal of Cleaner Production (2021). DOI: https://doi.org/10.1016/j.jclepro.2021.128587.

[24] A. Yumashev, B. Ślusarczyk, S. Kondrashev, and A. Mikhaylov. “Global indicators of sustainable development: Evaluation of the influence of the human development index on consumption and quality of energy”. Energies (2020). DOI: https://doi.org/10.3390/en13112768.

[25] M. A. ul Haq, G. S. Rao, M. Albassam, and M. Aslam. “Marshall–Olkin Power Lomax distribution for modeling of wind speed data”. Energy Reports 6.May (2020), pp. 1118–1123. DOI: https://doi.org/10.1016/j.egyr.2020.04.033.

[26] S. Kageyama, N. Mori, S. Mugikura, H. Tokunaga, and K. Takase. “Gaussian mixture model-based cluster analysis of apparent diffusion coefficient values: a novel approach to evaluate uterine endometrioid carcinoma grade”. European Radiology (2021). DOI: https://doi.org/10.1007/s00330-020-07047-6.

[27] M. Krit, O. Gaudoin, M. Xie, and E. Remy. “Simplified likelihood based goodness-of-fit tests for the Weibull distribution”. Communications in Statistics - Simulation and Computation 45.3 (2016), pp. 920–951. DOI: https://doi.org/10.1080/03610918.2013.879889.

[28] M. H. Ouahabi, H. Elkhachine, F. Benabdelouahab, and A. Khamlichi. “Comparative study of five different methods of adjustment by the Weibull model to determine the most accurate method of analyzing annual variations of wind energy in Tetouan - Morocco”. Procedia Manufacturing 46.2019 (2020), pp. 698–707. DOI: https://doi.org/10.1016/j.promfg.2020.03.099.

[29] P. A. Costa Rocha, R. C. de Sousa, C. F. de Andrade, and M. E. V. da Silva. “Comparison of seven numerical methods for determining Weibull parameters for wind energy generation in the northeast region of Brazil”. Applied Energy 89.1 (2012), pp. 395–400. DOI: https://doi.org/10.1016/j.apenergy.2011.08.003.

[30] A. K. Azad, M. G. Rasul, M. M. Alam, S. M. Ameer Uddin, and S. K. Mondal. “Analysis of wind energy conversion system using Weibull distribution”. Procedia Engineering 90 (2014), pp. 725–732. DOI: https://doi.org/10.1016/j.proeng.2014.11.803.

[31] M. Nassar, A. Alzaatreh, M. Mead, and O. Abo-Kasem. “Alpha power Weibull distribution: Properties and applications”. Communications in Statistics - Theory and Methods 46.20 (2017), pp. 10236–10252. DOI: https://doi.org/10.1080/03610926.2016.1231816.

[32] M. Krit, O. Gaudoin, and E. Remy. “Goodness-of-fit tests for the Weibull and extreme value distributions: A review and comparative study”. Communications in Statistics - Simulation and Computation 50.7 (2021), pp. 1888–1911. DOI: https://doi.org/10.1080/03610918.2019.1594292.

[33] P. K. Chaurasiya, S. Ahmed, and V. Warudkar. “Study of different parameters estimation methods of Weibull distribution to determine wind power density using ground based Doppler SODAR instrument”. Alexandria Engineering Journal 57.4 (2018), pp. 2299–2311. DOI: https://doi.org/10.1016/j.aej.2017.08.008.

[34] A. M. Basheer. “Alpha power inverse Weibull distribution with reliability application”. Journal of Taibah University for Science 13.1 (2019), pp. 423–432. DOI: https://doi.org/10.1080/16583655.2019.1588488.

[35] A. O. Lima, G. B. Lyra, M. C. Abreu, J. F. Oliveira-Júnior, M. Zeri, and G. Cunha-Zeri. “Extreme rainfall events over Rio de Janeiro State, Brazil: Characterization using probability distribution functions and clustering analysis”. Atmospheric Research 247.March 2020 (2021), p. 105221. DOI: https://doi.org/10.1016/j.atmosres.2020.105221.

[36] A. Sulis, R. Cozza, and A. Annis. “Extreme wave analysis methods in the gulf of Cagliari (South Sardinia, Italy)”. Ocean & Coastal Management 140 (2017), pp. 79–87. DOI: https://doi.org/10.1016/j.ocecoaman.2017.02.023.

[37] A. O. Lima, G. B. Lyra, M. C. Abreu, J. F. Oliveira-Júnior, M. Zeri, and G. Cunha-Zeri. “Extreme rainfall events over Rio de Janeiro State, Brazil: Characterization using probability distribution functions and clustering analysis”. Atmospheric Research 247.July 2020 (2021), p. 105221. DOI: https://doi.org/10.1016/j.atmosres.2020.105221.

[38] Q. Han, S. Ma, T. Wang, and F. Chu. “Kernel density estimation model for wind speed probability distribution with applicability to wind energy assessment in China”. Renewable and Sustainable Energy Reviews 115.September (2019), p. 109387. DOI: https://doi.org/10.1016/j.rser.2019.109387.

[39] M. H. Samuh and A. M. Salhab. “Distribution of squared sum of products of independent Nakagami-m random variables”. Communications in Statistics - Simulation and Computation (2023). DOI: https://doi.org/10.1080/03610918.2023.2234668.

[40] K. Mohammadi, O. Alavi, and J. G. McGowan. “Use of Birnbaum-Saunders distribution for estimating wind speed and wind power probability distributions: A review”. Energy Conversion and Management 143 (2017), pp. 109–122. DOI: https://doi.org/10.1016/j.enconman.2017.03.083.

[41] K. S. Guedes, C. F. de Andrade, P. A. C. Rocha, R. dos S. Mangueira, and E. P. de Moura. “Performance analysis of metaheuristic optimization algorithms in estimating the parameters of several wind speed distributions”. Applied Energy 268.March (2020), p. 114952. DOI: https://doi.org/10.1016/j.apenergy.2020.114952.

[42] C. Jung and D. Schindler. “Global comparison of the goodness-of-fit of wind speed distributions”. Energy Conversion and Management 133 (2017), pp. 216–234. DOI: https://doi.org/10.1016/j.enconman.2016.12.006.

[43] Y. M. Kantar and I. Usta. “Analysis of the upper-truncated Weibull distribution for wind speed”. Energy Conversion and Management 96 (2015), pp. 81–88. DOI: https://doi.org/10.1016/j.enconman.2015.02.063.

[44] K. Mohammadi, O. Alavi, A. Mostafaeipour, N. Goudarzi, and M. Jalilvand. “Assessing different parameters estimation methods of Weibull distribution to compute wind power density”. Energy Conversion and Management 108 (2016), pp. 322–335. DOI: https://doi.org/10.1016/j.enconman.2015.11.015.

[45] T. Arslan, Y. M. Bulut, and A. Altin Yavuz. “Comparative study of numerical methods for determining Weibull parameters for wind energy potential”. Renewable and Sustainable Energy Reviews 40 (2014), pp. 820–825. DOI: https://doi.org/10.1016/j.rser.2014.08.009.

[46] A. K. Mbah and A. Paothong. “Shapiro–Francia test compared to other normality test using expected p-value”. Journal of Statistical Computation and Simulation 85.15 (2015), pp. 3002–3016. DOI: https://doi.org/10.1080/00949655.2014.947986.

[47] M. M. Badr. “Goodness-of-fit tests for the Compound Rayleigh distribution with application to real data”. Heliyon 5.8 (2019), p. e02225. DOI: https://doi.org/10.1016/j.heliyon.2019.e02225.

[48] B. Yazici and S. Yolacan. “A comparison of various tests of normality”. Journal of Statistical Computation and Simulation 77.2 (2007), pp. 175–183. DOI: https://doi.org/10.1080/10629360600678310.

[49] H. A. Bayoud. “Tests of normality: new test and comparative study”. Communications in Statistics - Simulation and Computation 50.12 (2021), pp. 4442–4463. DOI: https://doi.org/10.1080/03610918.2019.1643883.

[50] S. Dey, D. Kumar, P. L. Ramos, and F. Louzada. “Exponentiated Chen distribution: Properties and estimation”. Communications in Statistics - Simulation and Computation 46.10 (2017), pp. 8118–8139. DOI: https://doi.org/10.1080/03610918.2016.1267752.

[51] I. Pobočíková, Z. Sedliačková, and M. Michalková. “Application of Four Probability Distributions for Wind Speed Modeling”. Procedia Engineering 192 (2017), pp. 713–718. DOI: https://doi.org/10.1016/j.proeng.2017.06.123.

[52] K. Aprianto. “Optimasi Kernel K-Means dalam Pengelompokan Kabupaten/Kota Berdasarkan Indeks Pembangunan Manusia di Indonesia”. Limits: Journal of Mathematics and Its Applications (2018). DOI: https://doi.org/10.12962/limits.v15i1.3408.

[53] K. D. R. Sianipar and I. Gunawan. “Algoritma K-Means Dalam Pengelompokan Kabupaten/Kota Berdasarkan Indeks Pembengunan Manusia Di Sumatera Utara”. Jurnal Infomedia (2021). DOI: https://doi.org/10.30811/jim.v6i2.2426.

[54] H. E. Prastyo and F. Ilfana. “Pengelompokan kabupaten dan kota di jawa timur berdasarkan indeks pembangunan manusia dengan menggunakan metode k-means tahun 2020-2021”. Jurnal Ilmiah Komputasi dan Statistika (2022).




DOI: https://doi.org/10.18860/cauchy.v10i2.32988

Refbacks

  • There are currently no refbacks.


Copyright (c) 2025 Alvian Sroyer, Henderina Morin, Felix Reba, Jonathan Wororomi, Agustinus Languwuyo

Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

Editorial Office
Mathematics Department,
Maulana Malik Ibrahim State Islamic University of Malang
Gajayana Street 50 Malang, East Java, Indonesia 65144
e-mail: cauchy@uin-malang.ac.id

Creative Commons License
CAUCHY: Jurnal Matematika Murni dan Aplikasi is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.