OPTIMIZATION OF DIABETES PREDICTION USING MACHINE LEARNING MODELS: THE ROLE OF PREPROCESSING AND TEST SIZE VARIATIONS

Irfan Kurniawan, Ema Utami

Abstract


Diabetes mellitus remains one of the fastest-growing non-communicable diseases in the world, with global prevalence pro-jected to rise from 463 million cases in 2019 to 578 million by 2030 and 700 million by 2045. Timely and accurate risk prediction is therefore essential for early clinical intervention. This study exam-ines how data preprocessing strategies and train–test split ratios jointly affect the predictive accuracy of three widely used machine learning classifiers Random Forest, Support Vector Machine (SVM), and Naïve Bayes applied to the Pima Indians Diabetes Database (PIDD). Physiologically implausible zero values in the Glucose, Blood Pressure, Skin Thickness, Insulin, and BMI attributes were treated as missing and imputed using the median owing to the pres-ence of substantial outliers. Three scaling techniques (Min-Max, Standard, and Robust Scaling) were then compared against the un-scaled dataset across three train–test split configurations (0.1, 0.2, and 0.3). The experimental results show that Random Forest con-sistently achieved the highest accuracy across every preprocessing scenario, peaking at 89.61% with Min-Max Scaling and a 0.1 test size. SVM benefited considerably from feature scaling, reaching 87.01% accuracy under the same configuration, while Naïve Bayes performed best on the untransformed data (81.82%) and slightly declined after scaling, confirming its comparative insensitivity to feature magnitude. Further analysis indicates that, because tree-splitting criteria are theoretically invariant to monotonic feature transformation, most of the accuracy gain attributed to “prepro-cessing” in Random Forest is more plausibly explained by the medi-an-imputation step than by scaling itself a distinction not previously articulated in this line of research. These findings demonstrate that the benefit of preprocessing is algorithm-dependent rather than uni-versal, and that smaller test sizes generally favor tree-based ensem-bles by providing more training instances. The study contributes empirical evidence and a practical decision framework for selecting preprocessing pipelines according to classifier characteristics, ra-ther than assuming a one-size-fits-all approach in diabetes predic-tion research.

Keywords


Diabetes Prediction; Machine Learning; Random Forest; Support Vector Machine; Naïve Bayes; Data Preprocessing; Train–Test Split Ratio

Full Text:

PDF

References


Althobaiti, T., Althobaiti, S., & Selim, M. M. (2024). An optimized diabetes mellitus detection model for improved prediction of accuracy and clinical decision-making. Alexandria Engineering Journal, 94(October 2023), 311–324. https://doi.org/10.1016/j.aej.2024.03.044

Breiman, L. (2001). Random Forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/A:1010933404324

Carpinteiro, C., Lopes, J., Abelha, A., & Santos, M. F. (2023). A Comparative Study of Classification Algorithms for Early Detection of Diabetes. Procedia Computer Science, 220, 868–873. https://doi.org/10.1016/j.procs.2023.03.117

Chai, J., Zeng, H., Li, A., & Ngai, E. W. T. (2021). Deep learning in computer vision: A critical review of emerging techniques and application scenarios. Machine Learning with Applications, 6, 100134. https://doi.org/https://doi.org/10.1016/j.mlwa.2021.100134

Chaki, J., Thillai Ganesh, S., Cidham, S. K., & Ananda Theertan, S. (2022). Machine learning and artificial intelligence based Diabetes Mellitus detection and self-management: A systematic review. Journal of King Saud University - Computer and Information Sciences, 34(6), 3204–3225. https://doi.org/10.1016/j.jksuci.2020.06.013

Cho, N. H., Shaw, J. E., Karuranga, S., Huang, Y., da Rocha Fernandes, J. D., Ohlrogge, A. W., & Malanda, B. (2018). IDF Diabetes Atlas: Global estimates of diabetes prevalence for 2017 and projections for 2045. Diabetes Research and Clinical Practice, 138, 271–281. https://doi.org/10.1016/j.diabres.2018.02.023

Cover, T. M. (1965). Geometrical and Statistical Properties of Systems of Linear Inequalities with Applications in Pattern Recognition. IEEE Trans. Electron. Comput., 14, 326–334. https://api.semanticscholar.org/CorpusID:18251470

Diabetes. (n.d.). https://www.who.int/health-topics/diabetes

El Mestari, S. Z., Lenzini, G., & Demirci, H. (2024). Preserving data privacy in machine learning systems. Computers and Security, 137(November 2023), 103605. https://doi.org/10.1016/j.cose.2023.103605

Eshun, R. B., Bikdash, M., & Islam, A. K. M. K. (2024). A deep convolutional neural network for the classification of imbalanced breast cancer dataset. Healthcare Analytics, 5(June 2023), 100330. https://doi.org/10.1016/j.health.2024.100330

Fan, C., Chen, M., Wang, X., Wang, J., & Huang, B. (2021). A Review on Data Preprocessing Techniques Toward Efficient and Reliable Knowledge Discovery From Building Operational Data. Frontiers in Energy Research, 9(March), 1–17. https://doi.org/10.3389/fenrg.2021.652801

Febrian, M. E., Ferdinan, F. X., Sendani, G. P., Suryanigrum, K. M., & Yunanda, R. (2022). Diabetes prediction using supervised machine learning. Procedia Computer Science, 216(2022), 21–30. https://doi.org/10.1016/j.procs.2022.12.107

Ghassemi, M., Naumann, T., Schulam, P., Beam, A. L., Chen, I. Y., & Ranganath, R. (2020). A Review of Challenges and Opportunities in Machine Learning for Health. AMIA Joint Summits on Translational Science Proceedings. AMIA Joint Summits on Translational Science, 2020, 191–200.

Khanam, J. J., & Foo, S. Y. (2021). A comparison of machine learning algorithms for diabetes prediction. ICT Express, 7(4), 432–439. https://doi.org/10.1016/j.icte.2021.02.004

Mamoshina, P., Vieira, A., Putin, E., & Zhavoronkov, A. (2016). Applications of Deep Learning in Biomedicine. Molecular Pharmaceutics, 13(5), 1445–1454. https://doi.org/10.1021/acs.molpharmaceut.5b00982

Maniruzzaman, M., Rahman, M. J., Ahammed, B., & Abedin, M. M. (2020). Classification and prediction of diabetes disease using machine learning paradigm. Health Information Science and Systems, 8(1), 1–14. https://doi.org/10.1007/s13755-019-0095-z

Mansuroglu, R., Eckstein, T., Nützel, L., Wilkinson, S. A., & Hartmann, M. J. (2023). Variational Hamiltonian simulation for translational invariant systems via classical pre-processing. Quantum Science and Technology, 8(2), 0–17. https://doi.org/10.1088/2058-9565/acb1d0

Mehedi Hassan, M., Mollick, S., & Yasmin, F. (2022). An unsupervised cluster-based feature grouping model for early diabetes detection. Healthcare Analytics, 2(July), 100112. https://doi.org/10.1016/j.health.2022.100112

Olisah, C. C., Smith, L., & Smith, M. (2022). Diabetes mellitus prediction and diagnosis from a data preprocessing and machine learning perspective. Computer Methods and Programs in Biomedicine, 220, 106773. https://doi.org/10.1016/j.cmpb.2022.106773

Ong, K. L., Stafford, L. K., McLaughlin, S. A., Boyko, E. J., Vollset, S. E., Smith, A. E., Dalton, B. E., Duprey, J., Cruz, J. A., Hagins, H., Lindstedt, P. A., Aali, A., Abate, Y. H., Abate, M. D., Abbasian, M., Abbasi-Kangevari, Z., Abbasi-Kangevari, M., Abd ElHafeez, S., Abd-Rabu, R., … Vos, T. (2023). Global, regional, and national burden of diabetes from 1990 to 2021, with projections of prevalence to 2050: a systematic analysis for the Global Burden of Disease Study 2021. The Lancet, 402(10397), 203–234. https://doi.org/10.1016/S0140-6736(23)01301-6

Saeedi, P., Petersohn, I., Salpea, P., Malanda, B., Karuranga, S., Unwin, N., Colagiuri, S., Guariguata, L., Motala, A. A., Ogurtsova, K., Shaw, J. E., Bright, D., & Williams, R. (2019). Global and regional diabetes prevalence estimates for 2019 and projections for 2030 and 2045: Results from the International Diabetes Federation Diabetes Atlas, 9th edition. Diabetes Research and Clinical Practice, 157, 107843. https://doi.org/10.1016/j.diabres.2019.107843

Shamreen Ahamed, B., Arya, M. S., & Nancy, A. O. (2023). Diabetes Mellitus Disease Prediction Using Machine Learning Classifiers and Techniques Using the Concept of Data Augmentation and Sampling. Lecture Notes in Networks and Systems, 516, 401–413. https://doi.org/10.1007/978-981-19-5221-0_40

Smith, J. W., Everhart, J. E., Dickson, W. C., Knowler, W. C., & Johannes, R. S. (1988). Using the ADAP Learning Algorithm to Forecast the Onset of Diabetes Mellitus. In Proceedings of the Annual Symposium on Computer Application in Medical Care (pp. 261–265).

Smola, A. J., & Schölkopf, B. (1998). Learning with kernels (Vol. 4). Citeseer.

Surden, H. (2014). Machine learning and law. Washington Law Review, 89(1), 87–115.

Werner de Vargas, V., Schneider Aranda, J. A., dos Santos Costa, R., da Silva Pereira, P. R., & Victória Barbosa, J. L. (2023). Imbalanced data preprocessing techniques for machine learning: a systematic mapping study. Knowledge and Information Systems, 65(1), 31–57. https://doi.org/10.1007/s10115-022-01772-8

Zheng, M., Wang, F., Hu, X., Miao, Y., Cao, H., & Tang, M. (2022). A Method for Analyzing the Performance Impact of Imbalanced Binary Data on Machine Learning Models. Axioms, 11(11). https://doi.org/10.3390/axioms11110607




DOI: https://doi.org/10.29100/jipi.v10i1.7454

Refbacks

  • There are currently no refbacks.


Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

JIPI (Jurnal Ilmiah Penelitian dan Pembelajaran Informatika)
ISSN 2540-8984
Published by
Prodi Pendidikan Teknologi Informasi
Universitas Bhinneka PGRI

Website :https://jurnal.stkippgritulungagung.ac.id/index.php/jipi/index
Email: jipistkippti@gmail.com


Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.