Machine Learning-Based Multi-Class Classification of Chronic Kidney Disease Stages Using XGBoost: A Feature-Driven Predictive Modelling Approach

  • Unique Paper ID: 198274
  • Volume: 12
  • Issue: 11
  • PageNo: 13423-13428
  • Abstract:
  • Background: Chronic kidney disease (CKD) is a major non-communicable disease burden affecting approximately 10–15% of the global adult population. Accurate and early classification of CKD progression stages is critical for timely clinical intervention and improved patient outcomes. Conventional staging relies on laboratory estimation of glomerular filtration rate (eGFR); however, machine learning methods offer an opportunity to classify multi-stage CKD from routinely collected clinical variables with high accuracy. Objective: This study proposes an XGBoost-based multi-class classification model to predict six CKD severity stages (G1–G5) using twelve clinically relevant biomarkers. The model's performance, generalisability, and interpretability are evaluated using standard machine learning metrics. Methods: A dataset of 10,000 patient records was used, covering CKD stages G1 (n=2,185), G2 (n=2,411), G3a (n=2,120), G3b (n=1,608), G4 (n=1,016), and G5 (n=660). Twelve features including serum creatinine, blood urea nitrogen (BUN), sex, haemoglobin, and blood pressure were selected for training. An XGBoost classifier was trained and evaluated using an 80/20 train-test split and validated with 5-fold cross-validation. Results: The model achieved a test accuracy of 91.20% and a 5-fold cross-validation accuracy of 91.12% ± 0.28%, indicating strong generalisation with minimal variance. Serum creatinine (importance score: ~0.42) and sex (~0.23) were identified as the two most influential features. Per-class F1-scores ranged from 0.84 (G4) to 0.96 (G1), with the model demonstrating consistently high performance across all six CKD stages. Conclusion: XGBoost provides a clinically viable, interpretable, and highly accurate framework for automated CKD staging from routine biochemical and demographic data. This approach has significant potential for integration into clinical decision support systems, particularly in settings where nephrology expertise is limited.

Copyright & License

Copyright © 2026 Authors retain the copyright of this article. This article is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

BibTeX

@article{198274,
        author = {Saarthak Dubey and Revanth Kumar and Ramarao and Suraj devamane and Dr Savitha G},
        title = {Machine Learning-Based Multi-Class Classification of Chronic Kidney Disease Stages Using XGBoost: A Feature-Driven Predictive Modelling Approach},
        journal = {International Journal of Innovative Research in Technology},
        year = {2026},
        volume = {12},
        number = {11},
        pages = {13423-13428},
        issn = {2349-6002},
        url = {https://ijirt.org/article?manuscript=198274},
        abstract = {Background: Chronic kidney disease (CKD) is a major non-communicable disease burden affecting approximately 10–15% of the global adult population. Accurate and early classification of CKD progression stages is critical for timely clinical intervention and improved patient outcomes. Conventional staging relies on laboratory estimation of glomerular filtration rate (eGFR); however, machine learning methods offer an opportunity to classify multi-stage CKD from routinely collected clinical variables with high accuracy.
Objective: This study proposes an XGBoost-based multi-class classification model to predict six CKD severity stages (G1–G5) using twelve clinically relevant biomarkers. The model's performance, generalisability, and interpretability are evaluated using standard machine learning metrics.
Methods: A dataset of 10,000 patient records was used, covering CKD stages G1 (n=2,185), G2 (n=2,411), G3a (n=2,120), G3b (n=1,608), G4 (n=1,016), and G5 (n=660). Twelve features including serum creatinine, blood urea nitrogen (BUN), sex, haemoglobin, and blood pressure were selected for training. An XGBoost classifier was trained and evaluated using an 80/20 train-test split and validated with 5-fold cross-validation.
Results: The model achieved a test accuracy of 91.20% and a 5-fold cross-validation accuracy of 91.12% ± 0.28%, indicating strong generalisation with minimal variance. Serum creatinine (importance score: ~0.42) and sex (~0.23) were identified as the two most influential features. Per-class F1-scores ranged from 0.84 (G4) to 0.96 (G1), with the model demonstrating consistently high performance across all six CKD stages.
Conclusion: XGBoost provides a clinically viable, interpretable, and highly accurate framework for automated CKD staging from routine biochemical and demographic data. This approach has significant potential for integration into clinical decision support systems, particularly in settings where nephrology expertise is limited.},
        keywords = {chronic kidney disease, XGBoost, machine learning, CKD staging, serum creatinine, eGFR, clinical prediction},
        month = {April},
        }

Cite This Article

Dubey, S., & Kumar, R., & Ramarao, , & devamane, S., & G, D. S. (2026). Machine Learning-Based Multi-Class Classification of Chronic Kidney Disease Stages Using XGBoost: A Feature-Driven Predictive Modelling Approach. International Journal of Innovative Research in Technology (IJIRT), 12(11), 13423–13428.

Related Articles