Machine-learning classification of Chronic Kidney Disease on the UCI CKD dataset (400 patients), evaluated with leak-free 5-fold stratified cross-validation.
- Go to colab.research.google.com
File ▸ Upload notebook→ uploadnotebooks/CKD_Classification_Colab.ipynb- Run the first cell and upload
data/kidney_disease.csvwhen prompted Runtime ▸ Run all
Colab already has pandas, numpy, scikit-learn, matplotlib and seaborn — nothing to install.
pip install pandas numpy scikit-learn matplotlib seaborn
python src/ckd_pipeline.py --experiment no_outlier # Experiment 1 (primary)
python src/ckd_pipeline.py --experiment outlier_capped # Experiment 2 (comparison)
PYTHONPATH=src python src/leakage_demo.py # leakage check
python src/ckd_pipeline.py --mode normal # single train/test split only
python src/ckd_pipeline.py --mode kfold # 5-fold cross-validation onlyBoth classification modes run by default (--mode both):
- normal — one stratified 80/20 train/test split. One confusion matrix, one ROC curve and one set of metrics per model, on 80 patients the model has never seen.
- kfold — 5-fold stratified cross-validation: mean accuracy across the folds, 5 ROC curves per model (one per fold plus their mean), and a pooled out-of-fold confusion matrix covering all 400 patients.
data/kidney_disease.csv raw dataset (400 rows)
notebooks/CKD_Classification_Colab.ipynb step-by-step teaching notebook (53 cells)
src/ckd_pipeline.py reproducible analysis script
src/leakage_demo.py measures leaky vs correct preprocessing
results/no_outlier/ Experiment 1 tables (CSV) + figures (PNG)
table_normal_classification.csv holdout metrics per model
table_model_summary.csv pooled out-of-fold + mean-accuracy columns
results/outlier_capped/ Experiment 2 tables + figures
results/table_experiment_comparison.csv side-by-side comparison
REPORT.md methodology, results, discussion, conclusion
| Model | Accuracy | Sensitivity | Specificity | Precision | F1 | AUC |
|---|---|---|---|---|---|---|
| Logistic Regression | 0.9875 | 0.9800 | 1.0000 | 1.0000 | 0.9899 | 1.0000 |
| Decision Tree | 0.9750 | 0.9800 | 0.9667 | 0.9800 | 0.9800 | 0.9733 |
| Random Forest | 1.0000 | 1.0000 | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| SVM (RBF) | 0.9750 | 0.9600 | 1.0000 | 1.0000 | 0.9796 | 1.0000 |
On 80 test patients one misclassification is worth 0.0125 accuracy, so the whole spread above is two patients. The cross-validated numbers below use all 400 and are the ones to quote.
| Model | Accuracy | Sensitivity | Specificity | Precision | F1 | AUC |
|---|---|---|---|---|---|---|
| Logistic Regression | 0.9925 | 0.9920 | 0.9933 | 0.9960 | 0.9940 | 0.9998 |
| Decision Tree | 0.9800 | 0.9880 | 0.9667 | 0.9802 | 0.9841 | 0.9773 |
| Random Forest | 0.9900 | 0.9960 | 0.9800 | 0.9881 | 0.9920 | 0.9997 |
| SVM (RBF) | 0.9950 | 0.9920 | 1.0000 | 1.0000 | 0.9960 | 1.0000 |
Positive class = CKD. Random Forest missed only 1 of 250 CKD patients (FN = 1) and is recommended for screening; SVM leads on accuracy and specificity.
Caveat: the top predictors are the defining diagnostic criteria of CKD, so this
benchmark is near-separable. These scores would not transfer to raw hospital data.
See REPORT.md for the full discussion.
Imputation, scaling, one-hot encoding and (in Experiment 2) IQR fences all live inside a
sklearn.pipeline.Pipeline, which is refitted from scratch inside every training fold.
Held-out folds are only .transform()-ed, never .fit()-ted on.