Machine learning · 2023
Kidney Stone Prediction
Kaggle Playground Series S3E12 — binary classification

End-to-end ML pipeline predicting kidney stones from urine physiology, using feature engineering and a CatBoost ensemble validated with stratified k-fold.
- 0.756
- private ROC AUC
- CatBoost · XGB
- models compared
#Project overview
I used a Python-based Kaggle kernel to run a full data analysis and predictive modelling workflow for a machine learning competition. The goal was to predict a binary target from a set of physiological features using an ensemble algorithm, CatBoost.
#Tools and libraries
- Python for data manipulation and modelling
- Pandas & NumPy for data processing and linear algebra
- scikit-learn for scaling and model evaluation
- Seaborn & Matplotlib for visual insight into the data
- CatBoost, an ensemble algorithm that is robust on categorical data
#Data handling
- Loading: imported the train and test sets within the Kaggle environment and explored their size and shape.
- Exploration: studied distributions and relationships using correlation matrices and plots.
- Visualisation: heatmaps to explore correlations between features, and scatter plots to compare key feature pairs across conditions.
#Feature engineering
I created ratios and products of existing features to uncover interaction effects, giving the model more nuanced information than the raw features alone.
#Modelling
- Preprocessing: normalised the feature set so every feature carried equal weight.
- Training:
CatBoostRegressorwith stratified k-fold cross-validation to make predictions more robust and reduce overfitting. - Evaluation: ROC AUC, a single metric that captures how well the model separates the classes.
#Results
Training was refined iteratively to optimise ROC AUC across the validation folds, and I compared CatBoost against an XGBoost classifier. Final predictions on the test set were formatted into a competition submission, scoring 0.756 ROC AUC on the private leaderboard.
#What it shows
- Managing and preprocessing data efficiently with industry-standard Python libraries
- Using ensemble learning for robust predictive modelling
- Analysing and visualising data to uncover underlying patterns
- Optimising and evaluating models with cross-validation and the right metrics