Back to Projects
F
Data Science / Analysis

Fibroid Risk Prediction

ML Engineer (Research)

XGBoostLightGBMSMOTESHAPScikit-learnPython

About the Project

A research project developing a machine learning model to predict uterine fibroid presence from clinical and demographic data, backing an academic paper of the same name. The dataset is 49 anonymized patient records — 37 negative and 12 positive, a 3:1 imbalance on top of an already tiny sample. Raw features (age, height, weight, blood pressure, symptom text) are engineered into structured predictors, then classified with gradient boosting under 5-fold cross-validation with SHAP interpretability for clinical transparency.

Key Highlights

  • Achieved 94% mean accuracy (±4.9%) and 0.88 mean F1 under 5-fold cross-validation on gradient-boosted trees
  • Engineered structured predictors out of raw clinical fields — deriving body-composition and blood-pressure features, and parsing free-text symptoms into usable signals
  • Applied SMOTE oversampling and reported original-versus-augmented model comparisons side by side rather than only the augmented result
  • Implemented SHAP explainability so a clinician can see which features drove an individual prediction
  • Generated the paper's methodology, results, and figures reproducibly from the pipeline itself

Technical Challenges

The honest difficulty here is the sample size, and I'd rather state it than hide behind the headline accuracy. With 12 positive cases, a single fold contains two or three of them — so precision swings between 0.67 and 1.00 across folds (±16%) even while mean accuracy sits at 94%. That variance is the real result: the model is a credible signal that these clinical features carry predictive information, not a deployable diagnostic. Reporting the fold-level spread and the original-versus-augmented comparison, rather than the best number, is what keeps the finding defensible.