SOFTWARE ENGINEER & SDET

Jennifer Montgomery

Backend · full-stack · quality engineering

COURSEWORK / RENEWIND

Model tuning for imbalanced data

How should a model flag generator failures when a missed failure costs more than an extra inspection?

← Coursework on resume

Completed UT Austin postgraduate coursework using a supplied scenario, dataset, and starter notebook. The charts below come from my completed notebook.

MODEL TUNING

ReneWind

COURSE FOCUS

Imbalanced classification and hyperparameter tuning

LIBRARIES USED

pandas, NumPy, scikit-learn, imbalanced-learn, XGBoost, Seaborn, Matplotlib

The supplied data

The supplied predictive-maintenance records contained 40 anonymized sensor-derived predictors and a rare generator-failure label, with separate training and test data. The scenario gave different costs for missed failures, unnecessary inspections, and planned repairs, so accuracy alone was an incomplete measure.

What I did

  • Checked the class imbalance, prepared missing values, and compared original, oversampled, and undersampled training data.
  • Compared logistic regression, decision trees, and ensemble models with cross-validation.
  • Tuned candidate models with randomized search, assembled a final pipeline, and evaluated the selection on the supplied held-out test set.
Histogram of the wind-turbine failure label showing many normal cases and relatively few failures.
From the notebook: the rare failure class shaped the choice of metrics and resampling methods. Open chart ↗
Horizontal feature-importance bars for 40 anonymized sensor variables, with V30, V18, and V12 highest.
From the notebook: V30, V18, and V12 led the selected AdaBoost model, but anonymized names limit physical interpretation. Open chart ↗

Finding on the course test split

I selected AdaBoost trained with oversampling. It reached 0.851 recall and 0.774 precision on the held-out course test set, emphasizing detection of failures while accepting some false alerts.

Learning reinforced

This project made model selection depend on the cost of each error, not accuracy alone. It also applied resampling, cross-validation, hyperparameter search, and pipeline construction to an imbalanced problem.