Completed UT Austin postgraduate coursework using a supplied scenario, dataset, and starter notebook. The charts below come from my completed notebook.
INN Hotels
Classification, model evaluation, and decision trees
pandas, NumPy, statsmodels, scikit-learn, Seaborn, Matplotlib
The supplied data
The course supplied reservation records labeled canceled or not canceled. Fields included lead time, room price, special requests, repeat-guest history, market segment, arrival timing, adults and children, and prior cancellations. The scenario asked which booking patterns related to cancellations and how a classifier might flag higher-risk reservations.
What I did
- Explored distributions and relationships between booking attributes and cancellation status.
- Prepared categorical fields and compared a logistic-regression baseline with decision-tree models.
- Examined confusion matrices, precision, recall, and F1, including how tree pruning changed the balance between fit and interpretability.


Finding in the course report
I favored a pre-pruned tree for the exercise's balance of performance and interpretability. Lead time, special requests, and repeat-guest patterns informed possible service and cancellation-policy questions, not a deployed model or proven causal effects.
Learning reinforced
The project brought together exploratory analysis, a baseline model, nonlinear classification, and the tradeoff between overfitting and useful generalization. It also made precision and recall concrete business choices.
Original notebook export
View the complete HTML report, including code, outputs, and the original written analysis.