
Customer Segmentation Using RFM Analysis and K-Means Clustering
Groups customers by recency, frequency and spend, so a marketing team can target retention and re-engagement instead of treating every customer the same.
Built on the UCI Online Retail dataset - about 540,000 transactions from a UK-based online retailer over a year. Each customer is scored on Recency, Frequency and Monetary value (RFM), first with a rule-based scoring system, then with K-Means clustering on the scaled RFM values.
The number of clusters (k=4) was chosen using the Elbow and Silhouette methods, and each resulting cluster is labeled by its average spend rather than a fixed index, so the labels (VIP, Loyal High-Spender, Mid-Value, At Risk) stay meaningful even if K-Means assigns cluster numbers differently on a future run.
The pipeline is modular and covered by pytest, and a saved model can score new customers without retraining.
- Python
- pandas
- scikit-learn
- matplotlib
- seaborn
- pytest

