2
votes

I work at an ecommerce company and I'm responsible for clustering our customers based on their transactional behavior. I've never worked with clustering before, so I'm having a bit of a rough time.

1st) I've gathered data on customers and I've chosen 12 variables that specify very nicely how these customers behave. Each line of the dataset represents 1 user, where the columns are the 12 features I've chosen.

2nd) I've removed some outliers and built a correlation matrix in order to check of redundant variables. Turns out some of them are highly correlated ( > 0.8 correlation)

3rd) I used sklearn's RobustScaler on all 12 variables in order to make sure the variable's variability doesn't change much (StandardScaler did a poor job with my silhouette)

4th) I ran KMeans on the dataset and got a very good result for 2 clusters (silhouette of >70%)

5th) I tried doing a PCA after scaling / before clustering to reduce my dimension from 12 to 2 and, to my surprise, my silhouette started going to 30~40% and, when I plot the datapoints, it's just a big mass at the center of the graph.

My question is:

1) What's the difference between RobustScaler and StandardScaler on sklearn? When should I use each?

2) Should I do : Raw Data -> Cleaned Data -> Normalization -> PCA/TSNE -> Clustering ? Or Should PCA come before normalization?

3) Is a 12 -> 2 dimension reduction through PCA too extreme? That might be causing the horrible silhouette score.

Thank you very much!

1

1 Answers

0
votes

Avoid comparing Silhouettes of different projections or scalings. Internal measures tend to be too sensitive.

Do not use tSNE for clustering (Google for the discussion on stats.SE, feel free to edit the link into this answer). It will cause false separation and false adjacency; it is a visualization technique.

PCA will scale down high variance axes, and scale up low variance directions. It is to be expected that this overall decreases the quality if the main axis is what you are interested in (and it is expected to help if it is not). But if PCA visualization shows only one big blob, then a Silhouette of 0.7 should not be possible. For such a high silhouette, the clusters should be separable in the PCA view.