I'm a spark newbie. I'm using pyspark for ALS recommendation. The fitting takes few minutes and runs fairly quickly. However the model.transform function takes a long time and requires significantly more nodes in the cluster.
- I was wondering if there are any optimization I can do to deal with the model.transform function?
- What is the method used underneath? Is it just simple matrix multiplication? If so can't I use another matrix multiplication library just for that?