According to Databricks best practices, Spark groupByKey should be avoided as Spark groupByKey processing works in a way that the information will be first shuffled across workers and then the processing will occur. Explanation
So, my question is, what are the alternatives for groupByKey in a way that it will return the following in a distributed and fast way?
// want this
{"key1": "1", "key1": "2", "key1": "3", "key2": "55", "key2": "66"}
// to become this
{"key1": ["1","2","3"], "key2": ["55","66"]}
Seems to me that maybe aggregateByKey or glom could do it first in the partition (map) and then join all the lists together (reduce).
groupByKeyis the most efficient choice (both time and storage) here. If it OOMs you simply need a larger cluster. - shuaiyuancn