I'm using from_json Pyspark SQL function as usual, e.g.:
>>> import pyspark.sql.types as t
>>> from pyspark.sql.functions import from_json
>>> df = sc.parallelize(['{"a":1}', '{"a":1, "b":2}', '{"a":1, "b":2, "c":3}']).toDF(t.StringType())
>>> df.show(3, False)
+---------------------+
|value |
+---------------------+
|{"a":1} |
|{"a":1, "b":2} |
|{"a":1, "b":2, "c":3}|
+---------------------+
>>> schema = t.StructType([t.StructField("a", t.IntegerType()), t.StructField("b", t.IntegerType()), t.StructField("c", t.IntegerType())])
>>> df.withColumn("json", from_json("value", schema)).show(3, False)
+---------------------+---------+
|value |json |
+---------------------+---------+
|{"a":1} |[1,,] |
|{"a":1, "b":2} |[1, 2,] |
|{"a":1, "b":2, "c":3}|[1, 2, 3]|
+---------------------+---------+
Please observe those keys not present in the JSON but specified in the schema have a parsed value of null (or some kind of empty value ?).
How can this be avoided? I mean, is there a way to set a default value to from_json? Or have I to add such a default value in a post-process of the dataframe?
Thanks!
[1,null,null]for the first row etc. so maybe you have some options set for the from_json which are different from default? - gawnullvalue to the list, a0.0value must be added. - frb