python - PySpark — UnicodeEncodeError: 'ascii' codec can't encode character

Question

Loading a dataframe with foreign characters (åäö) into Spark using spark.read.csv, with encoding='utf-8' and trying to do a simple show().

>>> df.show()

Traceback (most recent call last):
File "<stdin>", line 1, in <module>
File "/usr/lib/spark/python/pyspark/sql/dataframe.py", line 287, in show
print(self._jdf.showString(n, truncate))
UnicodeEncodeError: 'ascii' codec can't encode character u'\ufffd' in position 579: ordinal not in range(128)

I figure this is probably related to Python itself but I cannot understand how any of the tricks that are mentioned here for example can be applied in the context of PySpark and the show()-function.

@zero323 are there any other print-related commands that I could try? — salient
For starters you can try if df.rdd.map(lambda x: x).count() succeeds. — zero323
@zero323 – Yes, I have even successfully run some Spark SQL-queries — it's only this show()-function that fails on the encoding of the characters in strings. — salient
So rdd.take(20) for example executes without a problem? If so the problem may be a header. One way or another can you provide a minimal data sample which can be used to reproduce the problem? — zero323

Jussi Kujala Jussi Kujala · Accepted Answer · 2017-06-15T12:30:55

https://issues.apache.org/jira/browse/SPARK-11772 talks about this issue and gives a solution that runs:

export PYTHONIOENCODING=utf8

before running pyspark. I wonder why above works, because sys.getdefaultencoding() returned utf-8 for me even without it.

How to set sys.stdout encoding in Python 3? also talks about this and gives the following solution for Python 3:

import sys
sys.stdout = open(sys.stdout.fileno(), mode='w', encoding='utf8', buffering=1)

python - PySpark — UnicodeEncodeError: 'ascii' codec can't encode character

3 Answers