I have a csv which I have previously read to a dataframe without issue, but now is giving me the following error: UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 0: invalid start byte
df = pd.read_csv(r'\\blah\blah2\csv.csv')
I tried this:
df = pd.read_csv(r'\\blah\blah2\csv.csv', encoding = 'utf-8-sig')
but that gave me this error: UnicodeDecodeError: 'utf-8-sig' codec can't decode byte 0xff in position 10423: invalid start byte
So then I tried 'utf-16', but that gave me this error: UnicodeError: UTF-16 stream does not start with BOM
Then I tried this:
with open(r'\\blah\blah2\csv.csv', 'rb') as f:
contents = f.read()
and that worked, but I need that csv as a dataframe, so then I tried:
new_df = pd.DataFrame.to_string(contents)
but I got this error: AttributeError: 'bytes' object has no attribute 'columns'
Could someone please help me get my dataframe?
Thank you.
UPDATE:
This fixed it. It read the csv into a dataframe without the unicode errors.
df = pd.read_csv(r'\\blah\blah2\csv.csv', encoding='latin1')
'rb'opening, update your question to report whatprint(contents[:100])displays. - Mark Tolonenlatin1will decode anything because it maps every byte to a Unicode code point, but not necessarily the correct code point. - Mark Tolonen