python utf 8_用Python写入UTF-8文件

最新推荐文章于 2023-03-08 19:05:32 发布

weixin_39621819

最新推荐文章于 2023-03-08 19:05:32 发布

阅读量307

点赞数

文章标签： python utf 8

I'm really confused with the codecs.open function. When I do:

file = codecs.open("temp", "w", "utf-8")

file.write(codecs.BOM_UTF8)

file.close()

It gives me the error

UnicodeDecodeError: 'ascii' codec can't decode byte 0xef in position

0: ordinal not in range(128)

If I do:

file = open("temp", "w")

file.write(codecs.BOM_UTF8)

file.close()

It works fine.

Question is why does the first method fail? And how do I insert the bom?

If the second method is the correct way of doing it, what the point of using codecs.open(filename, "w", "utf-8")?

解决方案

I believe the problem is that codecs.BOM_UTF8 is a byte string, not a Unicode string. I suspect the file handler is trying to guess what you really mean based on "I'm meant to be writing Unicode as UTF-8-encoded text, but you've given me a byte string!"

Try writing the Unicode string for the byte order mark (i.e. Unicode U+FEFF) directly, so that the file just encodes that as UTF-8:

import codecs

file = codecs.open("lol", "w", "utf-8")

file.write(u'\ufeff')

file.close()

(That seems to give the right answer - a file with bytes EF BB BF.)

EDIT: S. Lott's suggestion of using "utf-8-sig" as the encoding is a better one than explicitly writing the BOM yourself, but I'll leave this answer here as it explains what was going wrong before.

weixin_39621819

关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
python utf 8_用Python写入UTF-8文件

I'm really confused with the codecs.open function. When I do:file = codecs.open("temp", "w", "utf-8")file.write(codecs.BOM_UTF8)file.close()It gives me the errorUnicodeDecodeError: 'ascii' codec can't...
复制链接

扫一扫