python的代码保存到文档中打不开怎么办,Python无法打开UTF-8编码的文本文件

最新推荐文章于 2022-08-19 18:31:49 发布

weixin_39638801

最新推荐文章于 2022-08-19 18:31:49 发布

阅读量369

点赞数

文章标签： python的代码保存到文档中打不开怎么办

I have .py script which contains following code to open specific text file (which was generated by Exchange Powershell):

with codecs.open("C:\\Temp\\myfile.txt",encoding="utf_8",mode="r",errors="replace") as myfile:

content = myfile.readlines() #here we convert lines to list

print(content)

however, i tried also utf-16-be and utf-16-le (and standard ASCII obviously), but the file output is still looking like this (this is just part of it):

['��\r', '\x00\n', '\x00D\x00o\x00m\x00a\x00i\x00n\x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00 \x00\r', '\x00\n', '\x00-\x00-\x00-\x00-\x00-\x00-\x00

the file which i am trying to open is located here

does anybody please know what am i doing wrong? Is this some different kind of encoding?

解决方案

First, this text is definitely not UTF-8, so that's why Python can't open it as a UTF-8-encoded text file.

Second, you claim you "tried also utf-16-be and utf-16-le", but didn't show how you did that, and I suspect you did it wrong.

From the output, this is very likely BOM-encoded UTF-16-LE.

The first two bytes—because of the way you've printed them, we can't tell which bytes they are, but this is what it looks like when you print out \xFF and \xFE bytes. And the rest of the strings are a bunch of NUL even bytes alternating with reasonable-looking bytes, which almost always means UTF-16-LE. Plus, most common two-byte with a BOM in the wild is UTF-16-LE, and the fact that you're using all Microsoft tools makes that even more likely.

So, if you'd really tried utf-16-le, you would almost certainly have gotten the right string, but with an extra \ufeff at the start.

But of course the right answer is to just decode it as 'utf-16', which will consume and use the BOM properly.

weixin_39638801

关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
python的代码保存到文档中打不开怎么办,Python无法打开UTF-8编码的文本文件

I have .py script which contains following code to open specific text file (which was generated by Exchange Powershell):with codecs.open("C:\\Temp\\myfile.txt",encoding="utf_8",mode="r",errors="replac...
复制链接

扫一扫