python禁止中国_Python中国的学习方式处理问题

最新推荐文章于 2023-12-10 11:27:54 发布

weixin_39755625

最新推荐文章于 2023-12-10 11:27:54 发布

阅读量76

点赞数

文章标签： python禁止中国

版权声明：本文为博主原创文章，遵循 CC 4.0 BY-SA 版权协议，转载请附上原文出处链接和本声明。

本文链接：https://blog.csdn.net/weixin_39755625/article/details/111847799

版权

a = '你们' 至 str 物

a = u'你们' 至 unicode 物

1.

>>> print 'u' + '你们'

>>> u欢

输出乱码

2.

>>> print 'u' + u'你'

>>> u你

正常

3.

>>> print 'u你'

>>> u浣

输出乱码

4.

>>> print 'u你' + 'u'

>>> u浣爑

输出乱码

5.

>>> print u'u你' + 'u'

>>> u你u

正常

6.

>>> print u'u你' + '你'

出现错误 UnicodeDecodeError: 'ascii' codec can't decode byte 0xe4 in position 0: ordinal not in range(128)

分析：'你'在内存中为 0xe4。而python默认的编码方案是ascii，ascii无法识别0xe4

7.

>>> print u'u你' + u'你'

>>> u你你

正常

8.

>>> print 'u你' + u'你'

出现错误 UnicodeDecodeError: 'ascii' codec can't decode byte 0xe4 in position 1: ordinal not in range(128)

9.

>>> print 'u你'.decode('utf-8') + u'你'

>>> u你你

正常

10.

而在处理由系统採集的含有中文的路径时，使用string.decode('utf-8')就不一定行了，由于中文简体的windows系统默认编码为gb2312，繁体中文版会採用Big5码

实验步骤例如以下：

file_from = sys.argv[1] 为由系统採集的包括中文的路径

file_to = file_from[:file_from.rfind('\\')+1].decode('utf-8') + u'你_' + file_from[file_from.rfind('\\')+1:].decode('utf-8')

print file_to

将出现错误：UnicodeDecodeError: 'utf8' codec can't decode byte 0xbb in position 24: invalid start byte

应该使用：decode('gb2312')

file_to = file_from[:file_from.rfind('\\')+1].decode('gb2312') + u'你_' + file_from[file_from.rfind('\\')+1:].decode('gb2312')

print file_to 正常

11.

而假设file_from是由你自己写入的包括中文的路径，如file_from = ‘c:\你.txt’

那么就应该用decode('utf-8')

能够參考上面的第7点和第9点

不足及错误之处，请批评指正！！谢谢！

。

參考文章：

版权声明：本文博主原创文章，博客，未经同意不得转载。

weixin_39755625

关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
python禁止中国_Python中国的学习方式处理问题

a = '你们'至 str 物a = u'你们'至 unicode 物1.>>> print 'u' + '你们'>>> u欢输出乱码2.>>> print 'u' + u'你'>>> u你正常3.>>> print 'u你'>>> u浣输出乱码4.>>> prin...
复制链接

扫一扫

评论

被折叠的条评论为什么被折叠?

到【灌水乐园】发言

查看更多评论

添加红包

成就一亿技术人!

hope_wisdom

发出的红包

实付元

使用余额支付

点击重新获取

扫码支付

钱包余额 0

抵扣说明：

1.余额是钱包充值的虚拟货币，按照1:1的比例进行支付金额的抵扣。
2.余额无法直接购买下载，可以购买VIP、付费专栏及课程。