python统计中文字数_【Python进阶】用 Python 统计字数

最新推荐文章于 2022-12-26 19:16:40 发布

weixin_39972519

最新推荐文章于 2022-12-26 19:16:40 发布

阅读量150

点赞数

文章标签： python统计中文字数

解决问题的思路：

1. 将字符串s进行空白符分割得到所有的单词列表split_s，如：['betty', 'bought', 'a', 'bit', 'of', 'butter', 'but', 'the', 'butter', 'was', 'bitter']

2. 建立maplist,将split_s转化为元素为元组的列表形式，如：[('betty', 1), ('bought', 1), ('a', 1), ('bit', 1), ('of', 1), ('butter', 1), ('but', 1), ('the', 1), ('butter', 1), ('was', 1), ('bitter', 1)]

3. 合并maplist中元素，元组的第一个索引值相同，则将其第二个索引值相加。

// 备注：准备采用defaultdict。得到的数据如下：{'betty': 1, 'bought': 1, 'a': 1, 'bit': 1, 'of': 1, 'butter': 2, 'but': 1, 'the': 1, 'was': 1, 'bitter': 1}

4. 进行排序，按照key进行字母排序，得到如下：[('a', 1), ('betty', 1), ('bit', 1), ('bitter', 1), ('bought', 1), ('but', 1), ('butter', 2), ('of', 1), ('the', 1), ('was', 1)]

5. 进行二次排序, 按照value进行排序，得到如下：[('butter', 2), ('a', 1), ('betty', 1), ('bit', 1), ('bitter', 1), ('bought', 1), ('but', 1), ('of', 1), ('the', 1), ('was', 1)]

6. 使用切片取出频率较高的*组数据

总结：在python3上不进行defaultdict进行排序结果也是正确的，python2上不正确。defaultdict本身是没有顺序的，要区分列表，所以必须进行排序。

也可尝试自己写，不借助第三方模块

weixin_39972519

关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
python统计中文字数_【Python进阶】用 Python 统计字数

解决问题的思路：1. 将字符串s进行空白符分割得到所有的单词列表split_s，如：['betty', 'bought', 'a', 'bit', 'of', 'butter', 'but', 'the', 'butter', 'was', 'bitter']2. 建立maplist,将split_s转化为元素为元组的列表形式，如：[('betty', 1), ('bought', 1), ('a...
复制链接

扫一扫