数据处理-------利用jieba对数据集进行分词和统计频数

最新推荐文章于 2022-10-08 20:53:37 发布

Ge1009

最新推荐文章于 2022-10-08 20:53:37 发布

阅读量2.9k

点赞数 2

文章标签：数据处理

本文链接：https://blog.csdn.net/Ge0022/article/details/84111058

版权

一，对txt文件中出现的词语的频数统计再找出出现频率多的
二，代码：

import re
from collections import Counter
import jieba


def cut_word(datapath):
    with open(datapath,'r',encoding='utf-8')as fp:
        string = fp.read()
        data = re.sub(r"[\s+\.\!\/_,$%^*(【】：\]\[\-:;+\"\']+|[+——！，。？、~@#￥%……&*（）]+|[0-9]+", "", string)
        word_list = jieba.cut(data)
        print(type(word_list))
        return word_list

def static_top_word(word_list,top=5):
    result = dict(Counter(word_list))
    print(result)
    sortlist = sorted(result.items(),key=lambda x:x[1],reverse=True)
    resultlist = []
    for i in range(0,top):
        resultlist.append(sortlist[i])
    return resultlist


def main():
    datapath = 'comment.txt'
    word_list = cut_word(datapath)
    Result = static_top_word(word_list)
    print(Result)
main()

三，用正则对特殊符号过滤,用re.sub()对字符进行空字符替换

确定要放弃本次机会？

福利倒计时

: :

立减 ¥

普通VIP年卡可用

立即使用

Ge1009

关注关注

2
点赞
踩
6

收藏

觉得还不错? 一键收藏
1
评论
数据处理-------利用jieba对数据集进行分词和统计频数

一，对txt文件中出现的词语的频数统计再找出出现频率多的二，代码：import refrom collections import Counterimport jiebadef cut_word(datapath): with open(datapath,'r',encoding='utf-8')as fp: string = fp.read() ...
复制链接

扫一扫