“21天好习惯”第一期-21

15083838632

于 2021-11-12 07:00:00 发布

阅读量144

点赞数

分类专栏：笔记文章标签： python

本文链接：https://blog.csdn.net/m0_63172885/article/details/121195451

版权

Python 词频统计文本处理停用词数据分析

关键词由CSDN通过智能技术生成

笔记专栏收录该内容

21 篇文章

订阅专栏

Python 文本词频统计

excludes={"the","and","of","you","my"}
def getText():
    txt=open(r"C:\Users\DELL\Desktop\英语短文.txt","r",encoding="UTF-8").read()
    txt=txt.lower()
    for ch in '"!@#$%&()+-,.:;?/{}[]~`':
        txt=txt.replace(ch," ")
    return txt
英语短文Txt=getText()
words=英语短文Txt.split()
counts={}
for word in words:
    counts[word]=counts.get(word,0)+1
for word in excludes:
    del(counts[word])
items=list(counts.items())
items.sort(key=lambda x:x[1],reverse=True)
for i in range(10):
    word,count=items[i]
    print("{0:<10}{1:>5}".format(word,count))