详解TF-IDF

最新推荐文章于 2021-03-26 04:46:10 发布

来自宇宙岛的海龟

最新推荐文章于 2021-03-26 04:46:10 发布

阅读量972

点赞数

分类专栏：算法 ML

本文链接：https://blog.csdn.net/weixin_44064649/article/details/103651607

版权

TF-IDF是一种统计方法，用于评估词在文档中的重要性。它由词频（TF）和逆文档频率（IDF）组成。TF衡量词在文档中的出现频率，IDF则考虑词在整个文档集合中的普遍性。TF-IDF常用于信息检索和文本挖掘，通过加权计算提高重要词的权重，降低常见词的影响。文章通过实例和代码展示了如何计算TF-IDF。

摘要由CSDN通过智能技术生成

什么是TF-IDF

TF-IDF（term frequency–inverse document frequency）是一种用于信息检索与数据挖掘的常用加权技术。TF意思是词频(Term Frequency)，IDF意思是逆文本频率指数(Inverse Document Frequency)。TF-IDF是一种统计方法，用以评估一字词对于一个文件集或一个语料库中的其中一份文件的重要程度。字词的重要性随着它在文件中出现的次数成正比增加，但同时会随着它在语料库中出现的频率成反比下降。

看看官网的解释：
Tf-idf stands for term frequency-inverse document frequency, and the tf-idf weight is a weight often used in information retrieval and text mining. This weight is a statistical measure used to evaluate how important a word is to a document in a collection or corpus. The importance increases proportionally to the number of times a word appears in the document but is offset by the frequency of the word in the corpus. Variations of the tf-idf weighting scheme are often used by search engines as a central tool in scoring and ranking a document’s relevance given a user query.
One of the simplest ranking functions is computed by summing the tf-idf for each query term; many more sophisticated ranking functions are variants of this simple model.Tf-idf can be successfully used for stop-words filtering in various subject fields including text summarization and classification.

怎么计算

Typically, the tf-idf weight is composed by two terms: the first computes the normalized Term Frequency (TF), aka. the number of times a word appears in a document, divided by the total number of words in that document; the second term is the Inverse Document Frequency (IDF), computed as the logarithm of the number of the documents in the corpus divided by the number of documents where the specific term appears.

TF: Term Frequency, which measures how frequently a term occurs in a document. Since every document is different in length, it is possible that a term would appear much more times in long documents than shorter ones. Thus, the term frequency is often divided by the docum

最低0.47元/天解锁文章

来自宇宙岛的海龟

关注

0
点赞
踩
2

收藏

觉得还不错? 一键收藏
0
评论
详解TF-IDF

目录什么是TF-IDF怎么计算举例例1例2再看代码什么是TF-IDFTF-IDF（term frequency–inverse document frequency）是一种用于信息检索与数据挖掘的常用加权技术。TF意思是词频(Term Frequency)，IDF意思是逆文本频率指数(Inverse Document Frequency)。TF-IDF是一种统计方法，用以评估一字词对于一个文件...
复制链接

扫一扫

专栏目录