自然语言处理_基础技术4_CountVectorizer

最新推荐文章于 2022-11-10 11:49:42 发布

dataastron

最新推荐文章于 2022-11-10 11:49:42 发布

阅读量295

点赞数

分类专栏：自然语言处理

本文链接：https://blog.csdn.net/dataastron/article/details/106306643

版权

自然语言处理专栏收录该内容

6 篇文章 1 订阅

订阅专栏

onehot编码是一种稀疏编码方式,如果词语越多,维度也越大.会出现维数灾难.
针对one-hot编码,sklearn中实现如下.CountVectorizer类(计数向量)
先用英文举例.
它会针对每个单词计数,丢失位置信息

from sklearn.feature_extraction.text import CountVectorizer
corpus = [
    'This is the first document.',
    'This document is the second document.',
    'And this is the third one.',
    'Is this the first document?',
]
vectorizer = CountVectorizer()
X = vectorizer.fit_transform(corpus)
print(vectorizer.get_feature_names())

print(X.toarray())

下面这个是ngram为2的编码方式

vectorizer2 = CountVectorizer(analyzer='word', ngram_range=(2, 2))
X2 = vectorizer2.fit_transform(corpus)
print(vectorizer2.get_feature_names())

确定要放弃本次机会？

福利倒计时

: :

立减 ¥

普通VIP年卡可用

立即使用

dataastron

关注关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
自然语言处理_基础技术4_CountVectorizer

onehot编码是一种稀疏编码方式,如果词语越多,维度也越大.会出现维数灾难.针对one-hot编码,sklearn中实现如下.CountVectorizer类(计数向量)先用英文举例.它会针对每个单词计数,丢失位置信息from sklearn.feature_extraction.text import CountVectorizercorpus = [ 'This is the first document.', 'This document is the second doc
复制链接

扫一扫