机器学习-中文文本特征抽取

最新推荐文章于 2023-01-01 18:15:00 发布

我有辣条你跟不跟我走_

最新推荐文章于 2023-01-01 18:15:00 发布

阅读量209

点赞数 1

文章标签：机器学习 python sklearn

本文链接：https://blog.csdn.net/m0_50843768/article/details/120891184

版权

对字符串列表 data = ['发表回复这件事', '飞机里面飞一杯飞机专属奶茶', '没有什么比在飞机上喝一杯飞机专属的飞机奶茶要更好了'] 进行中文文本特征抽取

import sklearn.feature_extraction.text as text
import jieba


transfer = text.CountVectorizer(stop_words=['vb'])


def count_chinese_demo2():
    data = ['发表回复这件事', '飞机里面飞一杯飞机专属奶茶', '没有什么比在飞机上喝一杯飞机专属的飞机奶茶要更好了']
    data_new = []
    # 中文文本分词
    for send in data:
        data_new.append(' '.join(list(jieba.cut(send))))
    print(data_new)

    # 文本特征提取
    data_final = transfer.fit_transform(data_new)
    print(data_final.toarray())
    print(transfer.get_feature_names())


if __name__ == "__main__":
    count_chinese_demo2()

输出：

['发表回复这件事', '飞机里面飞一杯飞机专属奶茶', '没有什么比在飞机上喝一杯飞机专属的飞机奶茶要更好了']
[[0 0 0 1 0 1 0 0 0 1 0 0]
[1 1 0 0 0 0 1 0 0 0 1 2]
[0 1 1 0 1 0 1 1 1 0 0 3]]
['一杯', '专属', '什么', '发表', '喝一杯', '回复', '奶茶', '更好', '没有', '这件', '里面', '飞机']

我有辣条你跟不跟我走_

关注

1
点赞
踩
2

收藏

觉得还不错? 一键收藏
0
评论
机器学习-中文文本特征抽取

对字符串列表 data = ['发表回复这件事', '飞机里面飞一杯飞机专属奶茶', '没有什么比在飞机上喝一杯飞机专属的飞机奶茶要更好了'] 进行中文文本特征抽取import sklearn.feature_extraction.text as textimport jiebatransfer = text.CountVectorizer(stop_words=['vb'])def count_chinese_demo2(): data = ['发表回复这件事', '飞机里面.
复制链接

扫一扫