LDA实例

最新推荐文章于 2024-03-28 19:04:18 发布

Love鱼小鱼

最新推荐文章于 2024-03-28 19:04:18 发布

阅读量412

点赞数

分类专栏： NLP 文章标签： nlp

本文链接：https://blog.csdn.net/qq_41864339/article/details/111770664

版权

NLP 专栏收录该内容

3 篇文章 0 订阅

订阅专栏

jieba+gensim 参考
 scikit-learn 参考

一. jieba + gensim

from gensim import corpora, models
import jieba.posseg as jp, jieba
# 文本集
texts = [
    '美国教练坦言，没输给中国女排，是输给了郎平',
    '美国无缘四强，听听主教练的评价',
    '中国女排晋级世锦赛四强，全面解析主教练郎平的执教艺术',
    '为什么越来越多的人买MPV，而放弃SUV？跑一趟长途就知道了',
    '跑了长途才知道，SUV和轿车之间的差距',
    '家用的轿车买什么好']
# 分词过滤条件
jieba.add_word('四强', 9, 'n')
flags = ('n', 'nr', 'ns', 'nt', 'eng', 'v', 'd')  # 词性
stopwords = ('没', '就', '知道', '是', '才', '听听', '坦言', '全面', '越来越', '评价', '放弃', '人')  # 停用词
# 分词
words_ls = []
for text in texts:
    words = [w.word for w in jp.cut(text) if w.flag in flags and w.word not in stopwords]
    # words = [(w.word, w.flag) for w in jp.cut(text)]
    words_ls.append(words)

print(words_ls)
dictionary = corpora.Dictionary(words_ls)
# 基于词典，使【词】→【稀疏向量】，并将向量放入列表，形成【稀疏向量集】
corpus = [dictionary.doc2bow(words) for words in words_ls]
print(corpus)
# lda模型，num_topics设置主题的个数
lda = models.ldamodel.LdaModel(corpus=corpus, id2word=dictionary, num_topics=2)
# 打印所有主题，每个主题显示5个词
for topic in lda.print_topics(num_words=5):
    print(topic)
# 推断每个语料库中的主题类别
print(lda.inference(corpus))
for e, values in enumerate(lda.inference(corpus)[0]):
    topic_val = 0
    topic_id = 0
    for tid, val in enumerate(values):
        if val > topic_val:
            topic_val = val
            topic_id = tid
    print(topic_id, '->', texts[e])

二. scikit-learn

from sklearn.feature_extraction.text import CountVectorizer
from sklearn.decomposition import LatentDirichletAllocation

# 加载数据
with open('./1.txt', 'r', encoding='utf-8') as f1:
    res1 = f1.read()
with open('./2.txt', 'r', encoding='utf-8') as f2:
    res2 = f2.read()
with open('./3.txt', 'r', encoding='utf-8') as f3:
    res3 = f3.read()

contents = [res1, res2, res3]
print(contents)
cntVector = CountVectorizer()
# 文本数据转换为向量的形式,要求给定的数据中单词是以空格隔开的
cnt = cntVector.fit_transform(contents)
print(cnt.shape)
print(cnt.toarray())
lda = LatentDirichletAllocation(n_components=2, learning_method='batch', random_state=28)
docs = lda.fit_transform(cnt)
# 打印文档和主题之间的相关性
print(docs)
# 打印主题和词之间的相关性
print(lda.components_)

Love鱼小鱼

关注

0
点赞
踩
4

收藏

觉得还不错? 一键收藏
1
评论
LDA实例

jieba+gensim 参考scikit-learn 参考一. jieba + gensimfrom gensim import corpora, modelsimport jieba.posseg as jp, jieba# 文本集texts = [ '美国教练坦言，没输给中国女排，是输给了郎平', '美国无缘四强，听听主教练的评价', '中国女排晋级世锦赛四强，全面解析主教练郎平的执教艺术', '为什么越来越多的人买MPV，而放弃SUV？跑一趟长途就知道.
复制链接

扫一扫