中文分词

最新推荐文章于 2023-12-18 15:49:45 发布

chuyijian7784

最新推荐文章于 2023-12-18 15:49:45 发布

阅读量83

点赞数

原文链接：https://my.oschina.net/u/3673704/blog/1570124

版权

#分词
#读分词词典,词典中最长词长度
def get_word_dict(dictpath):
    max_index = 1
    with open('dict.txt','r',encoding='utf-8') as f:
        dictwords =f.readlines()
    word_dict = set()
    for word in dictwords:
        word_dict.add(word.strip())
        if len(word)>max_index:
            max_index = len(word)
    return word_dict,max_index
#读取停用词词典
def get_stop_words(stopwordpath):
    with open('dict.txt','r',encoding='utf-8') as f:
        stopwords =f.readlines()
    stop_words = set()
    for word in stopwords:
        stop_words.add(word.strip())
    return stop_words

#分词，返回list
def cut(sentence,dictpath,stopwordpath='dict/stopwords.txt',del_stopword=False):
    start_index = 1
    end_index = len(sentence)

    word_dict, max_index = get_word_dict(dictpath)

    result_sentence = []
    while start_index > 0:
        for start_index in range(max(end_index - max_index, 0), end_index, 1):
            if del_stopword:
                stop_words = get_stop_words(stopwordpath)
                if sentence[start_index:end_index] in stop_words:
                    break
            if sentence[start_index:end_index] in word_dict or end_index == start_index + 1:
                str = sentence[start_index:end_index]
                result_sentence.append(str)
                break
        end_index = start_index

    return result_sentence

结巴分词使用：

http://blog.csdn.net/alis_xt/article/details/53522435

http://blog.csdn.net/wangpei1949/article/details/57077007

转载于:https://my.oschina.net/u/3673704/blog/1570124

chuyijian7784

关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
中文分词

#分词#读分词词典,词典中最长词长度def get_word_dict(dictpath): max_index = 1 with open('dict.txt','r',encoding='utf-8') as f: dictwords =f.readli...
复制链接

扫一扫