动手学习深度学习（Pytorch版）Task 2:文本预处理

最新推荐文章于 2022-09-08 18:09:27 发布

zpc1219

最新推荐文章于 2022-09-08 18:09:27 发布

阅读量253

点赞数

分类专栏：技术文章标签： python 正则表达式自然语言处理

本文链接：https://blog.csdn.net/zpc1219/article/details/104314895

版权

技术专栏收录该内容

17 篇文章 0 订阅

订阅专栏

文本预处理

文本是一类序列数据，一篇文章可以看作是字符或单词的序列，文本数据常见的预处理四个步骤如下：

读入文本
分词
建立字典，将每个词映射到一个唯一的索引（index）
将文本从词的序列转换为索引的序列，方便输入模型

读入文本

数据集：英文小说——H. G. Well的Time Machine

import collections
import re

def read_time_machine():
    #只读方式打开存放在与代码文件相同目录下的文本
    with open('timemachine.txt','r') as f:
        lines=[re.sub('[^a-z]+',' ',line.strip().lower()) for line in f]
    return lines

lines=read_time_machine()
print('# sentences %d'%len(lines))

# sentences 3221

lines[0]

'the time machine by h g wells'

要点
1、.strip():

str.strip([chars]);去除字符串前面和后面的所有设置的字符串，默认为空格.例子如下：

st=" wang ge "
st=st.strip()
print("lao"+st+"!")

laowang ge!

如果设置了字符序列的话，那么它会删除，字符串前后出现的所有序列中有的字符（包括两字符之间的序列）。但不会清除空格。

st=st.strip('l,o,e')
print(st)#l、a、o、e均被删除

wang g

2、.lower()
str.lower():把字符串中的大写字母变为小写。例子如下：

st="ABCDe"
print(st.lower())

abcde

3、re.sub()
替换字符串中的某些子串，可以用正则表达式来匹配被选子串。
re.sub(pattern, repl, string, count=0, flags=0)
pattern：表示正则表达式中的模式字符串；
repl：被替换的字符串（既可以是字符串，也可以是函数）；
string：要被处理的，要被替换的字符串；
count：匹配的次数, 默认是全部替换

import re
st='aswdefrgTHHGFsxgshfubRFGG,df gtb 5 $*/ghm'
print(re.sub('[^a-z]+','换',st.strip().lower()))#保留‘[^a-z]+’代表的所有字母，用‘换’字替换不是字母的它们，发现一串非要字母仅以一个‘换’代替

aswdefrgthhgfsxgshfubrfgg换df换gtb换ghm

分词

对每个句子进行分词，也就是将一个句子划分成若干个词（token），转换为一个词的序列。

def tokenize(sentences, token='word'):
    """Split sentences into word or char tokens"""
    if token == 'word':
        return [sentence.split(' ') for sentence in sentences]##sentence.split(' ')以空格为间隔符，划分单词，每次返回一个列表
    elif token == 'char':
        return [list(sentence) for sentence in sentences]
    else:
        print('ERROR: unkown token type '+token)

tokens = tokenize(lines)
tokens[0:2]

[['the', 'time', 'machine', 'by', 'h', 'g', 'wells', ''], ['']]

建立字典

为了方便模型处理，我们需要将字符串转换为数字。因此我们需要先构建一个字典（vocabulary），将每个词映射到一个唯一的索引编号。

class Vocab(object):
    def __init__(self, tokens, min_freq=0, use_special_tokens=False):
        counter = count_corpus(tokens)  # : 
        self.token_freqs = list(counter.items())
        self.idx_to_token = []
        if use_special_tokens:
            # padding, begin of sentence, end of sentence, unknown
            self.pad, self.bos, self.eos, self.unk = (0, 1, 2, 3)
            self.idx_to_token += ['', '', '', '']
        else:
            self.unk = 0
            self.idx_to_token += ['']
        self.idx_to_token += [token for token, freq in self.token_freqs
                        if freq >= min_freq and token not in self.idx_to_token]
        self.token_to_idx = dict()
        for idx, token in enumerate(self.idx_to_token):
            self.token_to_idx[token] = idx

    def __len__(self):
        return len(self.idx_to_token)

    def __getitem__(self, tokens):
        if not isinstance(tokens, (list, tuple)):
            return self.token_to_idx.get(tokens, self.unk)
        return [self.__getitem__(token) for token in tokens]

    def to_tokens(self, indices):
        if not isinstance(indices, (list, tuple)):
            return self.idx_to_token[indices]
        return [self.idx_to_token[index] for index in indices]

def count_corpus(sentences):
    tokens = [tk for st in sentences for tk in st]
    return collections.Counter(tokens)  # 返回一个字典，记录每个词的出现次数

我们看一个例子，这里我们尝试用Time Machine作为语料构建字典

vocab = Vocab(tokens)
print(list(vocab.token_to_idx.items())[0:10])

[('', 0), ('the', 1), ('time', 2), ('machine', 3), ('by', 4), ('h', 5), ('g', 6), ('wells', 7), ('i', 8), ('traveller', 9)]

将词转为索引

使用字典，我们可以将原文本中的句子从单词序列转换为索引序列

for i in range(8, 10):
    print('words:', tokens[i])
    print('indices:', vocab[tokens[i]])

words: ['the', 'time', 'traveller', 'for', 'so', 'it', 'will', 'be', 'convenient', 'to', 'speak', 'of', 'him', '']
indices: [1, 2, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 0]
words: ['was', 'expounding', 'a', 'recondite', 'matter', 'to', 'us', 'his', 'grey', 'eyes', 'shone', 'and']
indices: [20, 21, 22, 23, 24, 16, 25, 26, 27, 28, 29, 30]

用现有工具进行分词

我们前面介绍的分词方式非常简单，它至少有以下几个缺点:

标点符号通常可以提供语义信息，但是我们的方法直接将其丢弃了
类似“shouldn’t", "doesn’t"这样的词会被错误地处理
类似"Mr.", "Dr."这样的词会被错误地处理

我们可以通过引入更复杂的规则来解决这些问题，但是事实上，有一些现有的工具可以很好地进行分词，我们在这里简单介绍其中的两个：spaCy和NLTK。

下面是一个简单的例子：

text = "Mr. Chen doesn't agree with my suggestion."

spaCy:

import spacy
nlp = spacy.load('en_core_web_sm')
doc = nlp(text)
print([token.text for token in doc])

['Mr.', 'Chen', 'does', "n't", 'agree', 'with', 'my', 'suggestion', '.']

NLTK:

from nltk.tokenize import word_tokenize
from nltk import data
data.path.append('/home/kesci/input/nltk_data3784/nltk_data')
print(word_tokenize(text))

['Mr.', 'Chen', 'does', "n't", 'agree', 'with', 'my', 'suggestion', '.']

zpc1219

关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
动手学习深度学习（Pytorch版）Task 2:文本预处理

文本预处理文本是一类序列数据，一篇文章可以看作是字符或单词的序列，文本数据常见的预处理四个步骤如下：读入文本分词建立字典，将每个词映射到一个唯一的索引（index）将文本从词的序列转换为索引的序列，方便输入模型读入文本数据集：英文小说——H. G. Well的Time Machineimport collectionsimport redef read_time_mach...
复制链接

扫一扫