input_Ids, attention_mask...

最新推荐文章于 2024-05-28 12:31:11 发布

Zhiq_yuan

最新推荐文章于 2024-05-28 12:31:11 发布

阅读量1w

点赞数 12

文章标签：自然语言处理人工智能 nlp

本文链接：https://blog.csdn.net/weixin_44219178/article/details/121991171

版权

这里选用的模型是Roberta。对源码当中的self.embedding 的几个参数做一个说明

tokenizer = RobertaTokenizer.from_pretrained('microsoft/codebert-base-mlm')
model = RobertaForMaskedLM.from_pretrained('microsoft/codebert-base-mlm')

源码：

embedding_output = self.embeddings(
        input_ids=input_ids, position_ids=position_ids, 
        token_type_ids=token_type_ids, inputs_embeds=inputs_embeds
        )   
     
encoder_outputs = self.encoder(
            embedding_output,
            attention_mask=extended_attention_mask,
            head_mask=head_mask,
            encoder_hidden_states=encoder_hidden_states,
            encoder_attention_mask=encoder_extended_attention_mask,
            output_attentions=output_attentions,
            output_hidden_states=output_hidden_states,
            return_dict=return_dict,
        )

1, input_ids: 将输入到的词映射到模型当中的字典ID

input_ids = tokenizer.encode("I love China!", add_special_tokens=False)
# print:[100, 657, 436, 328]

# 可以使用 如下转回原来的词
tokens= tokenizer.convert_ids_to_tokens(input_ids)
# print: ['I', 'Ġlove', 'ĠChina', '!']. Note： Ġ 代码该字符的前面是一个空格

2，attention_mask: 有时，需要将多个不同长度的sentence，统一为同一个长度，例如128 dim. 此时我们会需要加padding，以此将一些长度不足的128的sentence，用1进行填充。为了让模型avoid performing attention on padding token indices. 所以这个需要加上这个属性。如果处理的文本是一句话，就可以不用了。如果不传入attention_mask时，模型会自动全部用1填充。

input_ids = tokenizer(["I love China","I love my family and I enjoy the time with my family"], padding=True)

# print:
# {
#'input_ids': [[0, 100, 657, 436, 2, 1, 1, 1, 1, 1, 1, 1, 1, 1],
#              [0, 100, 657, 127, 284, 8, 38, 2254, 5, 86, 19, 127, 284, 2]], 
#'attention_mask': [[1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0],
#                   [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1]]
#}

3、token_type_ids: 值得注意的是，由于roberta移除了NSP的任务，所以在直接使用tokenizer()时，只会返回前面两个值。

Segment token indices to indicate first and second portions of the inputs.

Indices are selected in ``[0, 1]``:

            - 0 corresponds to a `sentence A` token,
            - 1 corresponds to a `sentence B` token.

4、position_ids: 下图中的position_ids 当中1表示是padding出来的值，非1值是原先的word-index

if position_ids is None:
   if input_ids is not None:
      # Create the position ids from the input token ids. Any padded tokens remain padded.
      position_ids = create_position_ids_from_input_ids(input_ids, self.padding_idx).to(input_ids.device)
    else:
         position_ids = self.create_position_ids_from_inputs_embeds(inputs_embeds)

 #"position_ids'' : tensor([[ 2,  3,  4,  5,  6,  7,  8,  9, 10],
        [ 2,  3,  4,  5,  6,  1,  1,  1,  1]])

最后模型输出的向量相加得到模型最后的word embedding:

参考文献
bert模型简介、transformers中bert模型源码阅读、分类任务实战和难点总结_HUSTHY的博客-CSDN博客_bert encoder模块

Zhiq_yuan

关注

12
点赞
踩
29

收藏

觉得还不错? 一键收藏
0
评论
input_Ids, attention_mask...

input_Ids, attention_mask...
复制链接

扫一扫

input_Ids, attention_mask...

“相关推荐”对你有帮助么？