Word2vec 学习

最新推荐文章于 2022-04-05 19:27:48 发布

cjneo

最新推荐文章于 2022-04-05 19:27:48 发布

阅读量177

点赞数

分类专栏：机器学习

原文链接：https://www.jianshu.com/p/a1163174ebaf

版权

机器学习专栏收录该内容

14 篇文章 0 订阅

订阅专栏

https://blog.csdn.net/mr_tyting/article/details/80091842

这个是非常经典的word2vect 的论文翻译

另外一个作者的写的非常详细

https://www.jianshu.com/p/a1163174ebaf

https://www.jianshu.com/p/ed15e2adbfad

这个是重要的举措

negative sampling

nceloss 经典

https://blog.csdn.net/diamonjoy_zone/article/details/67638243

https://www.jianshu.com/p/fab82fa53e16

这里写图片描述

训练的两种方式:

1 输入：前后的若干个词 ;输出:单个单词

2 和1 输入输出互换

目的：

将高维的稀疏向量映射到一个新的低维稠密空间，且词向量之前的相似性表征词的关联性。

巧妙之处：

不需要标注数据，而是使用句子中距离相近的词的关联性来进行训练。

word2vec工作流程
1word2Vec只是一个三层的神经网络。
2 喂给模型一个word，然后用来预测它周边的词。
3 然后去掉最后一层，只保存input_layerinput 和 hidden_layer。
4 从词表中选取一个词，喂给模型，在hidden_layer 将会给出该词的embedding repesentation。

这里写图片描述

把datadatadata打印出来看看？

print(data)
[['he', 'is'],
 ['he', 'the'],
 ['is', 'he'],
 ['is', 'the'],
 ['is', 'king'],
 ['the', 'he'],
 ['the', 'is'],
 ['the', 'king'],
.
.
.
]

# function to convert numbers to one hot vectors
def to_one_hot(data_point_index, vocab_size):
    temp = np.zeros(vocab_size)
    temp[data_point_index] = 1
    return temp
x_train = [] # input word
y_train = [] # output word
for data_word in data:
    x_train.append(to_one_hot(word2int[ data_word[0] ], vocab_size))
    y_train.append(to_one_hot(word2int[ data_word[1] ], vocab_size))
# convert them to numpy arrays
x_train = np.asarray(x_train)
y_train = np.asarray(y_train)

利用tensorflowtensorflowtensorflow建立模型

# making placeholders for x_train and y_train
x = tf.placeholder(tf.float32, shape=(None, vocab_size))
y_label = tf.placeholder(tf.float32, shape=(None, vocab_size))

EMBEDDING_DIM = 5 # you can choose your own number
W1 = tf.Variable(tf.random_normal([vocab_size, EMBEDDING_DIM]))
b1 = tf.Variable(tf.random_normal([EMBEDDING_DIM])) #bias
hidden_representation = tf.add(tf.matmul(x,W1), b1)

W2 = tf.Variable(tf.random_normal([EMBEDDING_DIM, vocab_size]))
b2 = tf.Variable(tf.random_normal([vocab_size]))
prediction = tf.nn.softmax(tf.add( tf.matmul(hidden_representation, W2), b2))


sess = tf.Session()
init = tf.global_variables_initializer()
sess.run(init) #make sure you do this!
# define the loss function:
cross_entropy_loss = tf.reduce_mean(-tf.reduce_sum(y_label * tf.log(prediction), reduction_indices=[1]))
# define the training step:
train_step = tf.train.GradientDescentOptimizer(0.1).minimize(cross_entropy_loss)
n_iters = 10000
# train for n_iter iterations
for _ in range(n_iters):
    sess.run(train_step, feed_dict={x: x_train, y_label: y_train})
    print('loss is : ', sess.run(cross_entropy_loss, feed_dict={x: x_train, y_label: y_train}))

这里写图片描述

减少计算量：

Negative Sampling · 负采样

在训练神经网络时，每当接受一个训练样本，然后调整所有神经单元权重参数，来使神经网络预测更加准确。换句话说，每个训练样本都将会调整所有神经网络中的参数。
我们词汇表的大小决定了我们skip-gram 神经网络将会有一个非常大的权重参数，并且所有的权重参数会随着数十亿训练样本不断调整。

negative sampling 每次让一个训练样本仅仅更新一小部分的权重参数，从而降低梯度下降过程中的计算量。
如果 vocabulary 大小为1万时，当输入样本 ( "fox", "quick") 到神经网络时， “ fox” 经过 one-hot 编码，在输出层我们期望对应 “quick” 单词的那个神经元结点输出 1，其余 9999 个都应该输出 0。在这里，这9999个我们期望输出为0的神经元结点所对应的单词我们为 negative word. negative sampling 的想法也很直接，将随机选择一小部分的 negative words，比如选 10个 negative words 来更新对应的权重参数。

在论文中作者指出指出对于小规模数据集，建议选择 5-20 个 negative words，对于大规模数据集选择 2-5个 negative words.

如果使用了 negative sampling 仅仅去更新positive word- “quick” 和选择的其他 10 个negative words 的结点对应的权重，共计 11 个输出神经元，相当于每次只更新 300 x 11 = 3300 个权重参数。对于 3百万的权重来说，相当于只计算了千分之一的权重，这样计算效率就大幅度提高。