CS224n课程Assignment1参考答案

最新推荐文章于 2024-05-03 21:41:37 发布

Jonariguez

最新推荐文章于 2024-05-03 21:41:37 发布

阅读量1.8k

点赞数 2

分类专栏：深度学习自然语言处理文章标签： CS224n nlp stanford word2vec

本文链接：https://blog.csdn.net/u013068502/article/details/91897480

版权

$Assignment\#1-solution\quad By\; Jonariguez$

所有的代码题目对应的代码已上传至github/CS224n/Jonariguez

解：
$softmax(\mathbf{x})_i=\frac{e^{x_i}}{\sum_{j}{e^{x_j}}}=\frac{e^ce^{x_i}}{e^c\sum_{j}{e^{x_j}}}=\frac{e^{x_i+c}}{\sum_{j}{e^{x_j+c}}}=softmax(\mathbf{x}+c)_i$
即
$softmax(\mathbf{x})=softmax(\mathbf{x}+c)$
证毕

解：
直接在代码中利用numpy实现即可。注意要先从 $x$ 中减去每一行的最大值，这样在保证结果不变的情况下，所有的元素不大于0，不会出现上溢出，从而保证结果的正确性。具体可参考 http://www.hankcs.com/ml/computing-log-sum-exp.html

def softmax(x):
   """Compute the softmax function for each row of the input x.

   It is crucial that this function is optimized for speed because
   it will be used frequently in later code. You might find numpy
   functions np.exp, np.sum, np.reshape, np.max, and numpy
   broadcasting useful for this task.

   Numpy broadcasting documentation:
   http://docs.scipy.org/doc/numpy/user/basics.broadcasting.html

   You should also make sure that your code works for a single
   N-dimensional vector (treat the vector as a single row) and
   for M x N matrices. This may be useful for testing later. Also,
   make sure that the dimensions of the output match the input.

   You must implement the optimization in problem 1(a) of the
   written assignment!

   Arguments:
   x -- A N dimensional vector or M x N dimensional numpy matrix.

   Return:
   x -- You are allowed to modify x in-place
   """
   orig_shape = x.shape

   if len(x.shape) > 1:
       # Matrix
       # 每行减去该行的最大值
       x = x-np.max(x,axis=1).reshape(x.shape[0],1)
       # 然后进行softmax计算
       x = np.exp(x)/np.sum(np.exp(x),axis=1).reshape(x.shape[0],1)
   else:
       # Vector
       x = x-np.max(x)
       x = np.exp(x)/np.sum(np.exp(x))

   assert x.shape == orig_shape
   return x

解：
$\sigma'(x)=\frac{e^{-x}}{(1+e^{-x})^2}=\frac{1}{1+e^{-x}}\cdot\frac{e^{-x}}{1+e^{-x}}=\sigma(x)\cdot(1-\sigma(x))$

即 $s i g m o i d$ 函数的求导可以由其本身来表示。

解：
我们知道真实标记 $y$ 是one-hot向量，因此我们下面的推导都基于 $y_k=1$ ,且 $y_i=0,i\neq k$ ，即真实标记是 $k$ .

$\frac{\partial CE(y,\hat{y})}{\partial\theta}=\frac{\partial CE(y,\hat{y})}{\partial\hat{y}}\cdot\frac{\partial\hat{y}}{\partial\theta}$

其中：
$\frac{\partial CE(y,\hat{y})}{\partial\hat{y}}=-\sum_{i}{\frac{y_i}{\hat{y}_i}}=-\frac{1}{\hat{y}_k}$

接下来讨论 $\frac{\partial\hat{y}}{\partial\theta}$ :

$i = k$ :
$\frac{\partial\hat{y}}{\partial\theta_k}=\frac{\partial}{\partial\theta_k}(\frac{e^{\theta_k}}{\sum_{j}{e^{\theta_j}}})=\hat{y}_k\cdot(1-\hat{y}_k)$

则：
$\frac{\partial CE}{\theta_i}=\frac{\partial CE}{\partial\hat{y}}\frac{\partial\hat{y}}{\theta_i}=-\frac{1}{\hat{y}_k}\cdot\hat{y}_k\cdot(1-\hat{y}_k)=\hat{y}_i-1$

最低0.47元/天解锁文章

Jonariguez

关注

2
点赞
踩
2

收藏

觉得还不错? 一键收藏
0
评论
CS224n课程Assignment1参考答案

Assignment#1−solutionBy&ThickSpace;Jonariguez Assignment\#1-solution\quad By\; Jonariguez Assignment#1−solutionByJonariguez所有的代码题目对应的代码已上传至github/CS224n/Jonariguez解：softmax(x)i=exi∑jexj=ecexie...
复制链接

扫一扫

专栏目录