pytorch model.train() 和model.eval() 对 BN 层的影响

model.train()

  • BN做归一化时,使用的均值和方差是当前这个Batch的
  • 如果这时 track_running_stats=True, 则会更新running_meanrunning_var
  • 但是,running_meanrunning_var不用在训练阶段

model.eval()

  • BN 做归一化时,使用的均值和方差是BN存储的running_meanrunning_var
  • 不管这时track_running_stats 是 True 还是 False, 都不会更新 running_meanrunning_var

感兴趣可以在以下测试代码下调整测试

'''
Author: Chae Luv
Date: 2022-08-17 22:40:13
LastEditors: Chae Luv
LastEditTime: 2022-08-17 23:15:22
FilePath: /re-record-audio-watermark/10-base_model/test_bn.py
Description: 

Copyright (c) 2022 by Chae Luv/USTC, All Rights Reserved. 
'''
import torch
import torch.nn as nn

def create_inputs():
    return torch.randn(8, 3, 20, 20)


def simulated_bn_forward(x, bn_weight, bn_bias, eps, mean_val=None, var_val=None):
    if mean_val is None:
        mean_val = x.mean([0, 2, 3])
    if var_val is None:
        var_val = x.var([0, 2, 3], unbiased=False)

    x = x - mean_val[None, ..., None, None]
    x = x / torch.sqrt(var_val[None, ..., None, None] + eps)
    x = x * bn_weight[..., None, None] + bn_bias[..., None, None]
    return mean_val, var_val, x

pytorch_bn = nn.BatchNorm2d(num_features=3, momentum=None)
running_mean = torch.zeros(3)
running_var = torch.ones_like(running_mean)

# 切换到eval模式
pytorch_bn.train(mode=False)
test_input = create_inputs()
print(f'pytorch_bn running_mean is {pytorch_bn.running_mean}')
print(f'pytorch_bn running_var is {pytorch_bn.running_var}')
bn_outputs = pytorch_bn(test_input)
print(f'Now pytorch_bn running_mean is {pytorch_bn.running_mean}')
print(f'Now pytorch_bn running_var is {pytorch_bn.running_var}')
# 用之前统计的running_mean和running_var替代输入的running_mean和running_var
_, _, simulated_outputs = simulated_bn_forward(
    test_input, pytorch_bn.weight,
    pytorch_bn.bias, pytorch_bn.eps,
    running_mean, running_var)
assert torch.allclose(simulated_outputs, bn_outputs)

# 关闭track_running_stats后,即使在eval模式下,也会去计算输入的mean和var
pytorch_bn.train(mode=True)
pytorch_bn.track_running_stats = False
bn_outputs_notrack = pytorch_bn(test_input)
_, _, simulated_outputs_notrack = simulated_bn_forward(
    test_input, pytorch_bn.weight,
    pytorch_bn.bias, pytorch_bn.eps)

print(torch.sum(simulated_outputs_notrack - bn_outputs_notrack))
assert torch.allclose(simulated_outputs_notrack, bn_outputs_notrack)
assert not torch.allclose(bn_outputs, bn_outputs_notrack)




  • 5
    点赞
  • 5
    收藏
    觉得还不错? 一键收藏
  • 0
    评论
model.train()和model.eval()是pytorch中用于控制模型训练状态的方法。model.train()将模型设置为训练模式,而model.eval()将模型设置为评估模式。 在训练过程中,model.train()会启用Batch Normalization层(BN层)和Dropout层的计算,以便在每个batch的训练过程中进行正则化和随机失活。同时,它还会更新模型的参数,使其适应训练数据。 相反,model.eval()会将模型设置为评估模式,此时模型不会进行BN层和Dropout层的计算,因为在评估阶段不需要进行正则化和随机失活。此外,模型的参数也不会更新,因为评估阶段只是用来测试模型在新数据上的性能。 需要注意的是,使用model.eval()之后,需要手动使用torch.no_grad()上下文管理器来禁止梯度的计算。torch.no_grad()会包裹住的代码块不会被追踪梯度,也就是说不会记录计算过程,不能进行反向传播更新参数。 综上所述,model.train()用于模型训练阶段,开启BN层和Dropout层的计算并更新参数,而model.eval()用于模型评估阶段,关闭BN层和Dropout层的计算并不更新参数。<span class="em">1</span><span class="em">2</span><span class="em">3</span> #### 引用[.reference_title] - *1* *3* [pytorchmodel.trainmodel.eval](https://blog.csdn.net/dagouxiaohui/article/details/125620786)[target="_blank" data-report-click={"spm":"1018.2226.3001.9630","extra":{"utm_source":"vip_chatgpt_common_search_pc_result","utm_medium":"distribute.pc_search_result.none-task-cask-2~all~insert_cask~default-1-null.142^v93^chatsearchT3_1"}}] [.reference_item style="max-width: 50%"] - *2* [pytorch:model.trainmodel.eval用法及区别详解](https://download.csdn.net/download/weixin_38611254/12855267)[target="_blank" data-report-click={"spm":"1018.2226.3001.9630","extra":{"utm_source":"vip_chatgpt_common_search_pc_result","utm_medium":"distribute.pc_search_result.none-task-cask-2~all~insert_cask~default-1-null.142^v93^chatsearchT3_1"}}] [.reference_item style="max-width: 50%"] [ .reference_list ]

“相关推荐”对你有帮助么?

  • 非常没帮助
  • 没帮助
  • 一般
  • 有帮助
  • 非常有帮助
提交
评论
添加红包

请填写红包祝福语或标题

红包个数最小为10个

红包金额最低5元

当前余额3.43前往充值 >
需支付:10.00
成就一亿技术人!
领取后你会自动成为博主和红包主的粉丝 规则
hope_wisdom
发出的红包
实付
使用余额支付
点击重新获取
扫码支付
钱包余额 0

抵扣说明:

1.余额是钱包充值的虚拟货币,按照1:1的比例进行支付金额的抵扣。
2.余额无法直接购买下载,可以购买VIP、付费专栏及课程。

余额充值