pytorh .to(device) 和.cuda()的区别

最新推荐文章于 2025-02-24 10:28:45 发布

Golden-sun

最新推荐文章于 2025-02-24 10:28:45 发布

阅读量2.6w

点赞数 34

分类专栏： Pytorch训练技巧文章标签：深度学习 GPU CUDA DataParallel 模型部署

本文链接：https://blog.csdn.net/weixin_43402775/article/details/109223794

版权

Pytorch训练技巧专栏收录该内容

17 篇文章

订阅专栏

本文介绍了如何将深度学习模型部署到单个GPU或多个GPU环境中。使用`.to(device)`函数可以方便地在CPU和GPU之间切换，而`nn.DataParallel`用于实现模型的并行计算。在多GPU环境下，通过设置`CUDA_VISIBLE_DEVICES`环境变量来指定使用哪些GPU，并利用`nn.DataParallel`进行模型的分布式训练。

摘要生成于 C知道，由 DeepSeek-R1 满血版支持，前往体验 >

原理

.to(device) 可以指定CPU 或者GPU

device = torch.device("cuda:0" if torch.cuda.is_available() else "cpu") # 单GPU或者CPU
model.to(device)
#如果是多GPU
if torch.cuda.device_count() > 1:
  model = nn.DataParallel(model，device_ids=[0,1,2])
model.to(device)

.cuda() 只能指定GPU

#指定某个GPU
os.environ['CUDA_VISIBLE_DEVICE']='1'
model.cuda()
#如果是多GPU
os.environment['CUDA_VISIBLE_DEVICES'] = '0,1,2,3'
device_ids = [0,1,2,3]
net  = torch.nn.Dataparallel(net, device_ids =device_ids)
net  = torch.nn.Dataparallel(net) # 默认使用所有的device_ids 
net = net.cuda()

class DataParallel(Module):
    def __init__(self, module, device_ids=None, output_device=None, dim=0):
        super(DataParallel, self).__init__()

        if not torch.cuda.is_available():
            self.module = module
            self.device_ids = []
            return

        if device_ids is None:
            device_ids = list(range(torch.cuda.device_count()))
        if output_device is None:
            output_device = device_ids[0]