记录 torch.optim.LBFGS

最新推荐文章于 2025-01-18 09:55:26 发布

米饭的白色

最新推荐文章于 2025-01-18 09:55:26 发布

阅读量3.4k

点赞数 1

分类专栏： PyTorch Python 文章标签： pytorch

本文链接：https://blog.csdn.net/mifangdebaise/article/details/126380817

版权

Python 同时被 2 个专栏收录

56 篇文章

订阅专栏

PyTorch

8 篇文章

订阅专栏

本文详细介绍了PyTorch中的LBFGS优化器，它基于L-BFGS算法。该优化器不支持参数组和设备间的参数分布，并且内存需求较高。关键参数包括最大内部迭代次数`max_iter`和最大评估次数`max_eval`，后者决定外部程序何时终止。此外，还提到了学习率、梯度和变化容忍度等设置。对于内存有限的环境，可以调整历史大小来减少内存消耗。

摘要生成于 C知道，由 DeepSeek-R1 满血版支持，前往体验 >

主要记录一下 torch.optim.LBFGS 的用法

源码如下

class LBFGS(Optimizer):
    """Implements L-BFGS algorithm, heavily inspired by `minFunc
    <https://www.cs.ubc.ca/~schmidtm/Software/minFunc.html>`.

    .. warning::
        This optimizer doesn't support per-parameter options and parameter
        groups (there can be only one).

    .. warning::
        Right now all parameters have to be on a single device. This will be
        improved in the future.

    .. note::
        This is a very memory intensive optimizer (it requires additional
        ``param_bytes * (history_size + 1)`` bytes). If it doesn't fit in memory
        try reducing the history size, or use a different algorithm.

    Arguments:
        lr (float): learning rate (default: 1)
        max_iter (int): maximal number of iterations per optimization step
            (default: 20)
        max_eval (int): maximal number of function evaluations per optimization
            step (default: max_iter * 1.25).
        tolerance_grad (float): termination tolerance on first order optimality
            (default: 1e-5).
        tolerance_change (float): termination tolerance on function
            value/parameter changes (default: 1e-9).
        history_size (int): update history size (default: 100).
        line_search_fn (str): either 'strong_wolfe' or None (default: None).
    """

    def __init__(self,
                 params,
                 lr=1,
                 max_iter=20,
                 max_eval=None,
                 tolerance_grad=1e-7,
                 tolerance_change=1e-9,
                 history_size=100,
                 line_search_fn=None):
       ...
       ...