RTDETR结合SIoU、GIoU、CIoU、DIoU、WIoU比较和涨点

繁星似锁

已于 2024-07-18 14:31:16 修改

阅读量464

点赞数 18

分类专栏： RTDETR智能沙龙文章标签：目标跟踪人工智能计算机视觉

于 2024-07-18 13:45:00 首次发布

本文链接：https://blog.csdn.net/weixin_48113312/article/details/140520468

版权

RTDETR智能沙龙专栏收录该内容

1 篇文章 0 订阅

订阅专栏

引言

在目标检测领域，交并比（IoU）是衡量预测框与真实框之间重叠程度的重要指标。然而，传统的IoU在某些情况下存在局限性，因此衍生出了多种改进的IoU变体，包括SIoU、WIoU、GIoU、DIoU、EIoU和CIoU。这些变体旨在更准确地评估预测框与真实框之间的匹配程度，从而提高目标检测的精度。

1 原理介绍

SIoU (Simple IoU)

SIoU考虑了真实框与预测框之间的角度、距离以及形状，其损失函数由Angle cost、Distance cost、Shape cost和IoU cost四部分组成。这使得SIoU能够更快地收敛并防止网络在最优解附近跳动。然而，SIoU在不同数据集上的效果可能不一，且角度、距离以及形状的损失权重需要进一步优化。

WIoU（Weighted IoU）

WIoU通过引入权重系数，将目标的位置和大小信息考虑在内，从而更准确地衡量目标检测的性能。这种加权的方式使得WIoU在处理不同大小和位置的目标时更加灵活和准确。然而，WIoU的计算方式相对复杂，需要额外的计算资源。

GIoU（Generalized IoU）

GIoU是IoU的一个扩展，它解决了IoU在处理非重叠目标时的局限性。GIoU的计算方式不仅考虑了两边界框之间的交集面积，还加入了它们并集之外的区域。这使得GIoU在目标没有重叠时仍能提供一个有意义的损失值，从而引导模型进行优化。

DIoU（Distance IoU）

DIoU进一步考虑了预测框与真实框之间的中心点距离。它不仅关注重叠区域，还通过引入中心点距离来优化预测框的位置。这使得DIoU在处理目标位置偏差时更加敏感和准确。

EIoU（Embedding Intersection over Union）

EIoU是一种创新的IoU计算方法，它将目标和预测框映射到嵌入空间中，并利用嵌入空间的距离来衡量目标与预测框之间的相似性。这种方法在处理复杂场景和相似目标时具有较高的准确性。然而，EIoU的计算复杂度相对较高，需要额外的嵌入空间映射过程。

CIoU（Complete IoU）

CIoU综合考虑了重叠面积、中心点距离以及长宽比等因素，旨在更全面地评估预测框与真实框之间的匹配程度。CIoU在处理不同形状和大小的目标时具有较好的鲁棒性和准确性。然而，其计算过程也相对复杂，需要平衡各个因素的权重。

2 代码引用

import torch
import math
 
class WIoU_Scale:
    ''' monotonous: {
            None: origin v1
            True: monotonic FM v2
            False: non-monotonic FM v3
        }
        momentum: The momentum of running mean'''
    iou_mean = 1.
    monotonous = False
    _momentum = 1 - 0.5 ** (1 / 7000)
    _is_train = True
 
    def __init__(self, iou):
        self.iou = iou
        self._update(self)
 
    @classmethod
    def _update(cls, self):
        if cls._is_train: cls.iou_mean = (1 - cls._momentum) * cls.iou_mean + \
                                         cls._momentum * self.iou.detach().mean().item()
 
    @classmethod
    def _scaled_loss(cls, self, gamma=1.9, delta=3):
        if isinstance(self.monotonous, bool):
            if self.monotonous:
                return (self.iou.detach() / self.iou_mean).sqrt()
            else:
                beta = self.iou.detach() / self.iou_mean
                alpha = delta * torch.pow(gamma, beta - delta)
                return beta / alpha
        return 1
 
def bbox_iou(box1, box2, xywh=True, GIoU=False, DIoU=False, CIoU=False, EIoU=False, SIoU=False, WIoU=False, eps=1e-7):
    """
    Calculate Intersection over Union (IoU) of box1(1, 4) to box2(n, 4).
    Args:
        box1 (torch.Tensor): A tensor representing a single bounding box with shape (1, 4).
        box2 (torch.Tensor): A tensor representing n bounding boxes with shape (n, 4).
        xywh (bool, optional): If True, input boxes are in (x, y, w, h) format. If False, input boxes are in
                               (x1, y1, x2, y2) format. Defaults to True.
        GIoU (bool, optional): If True, calculate Generalized IoU. Defaults to False.
        DIoU (bool, optional): If True, calculate Distance IoU. Defaults to False.
        CIoU (bool, optional): If True, calculate Complete IoU. Defaults to False.
        EIoU (bool, optional): If True, calculate Efficient IoU. Defaults to False.
        SIoU (bool, optional): If True, calculate Scylla IoU. Defaults to False.
        eps (float, optional): A small value to avoid division by zero. Defaults to 1e-7.
    Returns:
        (torch.Tensor): IoU, GIoU, DIoU, or CIoU values depending on the specified flags.
    """
 
    if xywh:  # transform from xywh to xyxy
        (x1, y1, w1, h1), (x2, y2, w2, h2) = box1.chunk(4, -1), box2.chunk(4, -1)
        w1_, h1_, w2_, h2_ = w1 / 2, h1 / 2, w2 / 2, h2 / 2
        b1_x1, b1_x2, b1_y1, b1_y2 = x1 - w1_, x1 + w1_, y1 - h1_, y1 + h1_
        b2_x1, b2_x2, b2_y1, b2_y2 = x2 - w2_, x2 + w2_, y2 - h2_, y2 + h2_
    else:  # x1, y1, x2, y2 = box1
        b1_x1, b1_y1, b1_x2, b1_y2 = box1.chunk(4, -1)
        b2_x1, b2_y1, b2_x2, b2_y2 = box2.chunk(4, -1)
        w1, h1 = b1_x2 - b1_x1, b1_y2 - b1_y1 + eps
        w2, h2 = b2_x2 - b2_x1, b2_y2 - b2_y1 + eps
 
    # Intersection area
    inter = (b1_x2.minimum(b2_x2) - b1_x1.maximum(b2_x1)).clamp_(0) * \
            (b1_y2.minimum(b2_y2) - b1_y1.maximum(b2_y1)).clamp_(0)
    # Union Area
    union = w1 * h1 + w2 * h2 - inter + eps
    # IoU
    iou = inter / union
 
    if CIoU or DIoU or GIoU or EIoU or SIoU or WIoU:
        cw = b1_x2.maximum(b2_x2) - b1_x1.minimum(b2_x1)  # convex (smallest enclosing box) width
        ch = b1_y2.maximum(b2_y2) - b1_y1.minimum(b2_y1)  # convex height
        if CIoU or DIoU or EIoU or SIoU  or WIoU :  # Distance or Complete IoU https://arxiv.org/abs/1911.08287v1
            c2 = cw ** 2 + ch ** 2 + eps  # convex diagonal squared
            rho2 = ((b2_x1 + b2_x2 - b1_x1 - b1_x2) ** 2 + (b2_y1 + b2_y2 - b1_y1 - b1_y2) ** 2) / 4  # center dist ** 2
            if CIoU:  # https://github.com/Zzh-tju/DIoU-SSD-pytorch/blob/master/utils/box/box_utils.py#L47
                v = (4 / math.pi ** 2) * (torch.atan(w2 / h2) - torch.atan(w1 / h1)).pow(2)
                with torch.no_grad():
                    alpha = v / (v - iou + (1 + eps))
                return iou - (rho2 / c2 + v * alpha)  # CIoU
            elif EIoU:
                rho_w2 = ((b2_x2 - b2_x1) - (b1_x2 - b1_x1)) ** 2
                rho_h2 = ((b2_y2 - b2_y1) - (b1_y2 - b1_y1)) ** 2
                cw2 = cw ** 2 + eps
                ch2 = ch ** 2 + eps
                return iou - (rho2 / c2 + rho_w2 / cw2 + rho_h2 / ch2) # EIoU
            elif SIoU:
                # SIoU Loss https://arxiv.org/pdf/2205.12740.pdf
                s_cw = (b2_x1 + b2_x2 - b1_x1 - b1_x2) * 0.5 + eps
                s_ch = (b2_y1 + b2_y2 - b1_y1 - b1_y2) * 0.5 + eps
                sigma = torch.pow(s_cw ** 2 + s_ch ** 2, 0.5)
                sin_alpha_1 = torch.abs(s_cw) / sigma
                sin_alpha_2 = torch.abs(s_ch) / sigma
                threshold = pow(2, 0.5) / 2
                sin_alpha = torch.where(sin_alpha_1 > threshold, sin_alpha_2, sin_alpha_1)
                angle_cost = torch.cos(torch.arcsin(sin_alpha) * 2 - math.pi / 2)
                rho_x = (s_cw / cw) ** 2
                rho_y = (s_ch / ch) ** 2
                gamma = angle_cost - 2
                distance_cost = 2 - torch.exp(gamma * rho_x) - torch.exp(gamma * rho_y)
                omiga_w = torch.abs(w1 - w2) / torch.max(w1, w2)
                omiga_h = torch.abs(h1 - h2) / torch.max(h1, h2)
                shape_cost = torch.pow(1 - torch.exp(-1 * omiga_w), 4) + torch.pow(1 - torch.exp(-1 * omiga_h), 4)
                return iou - 0.5 * (distance_cost + shape_cost) + eps # SIoU
            elif WIoU:
                self = WIoU_Scale(1 - iou)
                dist = getattr(WIoU_Scale, '_scaled_loss')(self)
                return iou * dist  # WIoU https://arxiv.org/abs/2301.10051
            return iou - rho2 / c2  # DIoU
        c_area = cw * ch + eps  # convex area
        return iou - (c_area - union) / c_area  # GIoU https://arxiv.org/pdf/1902.09630.pdf
    return iou  # IoU

3 模型修改

第一步：利用以上代码替换ultralytics\utils\metrics.py中的bbox_iou全部代码

第二步：在ultralytics\utils\loss.py中ctrl+F搜索bbox_iou，大概76行左右，此时修改如下：采用SIOU损失函数

        iou = bbox_iou(pred_bboxes[fg_mask], target_bboxes[fg_mask], xywh=False, SIoU=True)

采用WIOU损失函数

        iou = bbox_iou(pred_bboxes[fg_mask], target_bboxes[fg_mask], xywh=False, WIoU=True)

采用EIOU损失函数

        iou = bbox_iou(pred_bboxes[fg_mask], target_bboxes[fg_mask], xywh=False, EIoU=True)

采用GIOU损失函数

        iou = bbox_iou(pred_bboxes[fg_mask], target_bboxes[fg_mask], xywh=False, GIoU=True)

采用DIOU损失函数

        iou = bbox_iou(pred_bboxes[fg_mask], target_bboxes[fg_mask], xywh=False, DIoU=True)

采用CIOU损失函数

        iou = bbox_iou(pred_bboxes[fg_mask], target_bboxes[fg_mask], xywh=False, CIoU=True)

4. 模型训练

5. 总结和展望

综上所述，这些IoU变种在目标检测领域各有优势和应用场景。SIoU通过考虑角度和距离来加速网络收敛；WIoU通过引入权重系数来提高检测精度；GIoU关注重叠和非重叠区域以更好地反映框的相交情况；DIoU通过最小化中心点距离来加速收敛；EIoU旨在提高效率并保持精度；而CIoU则通过考虑长宽比来优化预测框的形状和位置。在实际应用中，可以根据具体需求和数据集特性选择合适的IoU变种来优化目标检测模型的性能。

繁星似锁

关注

18
点赞
踩
20

收藏

觉得还不错? 一键收藏
0
评论
RTDETR结合SIoU、GIoU、CIoU、DIoU、WIoU比较和涨点

SIoU考虑了真实框与预测框之间的角度、距离以及形状，其损失函数由Angle cost、Distance cost、Shape cost和IoU cost四部分组成。这使得SIoU能够更快地收敛并防止网络在最优解附近跳动。然而，SIoU在不同数据集上的效果可能不一，且角度、距离以及形状的损失权重需要进一步优化。
复制链接

扫一扫