【目标检测】YOLOX网络结构基础模块详解及代码实现

行走的学习机器

已于 2024-01-04 16:46:43 修改

阅读量309

点赞数

文章标签：深度学习人工智能

于 2023-10-10 10:19:22 首次发布

本文链接：https://blog.csdn.net/weixin_44984705/article/details/133739332

版权

文章目录

前言
一、FOCUS模块
二、BaseConv
三、Bottleneck

前言

YOLOX可以看成是各基础模块堆叠而成，明白各基础模块如何形成，代码如何实现，就能搭建出网络结构。

YOLOX网络结构如下图所示。
在这里插入图片描述
图片转载自链接: 太阳花的小绿豆博客

一、FOCUS模块

在这里插入图片描述
代码实现：

class Focus(nn.Module):
    """Focus width and height information into channel space."""

    def __init__(self, in_channels, out_channels, ksize=1, stride=1, act="silu"):
        super().__init__()
        self.conv = BaseConv(in_channels * 4, out_channels, ksize, stride, act=act)

    def forward(self, x):
        # shape of x (b,c,w,h) -> y(b,4c,w/2,h/2)
        patch_top_left = x[..., ::2, ::2]
        patch_top_right = x[..., ::2, 1::2]
        patch_bot_left = x[..., 1::2, ::2]
        patch_bot_right = x[..., 1::2, 1::2]
        x = torch.cat(
            (
                patch_top_left,
                patch_bot_left,
                patch_top_right,
                patch_bot_right,
            ),
            dim=1,
        )
        return self.conv(x)

结合代码及图片可知，FOCUS是将原本的feature map分成四份，再从第一维度将其concat到一起后，再做一次卷积。

二、BaseConv

现在卷积层+BN+激活函数几乎成了标配。

class BaseConv(nn.Module):
    """A Conv2d -> Batchnorm -> silu/leaky relu block"""

    def __init__(
        self, in_channels, out_channels, ksize, stride, groups=1, bias=False, act="silu"
    ):
        super().__init__()
        # same padding
        pad = (ksize - 1) // 2
        self.conv = nn.Conv2d(
            in_channels,
            out_channels,
            kernel_size=ksize,
            stride=stride,
            padding=pad,
            groups=groups,
            bias=bias,
        )
        self.bn = nn.BatchNorm2d(out_channels)
        self.act = nn.SiLU(inplace=inplace)

    def forward(self, x):
        return self.act(self.bn(self.conv(x)))

三、Bottleneck

class Bottleneck(nn.Module):
    # Standard bottleneck
    def __init__(
        self,
        in_channels,
        out_channels,
        shortcut=True,
        expansion=0.5,
        depthwise=False,
        act="silu",
    ):
        super().__init__()
        hidden_channels = int(out_channels * expansion)
        Conv = DWConv if depthwise else BaseConv
        self.conv1 = BaseConv(in_channels, hidden_channels, 1, stride=1, act=act)
        self.conv2 = Conv(hidden_channels, out_channels, 3, stride=1, act=act)
        self.use_add = shortcut and in_channels == out_channels

    def forward(self, x):
        y = self.conv2(self.conv1(x))
        if self.use_add:
            y = y + x
        return y