Reads, Shreejal Trivedi
Category: Object Detection
First published: VisionWizard, Medium
S. Trivedi
16 Jul 2020

DetectoRS — A Comprehensive Review

Summary

A detailed breakdown of a new paper from Google Research and Johns Hopkins University — DetectoRS: Detecting Objects using Recursive Feature Pyramids and Switchable Atrous Convolutions.

Table of Contents

  1. 1. Introduction2
  2. 2. Proposed Workflow2
  3. 3. Results2
  4. 4. Conclusion2
  5. 5. References2
  6. 6. Connections2
TrivediInformational[Page 1]
S. TrivediJul 2020

There has been extensive research going on in finding novel techniques, algorithms, and new end-to-end trainable pipelines for Object Detection and Image Segmentation tasks in the field of Computer Vision.

Year by year, different research institutes/organizations come up with some new ideas to tackle the pertaining problems of these tasks with robust and real-time solutions.

Given the premise, 2020 is being one of the most exciting years for the field of Computer Vision and Deep Learning.

In June 2020, Google Research Team with Johns Hopkins University came up with an exciting new architecture design of Feature Pyramid Networks and modulation of existing convolutional operation chains in the paper DetectoRS: Detecting Objects using Recursive Feature Pyramids and Switchable Atrous Convolutions[1] that can be directly leveraged into the present neural network backbones such as ResNets, ResNeXts, etc.

This paper claimed to have a state-of-the-art performance on COCO Object Detection/Image Segmentation Dataset with an mAP of 54.7% and mask AP of 47.1% respectively.

In this article, we will go through the critical baselines mentioned in the paper. Let’s jump right in.

1. Introduction

  • Recently, the coupling of thinking and looking twice mechanism in state-of-the-art two-stage object detectors has proved to give promising results in image segmentation and object detection tasks.
  • Object detectors such as Faster RCNN first detects the object proposals based on the foreground-background classification, which are then used to extract rich semantic features for further classification/detection tasks in later stages.
  • Following thinking and looking twice regime, Cascade RCNN is designed to refine these obtained proposals from one stage and enrich these proposals with their multi-stage design, as shown in Fig. 1. As we continue to go deep inside the network, subsequent head classifiers/detectors are trained on more enhanced object features.
Faster R-CNNIconvH0C0B0poolH1C1B1Cascade R-CNNIconvB0poolH1C1B1poolH2C2B2poolH3C3B3
Fig. 1 (Left) Faster RCNN. (Right) Cascade RCNN(Source: Link). Drawn by drawings/detectors.py
  • Inspired by this, the authors introduced a novel approach to leverage this mechanism directly in the network backbone.
  • They consider the Hybrid Task Cascade network as their baseline architecture for comparison studies.

HTC is a version up of Cascade Faster-RCNN, but have an extra module for image segmentation mask. You can visualize it from Fig. 2.

Cascade Mask R-CNNFRPNpoolM1B1poolM2B2poolM3B3Hybrid Task CascadeFSRPNpoolpoolpoolpoolB1B2B3M1M2M3blue: mask information flow rust: semantic features
Fig. 2 (Left) Cascade Mask RCNN. (Right) Hybrid Task Cascade Architectural Design(Source: Link). Green boxes are the object masks. HTC has different information flow for improving the segmentation masks and bounding boxes with its interleaved design, unlike Cascade Mask RCNN, wherein masks are parallelly calculated at every stage. Drawn by drawings/detectors.py

The proposition of Switchable Atrous Convolutions and Recursive Feature Pyramid Networks at the macro(Neck) and micro(Backbone Convolutions) levels of the network respectively significantly improved the state-of-the-art network HTC with almost the same inference time[1].

MethodBackboneAP boxAP maskFPS
HTCResNet-5043.638.54.3
DetectoRSResNet-5051.344.43.9

Table. 1 Inference and mAP number comparison between HTC and DetectoRS.

2. Proposed Workflow

  • The authors were more focused on the present refinements of the convolutions operations and Feature Pyramid Networks to incorporate the stated regime in current models.
  • Recursive Feature Pyramid Networks and Switchable Atrous Convolution are two significant findings mentioned in the paper.
  • At the macro level, they introduced a new version of Feature Pyramids, by connecting the top-down level path predictions as a feedback loop to the bottom-up path, as shown in Fig. 3.
Fig.3 Feedback prediction feature maps to bottom-up layers(Source: [1]).
Fig.3 Feedback prediction feature maps to bottom-up layers(Source: [1]).

Unrolling the recursive structure to a sequential implementation, we obtain a backbone for an object detector that looks at the images twice or more. Similar to the cascaded detector heads in Cascade R-CNN trained with more selective examples, our RFP recursively enhances FPN to generate increasingly powerful representations[1].

  • At the micro-level, they propose SAC(Switchable Atrous Convolutions), which takes feature maps as input and convolves with different atrous rates based on the global context information and combines the output using the switch functions as shown in Fig. 4.
Fig. 4 Switch function for different atrous rates(Source: [1]).
Fig. 4 Switch function for different atrous rates(Source: [1]).

The switch functions are spatially dependent, i.e., each location of the feature map might have different switches to control the outputs of SAC. To use SAC in the detector, we convert all the standard 3x3 convolutional layers in the bottom-up backbone to SAC, which improves the detector performance by a large margin[1].

Recursive Feature Pyramid Networks

Original FPNBottom UpLayersx0B1x1B2x2B3x3B4x4B5x5F5F4F3F2PredictionStagef5f4f3f2Top DownFeatureMapsBottom Up Feature MapsBottom Up Downsample LayersTop Down Downsample LayersTop Down Feature Maps
Fig. 5 Structure of original FPN. Drawn by drawings/detectors.py
  • Let’s revisit the structure of the basics of Feature Pyramid Networks. Here, B(i) denotes the bottom-up downsampling layer chains. F(i) is the top-down FPN operation.
  • B(i) takes x(i-1) high resolution feature map. and convert it into low-resolution alternative x(i), as shown in Fig. 5.
  • Also, prediction feature maps f(i) are passed to the prediction head for further classification and regression. Top-down layer F(i) takes input the bottom-up feature map x(i) and low-resolution alternative f(i+1). All f(i)s are the set of prediction feature maps. It can be generalized by the equation as:-
fi=Fi(fi+1,xi),xi=Bi(xi−1)(1)

Source: [1]

  • Now, in Recursive Feature Pyramids, as the name depicts, each FPN layer is stacked T times. Output prediction maps f(i) in Fig 5. are used as feedback loops in the bottom-up downsampling layer chains B(i)(The feedback loop structure can be visualized from Fig. 3). Let’s magnify that image to see how the unrolling happens in RFP.
Recursive Feature Pyramid Networkt = 1x0B1x1B2x2B3x3B4x4B5x5F5F4F3F2Same Backbonet = 2x0B1x1B2x2B3x3B4x4B5x5F5F4F3F2ASPPf5FusionASPPf4FusionASPPf3FusionASPPf2Fusion
Fig. 6 Two-stage structure of Recursive Feature Pyramid. Drawn by drawings/detectors.py
  • The above figure shows the two-stage(t = 2) unrolling of the Recursive Feature Pyramids. Blue shaded region is the same backbone(or different backbone. We will see this in the implementation details section) used for the first stage bottom-up FPN operation.
  • Before passing the prediction feature maps f(i) to the backbone again, they are passed through the ASPP Module as shown. This is called transformation R(i) on feature map f(i). We can generalize this network design for different values of stacking stage t with the below-given equation.
fit=Fit(fi+1t,xit),xit=Bit(xi−1t,Rit(fit−1))(2)
  • Here, t ∈[2, 3, …, T] is the number of unrolling operation to be done. By default, T = 2.
  • There are two main pre-requisites or transformation parts used in RFP viz. ASPP(Atrous Spatial Pyramid Pooling) and Fusion Module.

ASPP(or Atrous Spatial Pyramid Pooling) module does the same thing to increase the receptive field of the network, but instead of using standard convolutions of different kernel sizes, dilated convolutions with different atrous rates are used. Its implementation in RFP can be visualized from the below-given figure.

Atrous Spatial Pyramid Pooling (ASPP) Modulef(i)(C, H, W)cConcatenationChannelwiseR(f(i))(C, H, W)Kernel Size = 1Dilation Value = 1Padding = 0Out Channels = C/4Kernel Size = 3Dilation Value = 3Padding = 3Out Channels = C/4Kernel Size = 3Dilation Value = 3Padding = 3Out Channels = C/4Global AveragePoolingKernel Size = 1, Dilation Value = 1Padding = 0, Out Channels = C/4Expand Spatially
Fig. 7 ASPP Module Design Drawn by drawings/detectors.py
  • ASPP module output features are directly passed to the first residual block of every layer of the ResNet. It takes both x(i) and R(f(i)) as an input to the residual block. Visualization is given below.
InputConv(1x1)Conv(3x3, s=2)Conv(1x1)Conv(1x1, s=2)RFP FeaturesConv(1x1)OutputResNetRFP
Fig. 8 Gray part is the default first residual block of ResNet. RFP Features are the outputs of the ASPP Module(Source: [1]). Drawn by drawings/detectors.py

The weight of RFP 1x1 convolutional layer is initialized with 0 to make sure it does not have any real effect when pretrained weights are used for training downstream tasks[1]. RFP Features are only applied on t = [2, …, T] stages. First stage is a default FPN operation.

  • The different combinations of feature fusion and attention layer techniques have performed very well in the deep networks to detect the varied objects at different scales with a decent confidence score.
  • The authors, to further improvise the information of the maps added Fusion Module that uses the top-down FPN operation feature maps(f(i)s as shown in Fig. 6) from different stages of FPN and combines them together with an implicit attention gate in their design. Visualization is given below.
f(i) at t+1f(i) at tConv(1x1)Sigmoidσ1 − σOutput
Fig. 9 Feature fusion with implicit attention gate mechanism(Source: [1]). Drawn by drawings/detectors.py

Feature fusion module is only applied at the unrolled steps from t=[2, …, T]. Module takes the present newly obtained top-down operation feature map f(i+1) at an unrolled step t + 1 and f(i) at t respectively. Sigmoid function generates an intermediary attention map for the present features and are used in the aggregation of both the inputs. Output of this module is the more focused feature map f(i+1) at unrolled step t + 1

What is the logic behind a cascaded FPN(or Recursive Feature Pyramid) structure?

  • We all know that as we go deep in the network, semantic information about the objects increases and the system tries to learn well while keeping a global context in mind. Also, earlier layers are crucial to learning low-level information such as edges, textures giving proper discriminative features.
  • Cascading and connecting the output maps to the backbone structures(Fig. 8) further at the unrolled stages can help to capture the missed information about the objects by keeping in mind the receptive field and context using ASPP, and remove the unwanted activated pixels values with the help of implicit attention layer in Fusion Stage.

This is the official RFP code on their repository. You can visit the code structure for getting implementation insights.

joe-siyuan-qiao/DetectoRS

Switchable Atrous Convolutions

  • Atrous(or Dilated) Convolutions are the operation that increases the field of view without any additional computations.
1-dilated convolutionsees 3 x 32-dilated convolutionsees 7 x 74-dilated convolutionsees 15 x 15
Fig. 10 Dilation value increasing going from left to right. With the increase in dilation-value increases the field of view(or receptive field) exponentially. Drawn by drawings/common.py
  • The receptive field is an essential concept in Convolutional Neural Networks. A large kernel size can increase the output context information of the feature map with the cost of increasing the computations.

Atrous convolution with atrous rate r introduces r − 1 zero between consecutive filter values, equivalently enlarging the kernel size of a k × k filter to knew = k + (k − 1)(r − 1) without increasing the number of parameters or the amount of computation[1].

  • As stated, a 3x3 convolutional kernel can give the same receptive as 5x5 with fixed computational multiplications by keeping atrous rate r = 2.
  • By using dilated convolutions with higher atrous rates can sometimes miss capturing the local information of the objects at smaller scales.
  • Conditional convolutions such as Selective Kernel Networks, Dynamic Convolutions, etc. are some of the present and well-known implementations of the dynamic, receptive field adjustment based on the scales and context of the objects. But these implementations are somewhat ineffective while using the pretrained network architectures.
  • So, the authors introduced a new concept of the conditional convolutions that can be used as a to-go choice for pretrained backbone networks.
InputPre-Global ContextGlobalAvgPoolConv(1x1)Switchable Atrous ConvolutionConv(3x3, atrous=1)Conv(3x3, atrous=3)AvgPool(5x5)Conv(1x1)sharedweightsS1-SPost-Global ContextGlobalAvgPoolConv(1x1)Output
Fig. 11 Breaking of standard 3x3 convolution operation into switchable atrous convolutions of rates r=1 and r=3, respectively. Weights are shared between two convolutions(Locking mechanism)(Source: [1]) Drawn by drawings/detectors.py
  • They also introduced a very intuitive switch and context mechanism in the convolution overview(shown in Fig. 11) to generate effective results.
ResNet Residual Block DesignBottleneck w/o Downsample LayerInput(W1, R, R)1x1 Conv+BN+ReLUbottleneckratio b3x3 Conv+BN+ReLUgroupsize g1x1 Conv+BN+ReLUAdd+ReLUOutput(W1, R, R)Bottleneck with Downsample LayerInput(W1, R, R)1x1 Conv+BN+ReLUbottleneckratio b3x3 Conv+BN+ReLUgroupsize g1x1 Conv+BN+ReLUAdd+ReLUOutput(W1, R/2, R/2)1x1 Conv+BN, Stride = 2
Fig. 12 Residual block with the bottleneck. Drawn by drawings/common.py
  • All 3x3 convolution operations in the residual blocks of the ResNet are changed to dilated convolutions with two different atrous rates, as shown in Fig. 11.

Both the convolution operations shares the same weight w.

Conv(x,w,1)→Convert to SACS(x)·Conv(x,w,1)+(1−S(x))·Conv(x,w+Δw,r)(3)

The above equation shows the switching mechanism between two convolution operations. Both share the same weights with a different trainable δw(Source: [1]).

  • S(x) is a switch function that consists of an Average Pooling Layer of 5x5 kernel followed by a 1x1 convolution block. This switch can gather some statistics of the feature map using pooling layers, which can be further used to detect objects at different scales.

A value obtained from the patch of an average pooling operation is a decision-maker to increase the field of view or not. If the pooling operation does not activate the feature map pixels then a convolution operation with a large field of view(or dilated convolution) might be more canonical to obtain better statistics of the object if present.

  • The global context mechanism is also added before and after each of the 3x3 blocks. These are lightweight 1x1 convolution kernels and a global average pooling layer. They state that using global information is more efficient for a switch function S(x) to stabilize the switching between two convolution operations.

Object detectors usually use pretrained checkpoints to initialize the weights. However, for a SAC layer converted from a standard convolutional layer, the weight for the larger atrous rate is missing. Since objects at different scales can be roughly detected by the same weight with different atrous rates, it is natural to initialize the missing weights with those in the pretrained model[1].

∇w is compensation parameter for the atrous convolutions and is initialized to zero. This weight differences are the small precision values for the original pretrained weights for atrous convolutions.

You can get a proper flow of SAC from this code.

joe-siyuan-qiao/DetectoRS

3
Atrous convolution. Filled cells are the nine pixels a 3×3 kernel reads at this rate. The outlined cell is the centre. Computed by figures/atrous.py.
The Python behind this figure

3. Results

  • The authors took the Hybrid Task Cascade framework as a baseline and added the above two novel designs to obtain DetectoRS.
  • Backbone architectures considered were the ResNet family variant such as ResNet-50 and ResNeXt-101–32x4d. Also, all the 3x3 convolution operations were replaced by Deformable Convolutions.
  • The offset values of the DCN and weight difference ∇w were initialized to 0.
  • The weights and biases of the Pre and Post Context 1x1 convolutional kernels were initialized to zeros and ones, respectively.
APAP50AP75APSAPMAPL
Baseline HTC42.060.845.523.745.556.4
RFP46.265.150.227.950.360.3
RFP + sharing45.464.149.426.549.060.0
RFP - aspp45.764.249.626.749.360.5
RFP - fusion45.964.750.027.050.160.1
RFP + 3X47.566.351.829.051.661.9
SAC46.365.850.227.850.662.4
SAC - DCN45.365.049.327.548.760.6
SAC - DCN - global44.363.748.225.748.059.6
SAC - DCN - locking44.764.448.726.048.759.0
SAC - DCN + DS45.164.649.026.349.360.1

Table. 1 Ablation Study for different variations of the parameters

  • We can see from the above table that different backbone weights(weights B(t=1) ≠weights B(t=2)) and three-stage unrolling(t=3) gave the best results.
  • Also, by using the locking mechanism of the weights between two atrous convolutions gave better results rather than just using ∇w for training.
Fig. 13 Training Loss comparison between HTC baseline and their novel alternatives.
Fig. 13 Training Loss comparison between HTC baseline and their novel alternatives.
MethodBackboneTTAAP bboxAP50AP75APSAPMAPL
YOLOv3DarkNet-5333.057.934.418.325.441.9
RetinaNetResNeXt-10140.861.144.124.144.251.2
RefineDetResNet-101✓41.862.945.725.645.154.1
CornerNetHourglass-104✓42.157.845.320.844.856.7
ExtremeNetHourglass-104✓43.760.547.024.146.957.6
FSAFResNeXt-101✓44.665.248.629.747.154.6
FCOSResNeXt-10144.764.148.427.647.555.6
CenterNetHourglass-104✓45.163.949.326.647.157.7
NAS-FPNAmoebaNet48.3-----
SEPCResNeXt-10150.169.854.331.353.363.7
SpineNetSpineNet-19052.171.856.535.455.063.6
EfficientDet-D7EfficientNet-B652.671.656.9---
Mask R-CNNResNet-10139.862.343.422.143.251.2
Cascade R-CNNResNet-10142.862.146.323.745.555.2
Libra R-CNNResNeXt-10143.064.047.025.345.654.6
DCN-v2ResNet-101✓46.067.950.827.849.159.5
PANetResNeXt-10147.467.251.830.151.760.0
SINPERResNet-101✓47.668.553.430.950.660.7
SNIPModel Ensemble✓48.369.753.731.451.660.7
TridentNetResNet-101✓48.469.753.531.851.360.3
Cascade Mask R-CNNResNeXt-152✓50.268.254.931.952.963.5
TSDSENet154✓51.271.956.033.854.864.2
MegDetModel Ensemble✓52.5-----
CBNetResNeXt-152✓53.371.958.535.555.866.7
HTCResNet-5043.662.647.424.846.055.9
HTCResNeXt-101-32x4d46.465.850.526.849.459.6
HTCResNeXt-101-64x4d47.266.551.427.750.160.3
HTC + DCN + mstrainResNeXt-101-64x4d50.870.355.231.154.164.8
DetectoRSResNet-5051.370.155.831.754.664.8
DetectoRSResNet-50✓53.072.257.835.955.664.6
DetectoRSResNeXt-101-32x4d53.371.658.533.956.566.9
DetectoRSResNeXt-101-32x4d✓54.773.560.137.457.366.4

Table 2. DetectoRS comparisons with different object detectors on COCO-test dev.

Official Github Code: [github]

4. Conclusion

  • The authors introduced the two novel techniques, viz. Switchable Atrous Convolutions and Recursive Feature Pyramid Networks.
  • They followed the regime of thinking and looking twice inspired from the network designs of Cascade RCNN, Hybrid Task Cascade, etc. by connecting the feedback loop in the unrolled cascade FPNs.
  • By leveraging the two novel methods with HTC as a baseline, it outperformed every present model in object detection, image, and panoptic segmentation with COCO mAP of 54.7%.

Different ablation studies conducted by the authors to find the legitness of the model spaces are skipped in this article. You can get to know about the different permutations of the network parameters from the paper. Image and Panoptic Segmentation are also not mentioned in this article as they were not intersecting with the scope of this article.

I hope you found the content meaningful. If you want to stay updated with the latest research in AI, please follow us at VisionWizard.

5. References

[1] DetectoRS: Detecting Objects with Recursive Feature Pyramid and Switchable Atrous Convolution

[2] Hybrid Task Cascade

[3] Deformable Convolution Networks

6. Connections

What this read builds on, and what builds on it. Hover a box for a preview.

This pageBUILDS ON THISTHIS BUILDS ONREADObject Detection — A …PAPERDetectoRS: Detecting …

Related

Go deeper

  1. Start with
    Read
  2. You are here
    DetectoRS — A Comprehensive Review
    Read
<- Older [Page 2] Newer ->