Reads, Shreejal Trivedi
Category: Yolo, Object Detection
First published: VisionWizard, Medium
S. Trivedi
28 May 2020

YOLOv4 — Version 4: Final Verdict

Summary

An Introductory Guide on the Fundamentals and Algorithmic Flows of YOLOv4 Object Detector

Table of Contents

  1. 1. Finalizing Bag of Freebies attributes2
  2. 2. Finalizing Bag of Specials attributes2
  3. 3. Effect of BoF + BoS on Training MiniBatch Size2
  4. 4. Results of Backbone + Neck + Head2
  5. 5. Comparison of mAP and FPS with different Object Detectors2
  6. 6. References2
  7. 7. Connections2
TrivediInformational[Page 1]
S. TrivediMay 2020
Source: Photo by Joanna Kosinska on Unsplash
Source: Photo by Joanna Kosinska on Unsplash

Welcome to the final part of YOLOv4¹ mini-series.

YOLOv4 — Version 0: Introduction

YOLOv4 — Version 1: Bag of Freebies

YOLOv4 — Version 2: Bag of Specials

YOLOv4 — Version 3: Proposed Workflow

YOLOv4 — Version 4: Final Verdict

I hope we were able to do a thorough walk through of all nuts and bolts of this amazing research.

This article’s main focus is on analytical results rather than any informative explanations. One last ride, let’s begin the finale.

This article will state the analytical comparisons between yolov4 and other object detectors.

1. Finalizing Bag of Freebies attributes

  • As discussed in the introduction of this series, many candidates were taken into consideration and were finalized to a small subset of them. We can analyze from the given tables below, how it affects the accuracy of the model.
  • Results given below in the form of tables are self-explanatory. The number’s speak for itself.
Bag of FreebiesUsed at training timeBackbone1. Class Label Smoothing2. Mosaic and CutMix Data Augmentations.3. DropBlock RegularizationDetector1. CIoU-loss2. Cross Minibatch Normalization.3. DropBlock regularization4. Mosaic data augmentation5. Self-Adversarial Training
Fig. 1 Final Candidates for Bag of Freebies Drawn by drawings/yolov4.py
  • S: Eliminate grid sensitivity the equation bx = σ(tx) + cx, by = σ(ty) + cy, where cx and cy are always whole numbers, is used in YOLOv3 for evaluating the object coordinates, therefore, extremely high tx absolute values are required for the bx value approaching the cx or cx + 1 values. We solve this problem through multiplying the sigmoid by a factor exceeding 1.0, so eliminating the effect of grid on which the object is undetectable.
  • M: Mosaic data augmentation - using the 4-image mosaic during training instead of single image
  • IT: IoU threshold - using multiple anchors for a single ground truth IoU (truth, anchor) > IoU_threshold
  • GA: Genetic algorithms - using genetic algorithms for selecting the optimal hyperparameters during network training on the first 10% of time periods
  • LS: Class label smoothing - using class label smoothing for sigmoid activation
  • CBN: CmBN - using Cross mini-Batch Normalization for collecting statistics inside the entire batch, instead of collecting statistics inside a single mini-batch
  • CA: Cosine annealing scheduler - altering the learning rate during sinusoid training
  • DM: Dynamic mini-batch size - automatic increase of mini-batch size during small resolution training by using Random training shapes
  • OA: Optimized Anchors - using the optimized anchors for training with the 512x512 network resolution
  • GIoU, CIoU, DIoU, MSE - using different loss algorithms for bounded box regression

Fig. 2: Bag of Freebies Abbreviations([1])

1.1 Results of BoF + Detector

SMITGALSCBNCADMOAlossAPAP50AP75
MSE38.0%60.0%40.8%
✓MSE37.7%59.9%40.5%
✓MSE39.1%61.8%42.0%
✓MSE36.9%59.7%39.4%
✓MSE38.9%61.7%41.9%
✓MSE33.0%55.4%35.4%
✓MSE38.4%60.7%41.3%
✓MSE38.7%60.7%41.9%
✓MSE35.3%57.2%38.0%
✓GIoU39.4%59.4%42.5%
✓DIoU39.1%58.8%42.1%
✓CIoU39.6%59.2%42.6%
✓✓✓✓CIoU41.5%64.0%44.8%
✓✓✓CIoU36.1%56.5%38.4%
✓✓✓✓✓MSE40.3%64.0%43.1%
✓✓✓✓✓GIoU42.4%64.4%45.9%
✓✓✓✓✓CIoU42.4%64.4%45.9%

Table. 1: Detector + BoF Ablation Study: Architecture used CSPResNext-50-SPP-PANet-512X512[1]

1.2 Results of BoF + Classifier

MixUpCutMixMosaicBluringLabel SmoothingSwishMishTop-1Top-5
77.9%94.0%
✓77.2%94.0%
✓78.0%94.3%
✓78.1%94.5%
✓77.5%93.8%
✓78.1%94.4%
✓64.5%86.0%
✓78.9%94.5%
✓✓✓78.5%94.8%
✓✓✓✓79.8%95.2%

Table. 2: BoF + Classifier Ablation Study. Architecture: CSPResNext-50[1].

MixUpCutMixMosaicBluringLabel SmoothingSwishMishTop-1Top-5
77.2%93.6%
✓✓✓77.8%94.4%
✓✓✓✓78.7%94.8%

Table.3 BoF + Classifier Ablation Study. Architecture: CSPDarkNet-53[1].

2. Finalizing Bag of Specials attributes

Bag of SpecialsUsed at inference timeBackbone1. Mish Activation2. Cross Stage Partial Connections3. Multi input Weighted Residual ConnectionsDetector1. Mish Activation2. Spatial Pyramid Pooling3. Spatial Attention Module4. Path Aggregation Networks5. DIoU-NMS
Fig. 3 Bag of Specials Finalized Candidates Drawn by drawings/yolov4.py
ModelAPAP50AP75
CSPResNeXt50-PANet-SPP42.4%64.4%45.9%
CSPResNeXt50-PANet-SPP-RFB41.8%62.7%45.1%
CSPResNeXt50-PANet-SPP-SAM42.7%64.6%46.3%
CSPResNeXt50-PANet-SPP-SAM-G41.6%62.7%45.0%
CSPResNeXt50-PANet-SPP-ASFF-RFB41.1%62.6%44.4%

Table. 4: Ablation Study of BoS on CSPResNext-50 architecture[1].

  • Mish and DIoU-NMS are taken into consideration during the inference stage.

3. Effect of BoF + BoS on Training MiniBatch Size

Model (without OA)SizeAPAP50AP75
CSPResNeXt50-PANet-SPP (without BoF/BoS, mini-batch 4)60837.159.239.9
CSPResNeXt50-PANet-SPP (without BoF/BoS, mini-batch 8)60838.460.641.6
CSPDarknet53-PANet-SPP (with BoF/BoS, mini-batch 4)51241.664.145.0
CSPDarknet53-PANet-SPP (with BoF/BoS, mini-batch 8)51241.764.245.2

Table. 5: Ablation study of different batch sizes with and without using BoF+BoS[1].

4. Results of Backbone + Neck + Head

Model (with optimal setting)SizeAPAP50AP75
CSPResNeXt50-PANet-SPP512x51242.464.445.9
CSPResNeXt50-PANet-SPP (BoF-backbone)512x51242.364.345.7
CSPResNeXt50-PANet-SPP (BoF-backbone + Mish)512x51242.364.245.8
CSPDarknet53-PANet-SPP (BoF-backbone)512x51242.464.546.0
CSPDarknet53-PANet-SPP (BoF-backbone + Mish)512x51243.064.946.5

Table. 6: Selection of Backbone, Neck, and Head of the detector. This table shows that Backbone: CSPDarkNet53, Neck: PAN + SPP, and Head: YOLOv3 outperforms all the models that were taken into consideration (BoF included)[1].

  • As discussed in the first article of these series, the architecture of CSPDarknet53 was proved to be most optimal in terms of the receptive field, FPS, FLOPs, etc.
  • After leveraging the techniques from Bag of Specials and Bag of Freebies with the given backbone CSPDarknet53 proved to give the best results of 43% AP on COCO test-2017.

5. Comparison of mAP and FPS with different Object Detectors

Fig. 4 Graph showing mAP vs FPS results for different object detectors.
Fig. 4 Graph showing mAP vs FPS results for different object detectors.

Yolov4 state-of-the-art detector which is faster (FPS)and more accurate (MS COCO AP[50…95] and AP50) than all available alternative real time detectors.

The original concept of one-stage anchor-based detectors has proven its viability. We have verified a large number of features, and selected for use such of them for improving the accuracy of both the classifier and the detector. [1]

Here ends the final part of this series. I know, it was a long journey, but I hope you now have a good grip in terms of the algorithmic perspective of YOLOv4.

Please check out other parts of our entire series on Yolov4 on our page VisionWizard.

It looks like you have a real interest in quality research work if you have reached till here. If you like our content and want more of it, please follow our page VisionWizard.

Do clap if you have learned something interesting and useful. It would motivate us to curate more quality content for you guys.

Thanks for your time :).

6. References

[1] YOLOv4

7. Connections

What this read builds on, and what builds on it. Hover a box for a preview.

This pageBUILDS ON THISTHIS BUILDS ONREADYOLOv4 — Version 2: B…PAPERYOLOv4: Optimal Speed…

Related

Go deeper

  1. Start with
    Read
  2. Read
  3. Read
  4. You are here
    YOLOv4 — Version 4: Final Verdict
    Read
<- Older [Page 2] Newer ->