Complex overlapping pedestrian target detection network based on the yolov3 model
DOI:
https://doi.org/10.62051/rpbbxx55Keywords:
yolov3, Computer vision, pedestrian target detection, deep learning.Abstract
This paper proposes a complex overlapping pedestrian target detection model based on yolov3 model by multi-scale feature fusion and context-aware mechanism. The SONY A7R3a camera shot the model on campus, and the data set was obtained after editing and collating. There were 358 high-definition videos with a resolution of 1920*1080, and the frame rate was 50HZ, about 179,000 frames. Through testing, this paper finds that compared with Single Shot Multibox Detector (SSD), the detection accuracy of the newly proposed model is slightly improved, the detection accuracy is the same as that of Faster R-CNN, and the detection accuracy of the newly proposed model is slightly worse than that of RetinaNet. However, the detection speed of Yolov3 is more than twice that of Single Shot Multibox Detector, RetinaNet and Faster R-CNN. The input size of Yolov3 is 320*320, and the processing of a single image only needs 22ms, so the detection speed of the simplified Yolov3 tiny is faster.
Downloads
References
Sengupta A, Ye Y, Wang R, et al. Going deeper in spiking neural networks: VGG and residual architectures. Frontiers in neuroscience, 2019, 13: 95.
Targ S, Almeida D, Lyman K. Resnet in resnet: Generalizing residual architectures. arXiv preprint arXiv:1603.08029, 2016.
Szegedy C, Vanhoucke V, Ioffe S, et al. Rethinking the inception architecture for computer vision, Proceedings of the IEEE conference on computer vision and pattern recognition. 2016: 2818-2826.
Tan M, Le Q. Efficientnet: Rethinking model scaling for convolutional neural networks[C]//International conference on machine learning. PMLR, 2019: 6105-6114.
Girshick R. Fast r-cnn: Proceedings of the IEEE international conference on computer vision. 2015: 1440-1448.
Redmon J, Divvala S, Girshick R, et al. You only look once: Unified, real-time object detection, Proceedings of the IEEE conference on computer vision and pattern recognition. 2016: 779-788.
Liu W, Anguelov D, Erhan D, et al. Ssd: Single shot multibox detector[C]//Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part I 14. Springer International Publishing, 2016: 21-37.
Fan Q, Zhuo W, Tang C K, et al. Few-shot object detection with attention-RPN and multi-relation detector, Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2020: 4013-4022.
Girshick R, Donahue J, Darrell T, et al. Rich feature hierarchies for accurate object detection and semantic segmentation, Proceedings of the IEEE conference on computer vision and pattern recognition. 2014: 580-587.
Girshick R. Fast r-cn, Proceedings of the IEEE international conference on computer vision. 2015: 1440-1448.
Ren S, He K, Girshick R, et al. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems, 2015, 28.
Wang Y, Wang C, Zhang H, et al. Automatic ship detection based on RetinaNet using multi-resolution Gaofen-3 imagery. Remote Sensing, 2019, 11(5): 531.
Iandola F N, Han S, Moskewicz M W, et al. SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and< 0.5 MB model size. arXiv preprint arXiv:1602.07360, 2016.
Uijlings J R R, Van De Sande K E A, Gevers T, et al. Selective search for object recognition. International journal of computer vision, 2013, 104: 154-171.
Qassim H, Verma A, Feinzimer D. Compressed residual-VGG16 CNN model for big data places image recognition,2018 IEEE 8th annual computing and communication workshop and conference (CCWC). IEEE, 2018: 169-175.
Redmon J, Farhadi A. YOLO9000: better, faster, stronger,Proceedings of the IEEE conference on computer vision and pattern recognition. 2017: 7263-7271.
Hamerly G, Elkan C. Learning the k in k-means. Advances in neural information processing systems, 2003, 16.
You Y, Zhang Z, Hsieh C J, et al. Imagenet training in minutes,Proceedings of the 47th International Conference on Parallel Processing. 2018: 1-10.
Lin T Y, Maire M, Belongie S, et al. Microsoft coco: Common objects in context,Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13. Springer International Publishing, 2014: 740-755.
Redmon J, Farhadi A. Yolov3: An incremental improvement.arXiv preprint arXiv:1804.02767, 2018.
Bodla N, Singh B, Chellappa R, et al. Soft-NMS--improving object detection with one line of code,Proceedings of the IEEE international conference on computer vision. 2017: 5561-5569.
Tzutalin D. LabelImg. GitHub repository, 2015, 6.
Downloads
Published
Conference Proceedings Volume
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.







