| Citation: | WANG D W,HU L C,FANG J,et al. Small target detection algorithm based on improved Double-Head RCNN for UAV aerial images[J]. Journal of Beijing University of Aeronautics and Astronautics,2024,50(7):2141-2149 (in Chinese) |
The feature information of small targets in unmanned aerial vehicle aerial images is small and easily interfered with by noise, which leads to the high missed detection and false detection rates of existing algorithms. To address these issues, a small target detection algorithm based on an improved Double-Head region-convolutional neural networks(RCNN)for unmanned aerial vehicle aerial images was proposed. Transformer and deformable convolution networks (DCN) modules were introduced on the backbone network ResNet-50 to extract small target feature information and semantic information more effectively. A feature pyramid network(FPN) structure based on content-aware reassembly of features (CARAFE) was proposed to solve the problem that the small target information is interfered with by the background noise, and the feature information is lost in the process of feature fusion. The generation scale of Anchor was reset according to the characteristics of small target scale distribution in the region proposal network to further improve the small target detection performance. The experimental results on the VisDrone-DET2021 dataset show that the proposed algorithm can extract feature and semantic information of small targets with representational capacity more effectively. Compared with the Double-Head RCNN algorithm, the parameter quantity of the proposed algorithm increases by 9.73×106, and the FPS loss is 0.6. However, AP, AP50, and AP75 increase by 2.6%, 6.2%, and 2.1% respectively, and APs increases by 3.1%.
| [1] |
LIN T Y, MAIRE M, BELONGIE S, et al. Microsoft COCO: Common objects in context[C]//Proceedings of the European Conference on Computer Vision. Berlin: Springer, 2014: 740-755.
|
| [2] |
WANG J W, YANG W, GUO H W, et al. Tiny object detection in aerial images[C]//Proceedings of the International Conference on Pattern Recognition. Piscataway: IEEE Press, 2021: 3791-3798.
|
| [3] |
陈映雪, 丁文锐, 李红光, 等. 基于视频帧间运动估计的无人机图像车辆检测[J]. 北京航空航天大学学报, 2020, 46(3): 634-642.
CHEN Y X, DING W R, LI H G, et al. Vehicle detection in UAV image based on video interframe motion estimation[J]. Journal of Beijing University of Aeronautics and Astronautics, 2020, 46(3): 634-642(in Chinese).
|
| [4] |
WANG W H, XIE E Z, LI X, et al. PVT v2: Improved baselines with pyramid vision transformer[J]. Computational Visual Media, 2022, 8(3): 415-424.
|
| [5] |
CARION N, MASSA F, SYNNAEVE G, et al. End-to-end object detection with transformers[C]//Proceedings of the European Conference on Computer Vision. Berlin: Springer, 2020: 213-229.
|
| [6] |
ZHU X Z, SU W J, LU L W, et al. Deformable DETR: Deformable transformers for end-to-end object detection[EB/OL]. (2021-03-18)[2022-05-18].
|
| [7] |
LIN T Y, GOYAL P, GIRSHICK R, et al. Focal loss for dense object detection[C]//Proceedings of the IEEE International Conference on Computer Vision. Piscataway: IEEE Press, 2017: 2999-3007.
|
| [8] |
GE Z, LIU S T, WANG F, et al. YOLOX: Exceeding YOLO series in 2021[EB/OL]. (2021-08-06)[2022-05-20].
|
| [9] |
WANG C Y, BOCHKOVSKIY A, LIAO H Y M. YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors[EB/OL]. (2022-07-01)[2022-07-03].
|
| [10] |
REN S Q, HE K M, GIRSHICK R, et al. Faster R-CNN: Towards real-time object detection with region proposal networks[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 39(6): 1137-1149.
|
| [11] |
PANG J M, CHEN K, SHI J P, et al. Libra R-CNN: Towards balanced learning for object detection[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2019: 821-830.
|
| [12] |
CAI Z W, VASCONCELOS N. Cascade R-CNN: Delving into high quality object detection[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2018: 6154-6162.
|
| [13] |
LU X, LI B Y, YUE Y X, et al. Grid R-CNN[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2019: 7355-7364.
|
| [14] |
GIRSHICK R. Fast R-CNN[C]//Proceedings of the IEEE International Conference on Computer Vision. Piscataway: IEEE Press, 2015: 1440-1448.
|
| [15] |
LIN T Y, DOLLÁR P, GIRSHICK R, et al. Feature pyramid networks for object detection[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2017: 936-944.
|
| [16] |
LIU S, QI L, QIN H F, et al. Path aggregation network for instance segmentation[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2018: 8759-8768.
|
| [17] |
TAN M X, PANG R M, LE Q V. EfficientDet: Scalable and efficient object detection[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2020: 10778-10787.
|
| [18] |
WU Y, CHEN Y P, YUAN L, et al. Rethinking classification and localization for object detection[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2020: 10183-10192.
|
| [19] |
DAI J F, QI H Z, XIONG Y W, et al. Deformable convolutional networks[C]//Proceedings of the IEEE International Conference on Computer Vision. Piscataway: IEEE Press, 2017: 764-773.
|
| [20] |
VASWANI A, SHAZEER N, PARMAR N, et al. Attention is all you need[C]//Proceedings of the International Conference on Neural Information Processing Systems. New York: ACM, 2017: 6000-6010.
|
| [21] |
WANG J Q, CHEN K, XU R, et al. CARAFE: Content-aware ReAssembly of FEatures[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Piscataway: IEEE Press, 2019: 3007-3016.
|
| [22] |
ZHU X Z, CHENG D Z, ZHANG Z, et al. An empirical study of spatial attention mechanisms in deep networks[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Piscataway: IEEE Press, 2019: 6687-6696.
|
| [23] |
CAO Y R, HE Z J, WANG L J, et al. VisDrone-DET2021: The vision meets drone object detection challenge results[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops. Piscataway: IEEE Press, 2021: 2847-2854.
|
| [24] |
ZHANG H Y, WANG Y, DAVOUB F, et al. VarifocalNet: An iou-aware dense object detector[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2021: 8510-8519.
|