Volume 50 Issue 7
Jul.  2024
Turn off MathJax
Article Contents
WANG D W,HU L C,FANG J,et al. Small target detection algorithm based on improved Double-Head RCNN for UAV aerial images[J]. Journal of Beijing University of Aeronautics and Astronautics,2024,50(7):2141-2149 (in Chinese)
Citation: WANG D W,HU L C,FANG J,et al. Small target detection algorithm based on improved Double-Head RCNN for UAV aerial images[J]. Journal of Beijing University of Aeronautics and Astronautics,2024,50(7):2141-2149 (in Chinese)

Small target detection algorithm based on improved Double-Head RCNN for UAV aerial images

doi: 10.13700/j.bh.1001-5965.2022.0591
Funds:

National Natural Science Foundation of China (62201454); Postgraduate Innovation Foundation of Xi’an University of Posts and Telecommunications (CXJJLY2021058) 

More Information
  • Corresponding author: E-mail:wangdianwei@xupt.edu.cn
  • Received Date: 05 Jul 2022
  • Accepted Date: 01 Nov 2022
  • Available Online: 13 Jan 2023
  • Publish Date: 10 Jan 2023
  • The feature information of small targets in unmanned aerial vehicle aerial images is small and easily interfered with by noise, which leads to the high missed detection and false detection rates of existing algorithms. To address these issues, a small target detection algorithm based on an improved Double-Head region-convolutional neural networks(RCNN)for unmanned aerial vehicle aerial images was proposed. Transformer and deformable convolution networks (DCN) modules were introduced on the backbone network ResNet-50 to extract small target feature information and semantic information more effectively. A feature pyramid network(FPN) structure based on content-aware reassembly of features (CARAFE) was proposed to solve the problem that the small target information is interfered with by the background noise, and the feature information is lost in the process of feature fusion. The generation scale of Anchor was reset according to the characteristics of small target scale distribution in the region proposal network to further improve the small target detection performance. The experimental results on the VisDrone-DET2021 dataset show that the proposed algorithm can extract feature and semantic information of small targets with representational capacity more effectively. Compared with the Double-Head RCNN algorithm, the parameter quantity of the proposed algorithm increases by 9.73×106, and the FPS loss is 0.6. However, AP, AP50, and AP75 increase by 2.6%, 6.2%, and 2.1% respectively, and APs increases by 3.1%.

     

  • loading
  • [1]
    LIN T Y, MAIRE M, BELONGIE S, et al. Microsoft COCO: Common objects in context[C]//Proceedings of the European Conference on Computer Vision. Berlin: Springer, 2014: 740-755.
    [2]
    WANG J W, YANG W, GUO H W, et al. Tiny object detection in aerial images[C]//Proceedings of the International Conference on Pattern Recognition. Piscataway: IEEE Press, 2021: 3791-3798.
    [3]
    陈映雪, 丁文锐, 李红光, 等. 基于视频帧间运动估计的无人机图像车辆检测[J]. 北京航空航天大学学报, 2020, 46(3): 634-642.

    CHEN Y X, DING W R, LI H G, et al. Vehicle detection in UAV image based on video interframe motion estimation[J]. Journal of Beijing University of Aeronautics and Astronautics, 2020, 46(3): 634-642(in Chinese).
    [4]
    WANG W H, XIE E Z, LI X, et al. PVT v2: Improved baselines with pyramid vision transformer[J]. Computational Visual Media, 2022, 8(3): 415-424.
    [5]
    CARION N, MASSA F, SYNNAEVE G, et al. End-to-end object detection with transformers[C]//Proceedings of the European Conference on Computer Vision. Berlin: Springer, 2020: 213-229.
    [6]
    ZHU X Z, SU W J, LU L W, et al. Deformable DETR: Deformable transformers for end-to-end object detection[EB/OL]. (2021-03-18)[2022-05-18].
    [7]
    LIN T Y, GOYAL P, GIRSHICK R, et al. Focal loss for dense object detection[C]//Proceedings of the IEEE International Conference on Computer Vision. Piscataway: IEEE Press, 2017: 2999-3007.
    [8]
    GE Z, LIU S T, WANG F, et al. YOLOX: Exceeding YOLO series in 2021[EB/OL]. (2021-08-06)[2022-05-20].
    [9]
    WANG C Y, BOCHKOVSKIY A, LIAO H Y M. YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors[EB/OL]. (2022-07-01)[2022-07-03].
    [10]
    REN S Q, HE K M, GIRSHICK R, et al. Faster R-CNN: Towards real-time object detection with region proposal networks[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 39(6): 1137-1149.
    [11]
    PANG J M, CHEN K, SHI J P, et al. Libra R-CNN: Towards balanced learning for object detection[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2019: 821-830.
    [12]
    CAI Z W, VASCONCELOS N. Cascade R-CNN: Delving into high quality object detection[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2018: 6154-6162.
    [13]
    LU X, LI B Y, YUE Y X, et al. Grid R-CNN[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2019: 7355-7364.
    [14]
    GIRSHICK R. Fast R-CNN[C]//Proceedings of the IEEE International Conference on Computer Vision. Piscataway: IEEE Press, 2015: 1440-1448.
    [15]
    LIN T Y, DOLLÁR P, GIRSHICK R, et al. Feature pyramid networks for object detection[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2017: 936-944.
    [16]
    LIU S, QI L, QIN H F, et al. Path aggregation network for instance segmentation[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2018: 8759-8768.
    [17]
    TAN M X, PANG R M, LE Q V. EfficientDet: Scalable and efficient object detection[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2020: 10778-10787.
    [18]
    WU Y, CHEN Y P, YUAN L, et al. Rethinking classification and localization for object detection[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2020: 10183-10192.
    [19]
    DAI J F, QI H Z, XIONG Y W, et al. Deformable convolutional networks[C]//Proceedings of the IEEE International Conference on Computer Vision. Piscataway: IEEE Press, 2017: 764-773.
    [20]
    VASWANI A, SHAZEER N, PARMAR N, et al. Attention is all you need[C]//Proceedings of the International Conference on Neural Information Processing Systems. New York: ACM, 2017: 6000-6010.
    [21]
    WANG J Q, CHEN K, XU R, et al. CARAFE: Content-aware ReAssembly of FEatures[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Piscataway: IEEE Press, 2019: 3007-3016.
    [22]
    ZHU X Z, CHENG D Z, ZHANG Z, et al. An empirical study of spatial attention mechanisms in deep networks[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Piscataway: IEEE Press, 2019: 6687-6696.
    [23]
    CAO Y R, HE Z J, WANG L J, et al. VisDrone-DET2021: The vision meets drone object detection challenge results[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops. Piscataway: IEEE Press, 2021: 2847-2854.
    [24]
    ZHANG H Y, WANG Y, DAVOUB F, et al. VarifocalNet: An iou-aware dense object detector[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2021: 8510-8519.
  • 加载中

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(7)  / Tables(2)

    Article Metrics

    Article views(1913) PDF downloads(176) Cited by()
    Proportional views
    Related

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return