留言板

尊敬的读者、作者、审稿人, 关于本刊的投稿、审稿、编辑和出版的任何问题, 您可以本页添加留言。我们将尽快给您答复。谢谢您的支持!

姓名
邮箱
手机号码
标题
留言内容
验证码

基于深度强化学习的高速无人机智能路由策略

寿一凡,  刘东升,  陈亚辉

寿一凡,刘东升,陈亚辉. 基于深度强化学习的高速无人机智能路由策略[J]. 北京航空航天大学学报,2026,52(9):3211-3222
引用本文: 寿一凡,刘东升,陈亚辉. 基于深度强化学习的高速无人机智能路由策略[J]. 北京航空航天大学学报,2026,52(9):3211-3222
Shou Y F,Liu D S,Chen Y H. Intelligent routing strategy for high-speed UAVs via deep reinforcement learning[J]. Journal of Beijing University of Aeronautics and Astronautics,2026,52(9):3211-3222 (in Chinese)
Citation: Shou Y F,Liu D S,Chen Y H. Intelligent routing strategy for high-speed UAVs via deep reinforcement learning[J]. Journal of Beijing University of Aeronautics and Astronautics,2026,52(9):3211-3222 (in Chinese)

基于深度强化学习的高速无人机智能路由策略

doi: 10.13700/j.bh.1001-5965.2025.0282
基金项目: 

浙江省特支持计划项目(2022R52048);浙江省科技计划项目(2024C01028,2024C01211)

详细信息
    通讯作者:

    E-mail:lds1118@zjgsu.edu.cn

  • 中图分类号: V279+.3

Intelligent routing strategy for high-speed UAVs via deep reinforcement learning

Funds: 

Zhejiang Province Special Support Program Project (2022R52048);Zhejiang Provincial Science and Technology Program (2024C01028,2024C01211)

More Information
  • 摘要:

    随着无人机技术的迅速发展,高速无人机集群在低空复杂场景中得到越来越广泛的应用,但同时也带来了传统路由协议难以满足高动态网络环境的问题。基于此,提出一种基于深度Q网络(DQN)深度强化学习模型的高速无人机智能路由策略,聚焦于高速、高动态场景下的通信保障,并进一步加强无人机节点的本地化独立决策能力,同时,构建用于衡量链路未来稳定性的Stable稳定度因子,使无人机能根据最新状态自适应地进行独立决策。实验结果表明:所提高速无人机智能路由策略在不同速度对比场景中,数据包抵达率均保持在85%以上,平均端到端时延、平均每跳时延相较传统路由协议平均降低23%以上。在高频率通信场景下,数据包抵达率较传统路由协议平均提升15%以上,满足高速无人机高动态通信组网对稳定性与通信效率的需求。

     

  • 图 1  无人机集群通信自组网建模

    Figure 1.  Modeling of self-organizing network communication for UAV swarms

    图 2  高速无人机集群智能路由策略框架

    Figure 2.  Intelligent routing policy framework for high-speed UAV swarms

    图 3  训练过程平均每步奖励对比

    Figure 3.  Comparison of the average rewards per step during the training process

    图 4  训练过程平均每跳时延对比

    Figure 4.  Comparison of average latency per jump during the training process

    图 5  不同数据包发送时间间隔下网络性能对比

    Figure 5.  Comparison of network performance at different packet sending intervals

    图 6  不同速度场景下网络性能对比

    Figure 6.  Comparison of network performance in different speed scenarios

    图 7  不同网络规模下网络性能对比

    Figure 7.  Comparison of network performance at different network scales

    图 8  不同节点密度下网络性能对比

    Figure 8.  Comparison of network performance under different node densities

    表  1  无人机集群通信自组网络仿真参数设定

    Table  1.   Simulation parameters setting for UAV swarm communication self-organizing networks

    集群场景
    大小/m
    无人机数量 最大通信
    距离/m
    无人机最大
    速度/(m·s−1)
    无人机最小
    速度/(m·s−1)
    数据包
    大小/bytes
    数据包发送
    间隔/ms
    Hello Package
    发送间隔/ms
    通信频率/GHz 最小可用
    带宽/(Mbit·s−1)
    1500×1500×500 30~70 150 10~160 5~80 1 024 20~80 150 2.4 1
    下载: 导出CSV

    表  2  深度强化学习模型参数设定

    Table  2.   Setting parameters for deep reinforcement learning models

    学习率训练轮数训练批量隐藏层维度$\gamma $$\tau $经验回放池容量$\varepsilon $ (DDQN)ε的最小值 (DDQN)N Steps网络评估频率网络评估次数
    0.000 01200 0002562560.990.005100 0000.90.0531 0003
    下载: 导出CSV
  • [1] Maza I, Caballero F, Capitán J, et al. Experimental results in multi-UAV coordination for disaster management and civil security applications[J]. Journal of Intelligent & Robotic Systems, 2011, 61(1): 563-585.
    [2] Keller J, Thakur D, Likhachev M, et al. Coordinated path planning for fixed-wing UAS conducting persistent surveillance missions[C]// Proceedings of the IEEE International Symposium on Safety, Security, and Rescue Robotics. Piscataway: IEEE Press, 2016: 1-6.
    [3] Khan A, Zhang J, Ahmad S, et al. Dynamic positioning and energy-efficient path planning for disaster scenarios in 5G-assisted multi-UAV environments[J]. Electronics, 2022, 11(14): 2197.
    [4] Dahmane S, Yagoubi M B, Brik B, et al. Multi-constrained and edge-enabled selection of UAV participants in federated learning process[J]. Electronics, 2022, 11(14): 2119.
    [5] Sharma R, Patel K, Shah S, et al. Aerial footage analysis using computer vision for efficient detection of points of interest near railway tracks[J]. Aerospace, 2022, 9(7): 370.
    [6] Zhang R, Li S, Ding Y M, et al. UAV path planning algorithm based on improved Harris Hawks optimization[J]. Sensors, 2022, 22(14): 5232.
    [7] Pang X, Liu M, Li Z C, et al. Geographic position based hopless opportunistic routing for UAV networks[J]. Ad Hoc Networks, 2021, 120: 102560.
    [8] Wheeb A H, Nordin R, Samah A A, et al. Topology-based routing protocols and mobility models for flying ad hoc networks: a contemporary review and future research directions[J]. Drones, 2022, 6(1): 9.
    [9] Hong L, Guo H Z, Liu J J, et al. Toward swarm coordination: topology-aware inter-UAV routing optimization[J]. IEEE Transactions on Vehicular Technology, 2020, 69(9): 10177-10187.
    [10] Shen H, Jiang Y J, Deng F M, et al. Task unloading strategy of multi UAV for transmission line inspection based on deep reinforcement learning[J]. Electronics, 2022, 11(14): 2188.
    [11] Perkins C E, Bhagwat P. Highly dynamic destination-sequenced distance-vector routing (DSDV) for mobile computers[C]//Proceedings of the Conference on Communications Architectures, Protocols and Applications. New York: ACM, 1994: 234-244.
    [12] Jacquet P, Muhlethaler P, Clausen T, et al. Optimized link state routing protocol for ad hoc networks[C]//Proceedings of IEEE International Multi Topic Conference, 2001. Technology for the 21st Century. Piscataway: IEEE Press, 2002: 62-68.
    [13] Johnson D B, Hu Y C, Maltz D A. The dynamic source routing protocol (DSR) for mobile Ad hoc networks for IPv4, RFC 4728[R]. Fremont: Internet Engineering Task Force, 2007.
    [14] Perkins C E, Belding-royer E M, Das S R. Ad hoc on-demand distance vector (AODV) routing[J]. RFC, 2003, 3561: 1-37.
    [15] Park V D, Corson M S. A highly adaptive distributed routing algorithm for mobile wireless networks[C]//Proceedings of INFOCOM '97. Piscataway: IEEE Press, 2002: 1405-1413.
    [16] 孟泠宇, 郭秉礼, 杨雯, 等. 基于深度强化学习的网络路由优化方法[J]. 系统工程与电子技术, 2022, 44(7): 2311-2318.

    Meng L Y, Guo B L, Yang W, et al. Network routing optimization approach based on deep reinforcement learning[J]. Systems Engineering and Electronics, 2022, 44(7): 2311-2318(in Chinese).
    [17] Karp B, Kung H T. GPSR: greedy perimeter stateless routing for wireless networks[C]//Proceedings of the 6th Annual International Conference on Mobile Computing and Networking. New York: ACM, 2000: 243-254.
    [18] Garg S, Ihler A, Bentley E S, et al. A hybrid reactive routing protocol for decentralized UAV networks[EB/OL]. (2025-01-23)[2025-05-08]. https://arxiv.org/abs/2407.02929.
    [19] Cao Y, Lien S Y, Liang Y C. Deep reinforcement learning for multi-user access control in non-terrestrial networks[J]. IEEE Transactions on Communications, 2021, 69(3): 1605-1619.
    [20] Lin D P, Peng T, Zuo P L, et al. Deep-reinforcement-learning-based intelligent routing strategy for FANETs[J]. Symmetry, 2022, 14(9): 1787.
    [21] Qiu X L, Yang Y W, Xu L, et al. Maintaining links in the highly dynamic FANET using deep reinforcement learning[J]. IEEE Transactions on Vehicular Technology, 2023, 72(3): 2804-2818.
    [22] Liu C Z, Wang Y X, Wang Q. PARouting: prediction-supported adaptive routing protocol for FANETs with deep reinforcement learning[J]. International Journal of Intelligent Networks, 2023, 4: 113-121.
    [23] Jung W S, Yim J, Ko Y B. QGeo: Q-learning-based geographic ad hoc routing protocol for unmanned robotic networks[J]. IEEE Communications Letters, 2017, 21(10): 2258-2261.
    [24] Qiu X, Xu L, Wang P, et al. A data-driven packet routing algorithm for an unmanned aerial vehicle swarm: a multi-agent reinforcement learning approach[J]. IEEE Wireless Communications Letters, 2022, 11(10): 2160-2164.
    [25] Bai Y J, Zhang X, Yu D J, et al. A deep reinforcement learning-based geographic packet routing optimization[J]. IEEE Access, 2022, 10: 108785-108796.
    [26] Wang Z, Li H X, Knoblock E J, et al. Deep reinforcement learning-based joint routing and capacity optimization in an aerial and terrestrial hybrid wireless network[J]. IEEE Access, 2024, 12: 132056-132069.
    [27] Ke Y Q, Huang K, Qiu X L, et al. Distributed routing optimization algorithm for FANET based on multiagent reinforcement learning[J]. IEEE Sensors Journal, 2024, 24(15): 24851-24864.
    [28] Mnih V, Kavukcuoglu K, Silver D, et al. Human-level control through deep reinforcement learning[J]. Nature, 2015, 518(7540): 529-533.
    [29] Van Hasselt H, Guez A, Silver D. Deep reinforcement learning with double Q-Learning[C]//Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence. New York: ACM, 2016: 2094-2100.
    [30] Hessel M, Modayil J, Van Hasselt H, et al. Rainbow: combining improvements in deep reinforcement learning[C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2018, 32(1): 5737-5744.
    [31] Hu G N, Zhang W, Zhu W H. Prioritized experience replay for continual learning[C]//Proceedings of the 2021 6th International Conference on Computational Intelligence and Applications. Piscataway: IEEE Press, 2021: 16-20.
    [32] Wang Z Y, Schaul T, Hessel M, et al. Dueling network architectures for deep reinforcement learning[C]//Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48. New York: ACM, 2016: 1995-2003.
    [33] Bellemare M G, Dabney W, Munos R. A distributional perspective on reinforcement learning[EB/OL]. (2017-01-27)[2025-05-08]. https://arxiv.org/abs/1707.06887.
    [34] Fortunato M, Azar M G, Piot B, et al. Noisy networks for exploration[EB/OL]. (2019-01-09)[2025-05-08]. https://arxiv.org/abs/1706.10295.
  • 加载中
图(8) / 表(2)
计量
  • 文章访问数:  400
  • HTML全文浏览量:  243
  • PDF下载量:  12
  • 被引次数: 0
出版历程
  • 收稿日期:  2025-05-09
  • 录用日期:  2025-07-25
  • 网络出版日期:  2025-07-31
  • 整期出版日期:  2026-09-01

目录

    /

    返回文章
    返回
    常见问答