亚信科技(中国)有限公司,北京 100193
常亮(1979− ),男,亚信科技(中国)有限公司专网产品产品经理,主要研究方向为5/6G网络管理和网络智能化技术。
王志刚(1979− ),男,亚信科技(中国)有限公司网络产品研发负责人,主要研究方向为5/6G/卫星网络和网络智能化技术。
王林颖(1983− ),男,亚信科技(中国)有限公司能源智能连接事业部总经理,主要研究方向为5/6G/卫星网络和网络智能化技术。
收稿:2026-04-11,
修回:2026-07-11,
录用:2026-07-13,
移动端阅览
常亮, 王志刚, 王林颖. 基于多智能体深度强化学习的空天地融合网络协同卸载与资源调度策略[J/OL]. 电信科学, 2026.
CHANG Liang, WANG Zhigang, WANG Linying. Multi-agent deep reinforcement learning-based collaborative offloading and resource scheduling strategy for space-air-ground integrated network[J/OL]. Telecommunications Science, 2026.
常亮, 王志刚, 王林颖. 基于多智能体深度强化学习的空天地融合网络协同卸载与资源调度策略[J/OL]. 电信科学, 2026. DOI: 10.11959/j.issn.1000-0801.DXKX260226.
CHANG Liang, WANG Zhigang, WANG Linying. Multi-agent deep reinforcement learning-based collaborative offloading and resource scheduling strategy for space-air-ground integrated network[J/OL]. Telecommunications Science, 2026. DOI: 10.11959/j.issn.1000-0801.DXKX260226.
针对空天地融合网络(SAGIN)中计算密集型任务对时延与能耗的严苛要求,本文深入探讨了如何将 SAGIN 复杂的异构物理特性(如高动态拓扑、卫星可视窗口及差异化链路增益)严谨地映射为多智能体马尔可夫决策过程(MDP)的状态空间、动作空间与归一化奖励函数。 在此基础上,设计了一种基于多头注意力机制的多智能体深度确定性策略梯度算法(AM-MADDPG),实现“端-边-云”三层协同计算卸载与多维资源调度。考虑到高动态环境易导致传统优化算法陷入维数灾难的问题,引入了“集中式训练、分布式执行”(CTDE)架构。各边缘智能体通过多头注意力机制在环境特征提取阶段精准识别近端碰撞风险,实现隐式协同,在线执行阶段仅凭局部观测即可输出决策。 仿真结果表明,该机制有效缓解了最大容忍时延约束下的系统压力;相较于主流基准(如 MAPPO)及基于局部邻域特征聚合的分布式图强化学习基准(DGN),本文算法在信令开销对等的前提下,展现出更优的能耗控制与系统开销优化能力,提升了极端高负载场景下的系统韧性。
Aiming at the stringent delay and energy consumption requirements of computation-intensive tasks in the Space-Air-Ground Integrated Network (SAGIN)
this paper explores how to rigorously map the complex heterogeneous physical characteristics of SAGIN into a Markov Decision Process (MDP) and introduces a multi-head attention mechanism to optimize the Multi-Agent Deep Deterministic Policy Gradient (AM-MADDPG) algorithm. Considering the curse of dimensionality in highly dynamic environments
a Centralized Training with Decentralized Execution (CTDE) architecture is introduced. Each edge agent utilizes multi-head attention during the feature extraction phase to accurately assess collision risks and achieves coordination with ultra-low delay. Simulation results show that the proposed mechanism effectively alleviates task delay constraints. Compared with mainstream benchmarks like MAPPO and distributed Graph Reinforcement Learning based on local neighborhood aggregation (DGN)
the proposed algorithm demonstrates superior cost reduction under the premise of equivalent signaling overhead
enhancing system resilience in extreme heavy-load scenarios.
张平 , 牛凯 , 贺志强 , 等 . 6G移动通信技术展望 [J ] . 通信学报 , 2019 , 40 ( 1 ): 141 - 148 .
ZHANG P , NIU K , HE Z Q , et al . 6G mobile communications technologies: A prospect [J ] . Journal on Communications , 2019 , 40 ( 1 ): 141 - 148 .
施巍松 , 孙辉 , 曹杰 , 等 . 边缘计算: 万物互联时代新型计算模型 [J ] . 计算机研究与发展 , 2017 , 54 ( 5 ): 907 - 924 .
SHI W S , SUN H , CAO J , et al . Edge computing: Vision and challenges [J ] . Journal of Computer Research and Development , 2017 , 54 ( 5 ): 907 - 924 .
WANG X , LI Y , ZHANG Z , et al . Advanced deep learning models for 6G: Overview, opportunities and challenges [J ] . IEEE Communications Surveys & Tutorials , 2024 , 26 ( 1 ): 112 - 145 .
曹阔 , 丁国如 , 郑宝玉 , 等 . 面向6G的空天地一体化网络架构与关键技术 [J ] . 通信学报 , 2020 , 41 ( 2 ): 134 - 145 .
CAO K , DING G R , ZHENG B Y , et al . Space-air-ground integrated network architecture and key technologies toward 6G [J ] . Journal on Communications , 2020 , 41 ( 2 ): 134 - 145 .
XIE W , CHEN C , JU Y , et al . Deep reinforcement learning-based computation offloading for space-air-ground integrated vehicle networks [J ] . IEEE Transactions on Intelligent Transportation Systems , 2025 , 26 ( 5 ): 5804 - 5815 .
TANG Q , FEI Z , LI B , et al . Computation offloading in LEO satellite networks with multicore setup [J ] . IEEE Communications Letters , 2021 , 25 ( 3 ): 906 - 910 .
DING C , WANG J B , ZHANG H , et al . Joint optimization of transmission and computation resources for satellite and UAV-assisted internet of things [J ] . IEEE Transactions on Wireless Communications , 2020 , 19 ( 8 ): 5512 - 5525 .
XIONG J , GUO H , LIU J . Task offloading in UAV-aided edge computing: Bit allocation and trajectory optimization [J ] . IEEE Communications Letters , 2019 , 23 ( 3 ): 538 - 541 .
ZHANG T , LEI J , LIU Y , et al . Trajectory optimization and resource allocation for UAV-assisted mobile edge computing [J ] . IEEE Wireless Communications Letters , 2020 , 9 ( 10 ): 1600 - 1604 .
WANG Y , FANG W , DING Y , et al . Computation offloading optimization for UAV-assisted mobile edge computing: A deep deterministic policy gradient approach [J ] . Wireless Networks , 2021 , 27 ( 4 ): 2991 - 3006 .
CUI J , LIU Y , NIYATO D . Multi-agent reinforcement learning-based resource allocation for UAV networks [J ] . IEEE Transactions on Wireless Communications , 2020 , 19 ( 1 ): 729 - 743 .
LI Z , ZHANG Y , WANG X , et al . Blockchain-aided digital twin offloading mechanism in space-air-ground networks [J ] . IEEE Transactions on Mobile Computing , 2025 , 24 ( 2 ): 880 - 895 .
刘明 , 邓晓燕 , 王卫东 . 面向空天地融合网络的协同计算卸载与资源分配算法 [J ] . 电子与信息学报 , 2022 , 44 ( 8 ): 2612 - 2620 .
LIU M , DENG X Y , WANG W D . Collaborative computation offloading and resource allocation algorithm for space-air-ground integrated networks [J ] . Journal of Electronics & Information Technology , 2022 , 44 ( 8 ): 2612 - 2620 .
LOWE R , WU Y , TAMAR A , et al . Multi-agent actor-critic for mixed cooperative-competitive environments [C ] // Advances in Neural Information Processing Systems (NeurIPS) . Long Beach : Curran Associates , 2017 : 6379 - 6390 .
赵成林 , 王文 , 李子涵 . 基于MADDPG的多无人机协同轨迹规划与资源分配 [J ] . 北京邮电大学学报 , 2022 , 45 ( 1 ): 1 - 7 .
ZHAO C L , WANG W , LI Z H . Multi-UAV collaborative trajectory planning and resource allocation based on MADDPG [J ] . Journal of Beijing University of Posts and Telecommunications , 2022 , 45 ( 1 ): 1 - 7 .
ZHANG Y , LIU J , WANG X , et al . Graphic deep reinforcement learning for dynamic resource allocation in space-air-ground integrated networks [J ] . IEEE Journal on Selected Areas in Communications , 2024 , 42 ( 4 ): 1021 - 1035 .
LI Y , LI J , YU C , et al . A hierarchical conflict resolution framework with graph transformer-based reinforcement learning for heterogeneous UAV networks [J ] . IEEE Internet of Things Journal , 2025 , 12 ( 3 ): 2150 - 2165 .
SINGH P , HAZARIKA B , HUANG W J . Graph-based federated multi-agent DRL for semantic and intent-aware V2X communication [J ] . IEEE Internet of Things Journal , 2025 , 12 ( 1 ): 450 - 462 .
MA X , HU J , LIANG S , et al . Federated learning and resource-aware graph neural network for intrusion detection in 6G-IoT driven healthcare system [J ] . IEEE Internet of Things Journal , 2025 , 12 ( 2 ): 1120 - 1135 .
LIU X , FISCHIONE C . Coordinated beamforming for multi-cell ISAC using graph neural networks [J ] . IEEE Transactions on Wireless Communications , 2025 , 24 ( 1 ): 310 - 324 .
JIANG J , LU Z . Graph convolutional reinforcement learning [C ] // Proceedings of the 7th International Conference on Learning Representations (ICLR) . New Orleans : OpenReview.net , 2019 : 1 - 15 .
0
浏览量
0
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621