中信数字科技集团有限公司,北京 100004
赵明(1980- ),男,中信数字科技集团公司高级工程师,主要研究方向为云计算、数据通信、网络信息与安全、人工智能、数据中心等。
收稿:2026-01-23,
修回:2026-02-27,
录用:2026-02-28,
网络首发:2026-07-22,
移动端阅览
赵明.智算网络中面向QoS感知的资源调度技术研究[J].电信科学,
Zhao Ming.Research on QoS-aware resource scheduling for AI computing networks[J].Telecommunications Science,
赵明.智算网络中面向QoS感知的资源调度技术研究[J].电信科学, DOI:10.11959/j.issn.1000−0801.DXKX260062.
Zhao Ming.Research on QoS-aware resource scheduling for AI computing networks[J].Telecommunications Science, DOI:10.11959/j.issn.1000−0801.DXKX260062.
为保障大语言模型推理与多智能体协同等新兴智能业务对服务质量(quality of service,QoS)的需求,构建高效的资源调度机制已成为智算网络中的核心问题。面向QoS多目标优化下的资源调度场景,设计了一种融合差异化智能业务建模、异构算网资源感知与高通量远程直接内存访问(remote direct memory access,RDMA)传输优化的调度框架;在该框架下提出基于强化学习(reinforcement learning,RL)的智能调度算法,在考虑业务优先级、算网资源负载与数据传输特性的基础上,自适应生成调度策略以提升QoS;引入网络实测数据与智算中心真实组网参数构建仿真平台,对所提算法进行性能验证。实验结果表明,该方法在任务完成率、平均时延和资源利用率等关键QoS指标上较现有方案具有明显优势。
To meet the quality of service (QoS) requirements of emerging intelligent services such as large language model inference and multi-agent collaboration
constructing an efficient resource scheduling mechanism has become a core issue in AI computing networks. For the resource scheduling scenario under QoS multi-objective optimization
a scheduling framework that integrates differentiated intelligent service modeling
heterogeneous computing-network resource awareness
and high-throughput remote direct memory access (RDMA) transmission optimization was proposed. Within this framework
a reinforcement learning (RL)-based intelligent scheduling algorithm was developed to adaptively generate scheduling policies by considering service priorities
computing-network resource loads
and data transmission characteristics
thereby improving QoS. Finally
a simulation platform was constructed using real-world network measurement data and actual topology parameters from an AI computing center to verify the performance of the proposed algorithm. Experimental results demonstrate that the proposed method achieves significant improvements over existing schemes in key QoS metrics such as task completion rate
average latency
and resource utilization.
Weisz J D , He J , Muller M , et al . Design principles for generative AI applications [C ] // Proceedings of the CHI Conference on Human Factors in Computing Systems . New York : ACM Press , 2024 : 1 - 22 .
Agrawal A , Kedia N , Panwar A , et al . Taming throughput-latency tradeoff in LLM inference with Sarathi-Serve [C ] // Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . Santa Clara : USENIX Association , 2024 : 117 - 134 .
Hu J B , Shen H Q , Liu X C , et al . RDMA transports in datacenter networks: survey [J ] . IEEE Network , 2024 , 38 ( 6 ): 380 - 387 .
Sun B , Huang Z , Zhao H , et al . . Llumnix: Dynamic scheduling for large language model serving [C ] // Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . Santa Clara : USENIX Association , 2024 : 173 - 191 .
Sun B Q , Theile M , Qin Z Y , et al . Edge generation scheduling for DAG tasks using deep reinforcement learning [J ] . IEEE Transactions on Computers , 2024 , 73 ( 4 ): 1034 - 1047 .
Huang J T , Golubchik L , Huang L B . When Lyapunov drift based queue scheduling meets adversarial bandit learning [J ] . IEEE/ACM Transactions on Networking , 2024 , 32 ( 4 ): 3034 - 3044 .
Ahmed G , Sheltami T , Ghaleb M , et al . Energy-efficient Internet of drones path-planning study using meta-heuristic algorithms [J ] . Applied Sciences , 2024 , 14 ( 6 ): 2418 .
Zhang Y N , Tzeng E , Du Y L , et al . Large-scale reinforcement learning for diffusion models [C ] // Computer Vision-ECCV 2024 . Cham : Springer Nature Switzerland , 2024 : 1 - 17 .
Paul A , Singh K , Kaushik A , et al . Quantum-enhanced DRL optimization for DoA estimation and task offloading in ISAC systems [J ] . IEEE Journal on Selected Areas in Communications , 2025 , 43 ( 1 ): 364 - 381 .
Gangidi A , Miao R , Zheng S B , et al . RDMA over Ethernet for distributed training at meta scale [C ] // Proceedings of the ACM SIGCOMM 2024 Conference . New York : ACM Press , 2024 : 57 - 70 .
Cai K , Wu Q W , Zhou M C , et al . Dynamically scheduling deadline-constrained interleaved workflows on heterogeneous computing systems [J ] . IEEE Transactions on Services Computing , 2025 , 18 ( 2 ): 758 - 769 .
Ma X , Zhou A , Zhang S , et al . Dynamic task scheduling in cloud-assisted mobile edge computing [J ] . IEEE Transactions on Mobile Computing , 2023 , 22 ( 4 ): 2116 - 2130 .
Lu Y Z , Yang L , Yang S X , et al . An intelligent deterministic scheduling method for ultralow latency communication in edge enabled industrial Internet of Things [J ] . IEEE Transactions on Industrial Informatics , 2023 , 19 ( 2 ): 1756 - 1767 .
Cheng Z R , Zhang W T , Yang D , et al . Intelligent end-to-end deterministic scheduling across converged networks [J ] . IEEE Transactions on Mobile Computing , 2025 , 24 ( 4 ): 2504 - 2518 .
Zhao G M , Wang J Z , Xu H L , et al . Joint request updating and elastic resource provisioning with QoS guarantee in clouds [J ] . IEEE/ACM Transactions on Networking , 2024 , 32 ( 1 ): 110 - 126 .
Ali J , Song H H , Roh B H . An SDN-based framework for E2E QoS guarantee in Internet of Things devices [J ] . IEEE Internet of Things Journal , 2025 , 12 ( 1 ): 605 - 622 .
Zhao L L , Chi X F , Ma S D . Delay-QoS-aware local-information-driven multiple access for MTC networks [J ] . IEEE Transactions on Mobile Computing , 2024 , 23 ( 5 ): 5848 - 5862 .
Zhu J , Wang S W . QoS-guaranteed resource allocation in mobile communications: a stochastic network calculus approach [J ] . IEEE/ACM Transactions on Networking , 2024 , 32 ( 6 ): 5159 - 5171 .
Peng Y , Tang X G , Zhou Y Q , et al . Computing and communication cost-aware service migration enabled by transfer reinforcement learning for dynamic vehicular edge computing networks [J ] . IEEE Transactions on Mobile Computing , 2024 , 23 ( 1 ): 257 - 269 .
Gao X Y , Sun Y P , Chen H , et al . Joint computing, pushing, and caching optimization for mobile-edge computing networks via soft actor-critic learning [J ] . IEEE Internet of Things Journal , 2024 , 11 ( 6 ): 9269 - 9281 .
Zhang X Y , Liu J , Zhang R , et al . Energy-efficient computation peer offloading in satellite edge computing networks [J ] . IEEE Transactions on Mobile Computing , 2024 , 23 ( 4 ): 3077 - 3091 .
Xu C , Zhang P F , Xia X F , et al . Digital-twin-assisted intelligent secure task offloading and caching in blockchain-based vehicular edge computing networks [J ] . IEEE Internet of Things Journal , 2025 , 12 ( 4 ): 4128 - 4143 .
Dai Y Y , Rao X Y , Gu B , et al . Graph learning-based multiuser multitask offloading in wireless computing power networks [J ] . IEEE Internet of Things Journal , 2025 , 12 ( 15 ): 29230 - 29239 .
Kuznetsov A , Shvechikov P , Grishin A , et al . Controlling overestimation bias with truncated mixture of continuous distributional quantile critics [C ] // Proceedings of the 37th International Conference on Machine Learning . New York : ACM Press , 2020 : 5556 - 5566 .
Haarnoja T , Zhou A , Abbeel P , et al . Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor [C ] // Proceedings of the 35th International Conference on Machine Learning (ICML 18) . Stockholm : PMLR , 2018 : 1861 - 1870 .
Fujimoto S , Van Hoof H , Meger D . Addressing function approximation error in actor-critic methods [C ] // Proceedings of the 35th International Conference on Machine Learning (ICML 18) . Stockholm : PMLR , 2018 : 1587 - 1596 .
Schulman J , Wolski F , Dhariwal P , et al . Proximal policy optimization algorithms [PP ] . V2. arXiv ( 2017-08-28 )[ 2026-01-0 ] arXiv: 1707.06347 , 2017 .
0
浏览量
0
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621