1.华北电力大学控制与计算机工程学院,北京 102206
2.交通运输部天津水运工程科学研究院,天津 300072
赵袁赫(2002- ),男,华北电力大学控制与计算机工程学院博士生,主要研究方向为电力信息安全、人工智能安全、多智能体协同。
王昱(1998- ),男,华北电力大学控制与计算机工程学院博士生,主要研究方向为电力信息安全、人工智能安全、多智能体协同。
李昊忱(1996- ),男,交通运输部天津水运工程科学研究院工程师,主要研究方向为计量测试。
杨士铎(1998- ),男,华北电力大学控制与计算机工程学院博士生,主要研究方向为电力信息安全、人工智能安全、多智能体协同。
郭柯楠(2000- ),男,华北电力大学控制与计算机工程学院硕士生,主要研究方向为电力信息安全、人工智能安全、多智能体协同。
高晨(2003- ),女,华北电力大学控制与计算机工程学院硕士生,主要研究方向为电力信息安全、人工智能安全、多智能体协同。
白小高(2003- ),男,华北电力大学控制与计算机工程学院硕士生,主要研究方向为电力信息安全、人工智能安全、多智能体协同。
管孜然(2002- ),女,华北电力大学控制与计算机工程学院硕士生,主要研究方向为电力信息安全、人工智能安全、多智能体协同。
王继业(1965- ),男,博士,华北电力大学控制与计算机工程学院教授、博士生导师、正高级工程师,主要研究方向为电力信息安全、人工智能安全、多智能体协同。
收稿:2026-02-25,
修回:2026-05-12,
录用:2026-06-08,
网络首发:2026-07-21,
纸质出版:2026-06-20
移动端阅览
赵袁赫,王昱,李昊忱等.基于冻结主干双回路进化的多智能体潜空间策略研究[J].电信科学,2026,42(06):78-91.
Zhao Yuanhe,Wang Yu,Li Haochen,et al.Research on multi-agent latent space strategies based on frozen main branch dual-loop evolution[J].Telecommunications Science,2026,42(06):78-91.
赵袁赫,王昱,李昊忱等.基于冻结主干双回路进化的多智能体潜空间策略研究[J].电信科学,2026,42(06):78-91. DOI: 10.11959/j.issn.1000-0801.DXKX260121.
Zhao Yuanhe,Wang Yu,Li Haochen,et al.Research on multi-agent latent space strategies based on frozen main branch dual-loop evolution[J].Telecommunications Science,2026,42(06):78-91. DOI: 10.11959/j.issn.1000-0801.DXKX260121.
针对大语言模型多智能体面临的“稳定性—可塑性困境”及微调引发的灾难性遗忘风险,提出一种基于冻结主干的双回路进化框架。该框架将通用推理与策略偏好解耦:行为回路利用离散价值迭代优化即时动作;语义回路引入预测误差调制的语义反思机制,将自然语言反馈转化为高维语义梯度,驱动外部潜向量在连续语义空间中演化。在GridWorld博弈环境实验中,该方法在完全冻结参数前提下实现了85.6%的任务成功率,路径加权成功率达到最高,且具备极低的推理延迟(2.1 s/step)和Token消耗(3.2 k/session)。实验还观察到基于隐式因果推理的社会协作行为,验证了该范式在低资源消耗与高可解释性方面的显著潜力。
To address the “stability-plasticity dilemma” faced by multi-agent systems based on large language models and the risk of catastrophic forgetting caused by fine-tuning
a frozen main branch dual-loop evolution framework was proposed. General reasoning was decoupled from policy preferences by this framework. Discrete value iteration was utilized to optimize immediate actions by the behavioral loop. A prediction error-modulated semantic reflection mechanism was introduced by the semantic loop that transformed natural language feedback into high-dimensional semantic gradients
driving the evolution of external latent vectors in a continuous semantic space. In experiments conducted in the GridWorld game environment
this method achieves a 85.6% task success rate under fully frozen parameters
with the highest path-weighted success rate
while exhibiting extremely low inference latency (2.1 s/step) and token consumption (3.2 k/session). The experiments also observe socially cooperative behavior based on implicit causal reasoning
validating the significant potential of this paradigm in terms of low resource consumption and high interpretability.
Xi Z H , Chen W X , Guo X , et al . The rise and potential of large language model based agents: a survey [J ] . Science China Information Sciences , 2025 , 68 ( 2 ): 121101 .
Wang L , Ma C , Feng X Y , et al . A survey on large language model based autonomous agents [J ] . Frontiers of Computer Science , 2024 , 18 ( 6 ): 186345 .
Chen X . A survey on large language model based autonomous agents [J ] . Computational Linguistics , 2024 ( 2 ): 141 - 150 .
Guo T , Chen X , Wang Y , et al . Large language model based multi-agents: a survey of progress and challenges [C ] // Proceedings of the 33rd International Joint Conference on Artificial Intelligence (IJCAI-24) . San Francisco : Morgan Kaufmann , 2024 : 8048 - 8057 .
李文斌 , 熊亚锟 , 范祉辰 , 等 . 持续学习的研究进展与趋势 [J ] . 计算机研究与发展 , 2024 , 61 ( 6 ): 1476 - 1496 .
Li W B , Xiong Y K , Fan Z C , et al . Advances and trends of continual learning [J ] . Journal of Computer Research and Development , 2024 , 61 ( 6 ): 1476 - 1496 .
Bosma M , Chi E , Ichter B , et al . Chain-of-thought prompting elicits reasoning in large language models [C ] // Proceedings of the Advances in Neural Information Processing Systems 35 . Red Hook : Curran Associates , 2022 : 24824 - 24837 .
Yao S , Zhao J , Yu D , et al . ReAct: synergizing reasoning and acting in language models [C ] // Proceedings of the Eleventh International Conference on Learning Representations . Amherst, Massachusetts : OpenReview , 2023 .
Wang G , Xie Y , Jiang Y , et al . Voyager: an open-ended embodied agent with large language models [J ] . Transactions on Machine Learning Research , 2024 .
Park J S , O' Brien J , Cai C J , et al . Generative agents: interactive simulacra of human behavior [C ] // Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology . New York : ACM Press , 2023 : 1 - 22 .
Cassano F , Gopinath A , Narasimhan K , et al . Reflexion: language agents with verbal reinforcement learning [C ] // Proceedings of the Advances in Neural Information Processing Systems 36 . Red Hook : Curran Associates , 2023 : 8634 - 8652 .
Sumers T , Yao S , Narasimhan K R , et al . Cognitive architectures for language agents [J ] . Transactions on Machine Learning Research , 2023 .
Liu N F , Lin K , Hewitt J , et al . Lost in the middle: how language models use long contexts [J ] . Transactions of the Association for Computational Linguistics , 2024 , 12 : 157 - 173 .
Chen B D , Chen R J , Liu S W , et al . Found in the middle: how language models use long contexts better via plug-and-play positional encoding [C ] // Proceedings of the Advances in Neural Information Processing Systems 37 . Red Hook : Curran Associates , 2024 : 60755 - 60775 .
Lester B , Al-Rfou R , Constant N . The power of scale for parameter-efficient prompt tuning [C ] // Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . Stroudsburg, PA : ACL Press , 2021 : 3045 - 3059 .
Liu X , Zheng Y N , Du Z X , et al . GPT understands, too [J ] . AI Open , 2024 , 5 : 208 - 215 .
Luo Y , Yang Z , Meng F D , et al . An empirical study of catastrophic forgetting in large language models during continual fine-tuning [J ] . IEEE Transactions on Audio, Speech and Language Processing , 2025 , 33 : 3776 - 3786 .
Huang J H , Cui L Y , Wang A T , et al . Mitigating catastrophic forgetting in large language models with self-synthesized rehearsal [C ] // Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Stroudsburg, PA : ACL Press , 2024 : 1416 - 1428 .
Li H Y , Ding L , Fang M , et al . Revisiting catastrophic forgetting in large language model tuning [C ] // Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2024 . Stroudsburg, PA : ACL Press , 2024 : 4297 - 4308 .
Sun T , Shao Y , Qian H , et al . Black-box tuning for language-model-as-a-service [C ] // Proceedings of the International Conference on Machine Learning . New York : PMLR Press , 2022 : 20841 - 20855 .
Hayes C F , Rădulescu R , Bargiacchi E , et al . A practical guide to multi-objective reinforcement learning and planning [J ] . Autonomous Agents and Multi-Agent Systems , 2022 , 36 : 26 .
Song Y F , Yin D , Yue X , et al . Trial and error: exploration-based trajectory optimization of LLM agents [C ] // Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Stroudsburg, PA : ACL Press , 2024 : 7584 - 7600 .
Xu Z , Gu W , Yu C , et al . Learning strategic language agents in the werewolf game with iterative latent space policy optimization [C ] // Proceedings of the 42nd International Conference on Machine Learning . New York : PMLR Press , 2025 .
王若男 , 董琦 . 基于学习机制的多智能体强化学习综述 [J ] . 工程科学学报 , 2024 , 46 ( 7 ): 1251 - 1268 .
Wang R N , Dong Q . Multiagent game decision-making method based on the learning mechanism [J ] . Chinese Journal of Engineering , 2024 , 46 ( 7 ): 1251 - 1268 .
Lei M , Cai H , Cui Z , et al . RoboMemory: a brain-inspired multi-memory agentic framework for lifelong learning in physical embodied systems [C ] // Proceedings of the NeurIPS 2025 Workshop on Space in Vision, Language, and Embodied AI . Red Hook : Curran Associates , 2025 .
Park C , Han S , Guo X Z , et al . MAPoRL: multi-agent post-co-training for collaborative large language models with reinforcement learning [C ] // Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Stroudsburg, PA : ACL Press , 2025 : 30215 - 30248 .
Zhou A , Yan K , Shlapentokh-Rothman M , et al . Language agent tree search unifies reasoning, acting, and planning in language models [C ] // Proceedings of the 41st International Conference on Machine Learning . New York : PMLR Press , 2024 : 1 - 23 .
Liu X , Yu H , Zhang H , et al . AgentBench: evaluating LLMs as agents [C ] // Proceedings of the Twelfth International Conference on Learning Representations . Amherst, Massachusetts : OpenReview , 2024 .
Anderson P , Chang A , Chaplot D S , et al . On evaluation of embodied navigation agents [PP ] . V1. arXiv ( 2018-07-18 )[ 2026-06-10 ] . arXiv: 1807.06757 .
Wu Y , Tang X , Mitchell T , et al . SmartPlay: abenchmark for LLMs as intelligent agents [C ] // Proceedings of the 12th International Conference on Learning Representations . Amherst, Massachusetts : OpenReview , 2024 .
Ethayarajh K . How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT-2 embeddings [C ] // Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) . Stroudsburg, PA : ACL Press , 2019 : 55 - 65 .
Hong S , Zhuge M , Chen J , et al . MetaGPT: meta programming for a multi-agent collaborative framework [C ] // Proceedings of the 12th International Conference on Learning Representations . Amherst, Massachusetts : OpenReview , 2024 .
Ghanem B , Hammoud H , Itani H , et al . CAMEL: communicative agents for "mind" exploration of large language model society [C ] // Proceedings of the Advances in Neural Information Processing Systems 36 . Red Hook : Curran Associates , 2023 : 51991 - 52008 .
0
浏览量
0
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621