{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,1]],"date-time":"2026-07-01T06:25:20Z","timestamp":1782887120698,"version":"3.54.5"},"reference-count":44,"publisher":"Institution of Engineering and Technology (IET)","issue":"4","license":[{"start":{"date-parts":[[2024,3,28]],"date-time":"2024-03-28T00:00:00Z","timestamp":1711584000000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61872171"],"award-info":[{"award-number":["61872171"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["ietresearch.onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["CAAI Trans on Intel Tech"],"published-print":{"date-parts":[[2024,8]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Multi\u2010agent reinforcement learning relies on reward signals to guide the policy networks of individual agents. However, in high\u2010dimensional continuous spaces, the non\u2010stationary environment can provide outdated experiences that hinder convergence, resulting in ineffective training performance for multi\u2010agent systems. To tackle this issue, a novel reinforcement learning scheme, Mutual Information Oriented Deep Skill Chaining (MioDSC), is proposed that generates an optimised cooperative policy by incorporating intrinsic rewards based on mutual information to improve exploration efficiency. These rewards encourage agents to diversify their learning process by engaging in actions that increase the mutual information between their actions and the environment state. In addition, MioDSC can generate cooperative policies using the options framework, allowing agents to learn and reuse complex action sequences and accelerating the convergence speed of multi\u2010agent learning. MioDSC was evaluated in the multi\u2010agent particle environment and the StarCraft multi\u2010agent challenge at varying difficulty levels. The experimental results demonstrate that MioDSC outperforms state\u2010of\u2010the\u2010art methods and is robust across various multi\u2010agent system tasks with high stability.<\/jats:p>","DOI":"10.1049\/cit2.12322","type":"journal-article","created":{"date-parts":[[2024,3,28]],"date-time":"2024-03-28T05:44:56Z","timestamp":1711604696000},"page":"1014-1030","update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Mutual information oriented deep skill chaining for multi\u2010agent reinforcement learning"],"prefix":"10.1049","volume":"9","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1637-1511","authenticated-orcid":false,"given":"Zaipeng","family":"Xie","sequence":"first","affiliation":[{"name":"Key Laboratory of Water Big Data Technology of Ministry of Water Resources Hohai University  Nanjing China"},{"name":"College of Computer and Information Hohai University  Nanjing China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Cheng","family":"Ji","sequence":"additional","affiliation":[{"name":"College of Computer and Information Hohai University  Nanjing China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-9983-5736","authenticated-orcid":false,"given":"Chentai","family":"Qiao","sequence":"additional","affiliation":[{"name":"College of Computer and Information Hohai University  Nanjing China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"WenZhan","family":"Song","sequence":"additional","affiliation":[{"name":"Center for Cyber\u2010Physical Systems University of Georgia  Athens Georgia USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6593-0987","authenticated-orcid":false,"given":"Zewen","family":"Li","sequence":"additional","affiliation":[{"name":"Information Networking Institute Carnegie Mellon University  Pittsburgh Pennsylvania USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yufeng","family":"Zhang","sequence":"additional","affiliation":[{"name":"College of Computer and Information Hohai University  Nanjing China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yujing","family":"Zhang","sequence":"additional","affiliation":[{"name":"Department of Electrical and Systems Engineering University of Pennsylvania  Philadelphia Pennsylvania USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"265","published-online":{"date-parts":[[2024,3,28]]},"reference":[{"key":"e_1_2_10_2_1","first-page":"1","article-title":"Multi\u2010agent deep reinforcement learning: a survey","author":"Gronauer S.","year":"2022","journal-title":"Artif. Intell. Rev."},{"key":"e_1_2_10_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3128584"},{"key":"e_1_2_10_4_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10458\u2010019\u201009421\u20101"},{"key":"e_1_2_10_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/comst.2022.3160697"},{"key":"e_1_2_10_6_1","first-page":"11352","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Su J.","year":"2021"},{"key":"e_1_2_10_7_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-28929-8"},{"key":"e_1_2_10_8_1","volume-title":"8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia","author":"Baker B.","year":"2020"},{"key":"e_1_2_10_9_1","first-page":"330","volume-title":"Proceedings of the Tenth International Conference on Machine Learning","author":"Tan M.","year":"1993"},{"key":"e_1_2_10_10_1","doi-asserted-by":"publisher","DOI":"10.1017\/s0269888912000057"},{"key":"e_1_2_10_11_1","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Foerster J.","year":"2018"},{"key":"e_1_2_10_12_1","article-title":"Multi\u2010agent actor\u2010critic for mixed cooperative\u2010competitive environments","volume":"30","author":"Lowe R.","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_2_10_13_1","article-title":"A unified game\u2010theoretic approach to multiagent reinforcement learning","volume":"30","author":"Lanctot M.","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_2_10_14_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2016.01.031"},{"issue":"1","key":"e_1_2_10_15_1","first-page":"7234","article-title":"Monotonic value function factorisation for deep multi\u2010agent reinforcement learning","volume":"21","author":"Rashid T.","year":"2020","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_2_10_16_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i13.17353"},{"key":"e_1_2_10_17_1","first-page":"1015","volume-title":"Advances in Neural Information Processing Systems, NeurIPS","author":"Konidaris G.D.","year":"2009"},{"key":"e_1_2_10_18_1","volume-title":"International Conference on Learning Representations, ICLR","author":"Bagaria A.","year":"2020"},{"key":"e_1_2_10_19_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20503-3_17"},{"key":"e_1_2_10_20_1","first-page":"2085","volume-title":"Proceedings of International Conference on Autonomous Agents and MultiAgent Systems, Ser. AAMAS","author":"Sunehag P.","year":"2018"},{"key":"e_1_2_10_21_1","article-title":"Playing atari with deep reinforcement learning","volume":"1312","author":"Mnih V.","year":"2013","journal-title":"CoRR"},{"key":"e_1_2_10_22_1","first-page":"457","volume-title":"Proceedings of the 28th International Joint Conference on Artificial Intelligence, IJCAI","author":"Liu Y.","year":"2019"},{"key":"e_1_2_10_23_1","first-page":"11308","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Phan T.","year":"2021"},{"key":"e_1_2_10_24_1","first-page":"125","volume-title":"Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI","author":"Danassis P.","year":"2021"},{"key":"e_1_2_10_25_1","first-page":"5887","volume-title":"Proceedings of the 36th International Conference on Machine Learning, ICML","author":"Son K.","year":"2019"},{"key":"e_1_2_10_26_1","first-page":"1726","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Bacon P.","year":"2017"},{"key":"e_1_2_10_27_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0004-3702(99)00052-1"},{"key":"e_1_2_10_28_1","first-page":"1566","volume-title":"Proceedings of the 19th International Conference on Autonomous Agents and MultiAgent Systems","author":"Yang J.","year":"2020"},{"key":"e_1_2_10_29_1","first-page":"3675","volume-title":"Advances in Neural Information Processing Systems, NeurIPS","author":"Kulkarni T.D.","year":"2016"},{"key":"e_1_2_10_30_1","first-page":"3540","volume-title":"International Conference on Machine Learning, PMLR","author":"Vezhnevets A.S.","year":"2017"},{"key":"e_1_2_10_31_1","first-page":"9041","volume-title":"Proceedings of the 39th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research","author":"Hu S.","year":"2022"},{"key":"e_1_2_10_32_1","volume-title":"International Conference on Learning Representations, ICLR","author":"Sharma A.","year":"2019"},{"key":"e_1_2_10_33_1","first-page":"2950","volume-title":"Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI 2021","author":"Seurin M.","year":"2021"},{"key":"e_1_2_10_34_1","first-page":"6826","volume-title":"International Conference on Machine Learning","author":"Liu I.\u2010J.","year":"2021"},{"key":"e_1_2_10_35_1","first-page":"15","article-title":"Updet: universal multi\u2010agent reinforcement learning via policy decoupling with transformers","author":"Hu S.","year":"2021","journal-title":"International Conference on Representation Learning"},{"key":"e_1_2_10_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/tnnls.2023.3236361"},{"key":"e_1_2_10_37_1","first-page":"1928","volume-title":"Proceedings of the International Conference on Machine Learning, ICML","author":"Mnih V.","year":"2016"},{"key":"e_1_2_10_38_1","volume-title":"International Conference on Learning Representations, ICLR","author":"Engstrom L.","year":"2020"},{"key":"e_1_2_10_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/584091.584093"},{"key":"e_1_2_10_40_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00145\u2010010\u20109084\u20108"},{"key":"e_1_2_10_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01102"},{"key":"e_1_2_10_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISIT.2008.4595271"},{"key":"e_1_2_10_43_1","unstructured":"Bakerjc\u2010bgner:MioDSC\u2010with\u2010StarCraft\u2010environment(2023). [Online].https:\/\/github.com\/Bakerjc\u2010bgner\/MioDSC"},{"key":"e_1_2_10_44_1","first-page":"2186","volume-title":"Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems, AAMAS \u201919, Montreal, QC, Canada","author":"Samvelyan M.","year":"2019"},{"key":"e_1_2_10_45_1","unstructured":"Brockman G.et\u00a0al.:OpenAI Gym. arXiv preprint arXiv:1606.01540 (2016)"}],"container-title":["CAAI Transactions on Intelligence Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/pdf\/10.1049\/cit2.12322","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,28]],"date-time":"2025-10-28T06:48:01Z","timestamp":1761634081000},"score":1,"resource":{"primary":{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/10.1049\/cit2.12322"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,3,28]]},"references-count":44,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2024,8]]}},"alternative-id":["10.1049\/cit2.12322"],"URL":"https:\/\/doi.org\/10.1049\/cit2.12322","archive":["Portico"],"relation":{},"ISSN":["2468-6557","2468-2322"],"issn-type":[{"value":"2468-6557","type":"print"},{"value":"2468-2322","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,3,28]]},"assertion":[{"value":"2023-04-29","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-08-09","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-03-28","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}