{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,14]],"date-time":"2026-07-14T15:59:06Z","timestamp":1784044746300,"version":"3.55.0"},"reference-count":29,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2024,11,29]],"date-time":"2024-11-29T00:00:00Z","timestamp":1732838400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Fundamental Research Funds for the Provincial Universities of Zhejiang","award":["GK249909299001-012"],"award-info":[{"award-number":["GK249909299001-012"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>The traditional maneuver decision-making approaches are highly dependent on accurate and complete situation information, and their decision-making quality becomes poor when opponent information is occasionally missing in complex electromagnetic environments. In order to solve this problem, an autonomous maneuver decision-making approach is developed based on deep reinforcement learning (DRL) architecture. Meanwhile, a Transformer network is integrated into the actor and critic networks, which can find the potential dependency relationships among the time series trajectory data. By using these relationships, the information loss is partially compensated, which leads to maneuvering decisions being more accurate. The issues of limited experience samples, low sampling efficiency, and poor stability in the agent training state appear when the Transformer network is introduced into DRL. To address these issues, the measures of designing an effective decision-making reward, a prioritized sampling method, and a dynamic learning rate adjustment mechanism are proposed. Numerous simulation results show that the proposed approach outperforms the traditional DRL algorithms, with a higher win rate in the case of opponent information loss.<\/jats:p>","DOI":"10.3390\/e26121036","type":"journal-article","created":{"date-parts":[[2024,12,2]],"date-time":"2024-12-02T03:50:57Z","timestamp":1733111457000},"page":"1036","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["An Intelligent Maneuver Decision-Making Approach for Air Combat Based on Deep Reinforcement Learning and Transformer Networks"],"prefix":"10.3390","volume":"26","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-9543-6435","authenticated-orcid":false,"given":"Wentao","family":"Li","sequence":"first","affiliation":[{"name":"School of Automation, Hangzhou Dianzi University, Hangzhou 310018, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Feng","family":"Fang","sequence":"additional","affiliation":[{"name":"School of Automation, Hangzhou Dianzi University, Hangzhou 310018, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dongliang","family":"Peng","sequence":"additional","affiliation":[{"name":"School of Automation, Hangzhou Dianzi University, Hangzhou 310018, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shuning","family":"Han","sequence":"additional","affiliation":[{"name":"China Academy of Launch Vehicle Technology, Beijing 100076, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2024,11,29]]},"reference":[{"key":"ref_1","first-page":"39","article-title":"Development and illustrative applications of an air combat engagement database","volume":"44","author":"Ma","year":"2023","journal-title":"Chin. J. Aeronaut."},{"key":"ref_2","first-page":"47","article-title":"Continuous Decision-making Method for Autonomous Air Combat","volume":"13","author":"Shan","year":"2022","journal-title":"Adv. Aeronaut. Sci. Eng."},{"key":"ref_3","first-page":"15","article-title":"New Concepts of Future Air Warfare and the Challenges for Its Realization","volume":"27","author":"Fan","year":"2020","journal-title":"Aero Weapon."},{"key":"ref_4","first-page":"71","article-title":"Research on UAV Air Combat Decision Making Based on DRL and Differential Games","volume":"46","author":"Yang","year":"2021","journal-title":"Fire Control Command Control"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"405","DOI":"10.1007\/BF00934680","article-title":"A stochastic homicidal chauffeur pursuit-evasion differential game","volume":"34","author":"Pachter","year":"1981","journal-title":"J. Optim. Theory Appl."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"65","DOI":"10.1016\/S1474-6670(17)55066-1","article-title":"Unified approach for two-target game analysis","volume":"20","author":"Shinar","year":"1987","journal-title":"IFAC Proc. Vol."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"204","DOI":"10.5139\/IJASS.2016.17.2.204","article-title":"Differential game based air combat maneuver generation using scoring function matrix","volume":"17","author":"Park","year":"2016","journal-title":"Int. J. Aeronaut. Space Sci."},{"key":"ref_8","first-page":"173","article-title":"Attact-defense Confrontation Simulation of Air Combat Based on Game-matrix Approach","volume":"33","author":"Che","year":"2015","journal-title":"Flight Dyn."},{"key":"ref_9","first-page":"77","article-title":"Decision-Making of Air Combat Maneuvering Based on APF and PSO","volume":"20","author":"Zhang","year":"2013","journal-title":"Electron. Opt. Control"},{"key":"ref_10","first-page":"2447","article-title":"Maneuvering Decision-making Method of UAV Based on Approximate Dynamic Programming","volume":"40","author":"Huang","year":"2018","journal-title":"J. Electron. Inf. Technol."},{"key":"ref_11","first-page":"1613","article-title":"Maneuver Decision of Autonomous Air Combat of Unmanned Combat Aerial Vehicle Based on Deep Neural Network","volume":"41","author":"Zhang","year":"2020","journal-title":"Acta Armamentarii"},{"key":"ref_12","first-page":"14","article-title":"A Decision-making of Advanced Fighter Cooperative Attack and Defense Based on Fuzzy Genetic Algorithm","volume":"45","author":"Xu","year":"2020","journal-title":"Fire Control Command Control"},{"key":"ref_13","first-page":"1063","article-title":"UAV Air Combat Maneuvering Decision Based on Intuitionistic Fuzzy Game Theory","volume":"41","author":"Li","year":"2019","journal-title":"Syst. Eng. Electron."},{"key":"ref_14","first-page":"33","article-title":"Maneuver decision of UCAV in air combat based on deep reinforcement learning","volume":"53","author":"Li","year":"2021","journal-title":"J. Harbin Inst. Technol."},{"key":"ref_15","first-page":"88","article-title":"Air combat maneuvering decision method based on APF-DQN","volume":"39","author":"Zhang","year":"2021","journal-title":"Flight Dyn."},{"key":"ref_16","first-page":"352","article-title":"Maneuvering strategy generation algorithm for multi-UAV in close-range air combat based on deep reinforcement learning and self-play","volume":"39","author":"Kong","year":"2022","journal-title":"Control Theory Appl."},{"key":"ref_17","first-page":"24","article-title":"Autonomous Air Combat Decision-making Algorithm of UAVs Based on SAC algorithm","volume":"44","author":"Li","year":"2022","journal-title":"Command Control Simul."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"64","DOI":"10.1049\/cit2.12109","article-title":"Autonomous air combat decision-making of UAV based on parallel self play reinforcement learning","volume":"8","author":"Li","year":"2023","journal-title":"CAAI Trans. Intell. Technol."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Fan, Z., Xu, Y., Kang, Y., and Luo, D. (2022). Air Combat Maneuver Decision method Based on A3C Deep Reinforcement Learning. Machines, 10.","DOI":"10.3390\/machines10111033"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"1421","DOI":"10.23919\/JSEE.2021.000121","article-title":"UAV cooperative air combat maneuver decision based on multi-agent reinforcement learning","volume":"32","author":"Zhang","year":"2021","journal-title":"J. Syst. Eng. Electron."},{"key":"ref_21","first-page":"19","article-title":"Maneuvering Decision of UCAV in Close Air Combat Based on LSTM-PPO Algorithm","volume":"23","author":"Ding","year":"2022","journal-title":"J. Air Force Eng. Univ. (Nat. Sci. Ed.)"},{"key":"ref_22","first-page":"97","article-title":"Intelligent Maneuvering Decision of Unmanned Combat Aircraft Based on LSTM-Dueling DQN","volume":"6","author":"Hu","year":"2021","journal-title":"Tactical Missile Technol."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"32282","DOI":"10.1109\/ACCESS.2021.3060426","article-title":"Application of Deep Reinforcement Learning in Maneuver Planning of Beyond-Visual-Range Air Combat","volume":"9","author":"Hu","year":"2021","journal-title":"IEEE Access"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"30819","DOI":"10.1109\/ACCESS.2023.3262023","article-title":"Long Short-Term Memory-Based Neural Networks for Missile Maneuvers Trajectories Prediction","volume":"11","author":"Lui","year":"2023","journal-title":"IEEE Access"},{"key":"ref_25","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems 30, Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, 4\u20139 December 2017, Curran Associates Inc."},{"key":"ref_26","unstructured":"Zambaldi, V., Raposo, D., Santoro, A., Bapst, V., Li, Y., Babuschkin, I., Tuyls, K., Reichert, D., Lillicrap, T., and Lockhart, E. (May, January 30). Deep reinforcement learning with relational inductive biases. Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canada."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"350","DOI":"10.1038\/s41586-019-1724-z","article-title":"Grandmaster level in Starcraft II using multi-agent reinforcement learning","volume":"575","author":"Vinyals","year":"2019","journal-title":"Nature"},{"key":"ref_28","unstructured":"Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I. (2021). Decision transformer: Reinforcement learning via sequence modeling. Advances in Neural Information Processing Systems 34, Proceedings of the 35th Conference on Neural Information Processing Systems, Online, 6\u201314 December 2021, Curran Associates Inc."},{"key":"ref_29","first-page":"797","article-title":"LSTM-MADDPG multi-agent cooperative decision algorithm based on asynchronous collaborative update","volume":"54","author":"Gao","year":"2024","journal-title":"J. Jilin Univ. (Eng. Technol. Ed.)"}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/26\/12\/1036\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T16:43:09Z","timestamp":1760114589000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/26\/12\/1036"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,11,29]]},"references-count":29,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2024,12]]}},"alternative-id":["e26121036"],"URL":"https:\/\/doi.org\/10.3390\/e26121036","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,11,29]]}}}