Gleb Merkulov, Eran Iceland, Shay Michaeli, Oren Gal, Ariel Barel, Tal Shima
This paper considers a multi-agent shoot–shoot–look engagement scenario, in which multiple pursuers act against multiple targets with predefined motion. The pursuers are arranged in two successive waves—the first wave engages the a priori allocated targets directly, and the second wave trails behind to assist with the pursuit of the targets that survive the first wave. The objective is to maximize the number of intercepted targets under the given time constraints. To facilitate the intercept performance, we propose a methodology in which the second-wave pursuers are guided to intermediate virtual targets that allow several reallocation options when the actual first-wave engagement outcomes become known. The choice of the virtual targets affects the time and maneuver required from the pursuers during the engagement, which influence the intercept probabilities. We formulate the problem of second-wave allocation as a stochastic Markov decision process. Greedy and reinforcement-learning-based solutions are proposed for decentralized sequential allocation. The simulation results show that both proposed solutions are close, suggesting that planning multiple steps ahead, as done in the reinforcement-learning-based solution, is not necessary in this class of problems. The analysis of theoretical bounds further supports this claim.