A Distributed and Scalable Testbed for Developing Intelligent UAV Swarm Controllers Based on Curiosity-Driven Machine Learning
This research project developed a novel distributed networking-based Unmanned Aerial Vehicle (UAV) intelligent system to enable real-time swarm UAV target tracking. The Proximal Policy Optimization (PPO) model, along with Intrinsic Curiosity Module (ICM), was developed and integrated with JSBSim for controlling Reinforcement Learning (RL) agents in the simulated environment of FlightGear. The PPO-ICM model was interfaced with the testbed’s Asynchronous Advantage Actor-Critic (A3C) model, facilitating the control of multiple tracking UAVs as a coordinated swarm. The target UAV, in contrast, is operated by the Advantage Actor-Critic (A2C) model. This study incorporates a distributed framework that enables parallel execution of an experimental simulation across multiple systems interconnected via computer networking. The effectiveness of this approach is demonstrated using four interconnected systems, ensuring efficient UAV coordination and scalability. This methodology has overcome the problem of policy degeneration of RL agents. The graphical results evaluate the performance of PPO-ICM, showing significant improvements in scalability, adaptability, and computational efficiency. This research presents a novel scalable and distributed system for swarm UAV target tracking by applying novel Machine Learning (ML) techniques in dynamic and distributed environments.