Yuqi Ping, Tianhao Liang, Huahao Ding, Guangyu Lei, Junwei Wu, Xuan Zou, Kuan Shi, Rui Shao, Chiya Zhang, Weizheng Zhang, Weijie Yuan, Tingting Zhang
Recent breakthroughs in multimodal large language models (MLLMs) have endowed AI systems with unified perception, reasoning, and natural-language interaction across text, image, and video modalities. Meanwhile, unmanned aerial vehicle (UAV) swarms are increasingly deployed in dynamic, safety-critical missions that demand rapid situational awareness and autonomous adaptation. This paper explores potential solutions for integrating MLLMs with UAV swarms to enhance intelligence and adaptability across diverse tasks. Specifically, we first outline the fundamental architectures and functions of UAVs and MLLMs. Then, we present a comprehensive framework for an MLLM-enabled UAV swarm system and discuss the opportunities it offers. Next, we demonstrate the capabilities of the proposed framework through a forest firefighting case study that includes both simulation and real-world experiments. Finally, we discuss the challenges and future research directions for MLLM-enabled UAV swarms. An illustrative video of the experiments is available athttps://youtu.be/zwnB9ZSa5A4.