Guangyuan Liu, Yinqiu Liu, Ruichen Zhang, Hongyang Du, Dusit Niyato, Zehui Xiong, Sumei Sun, Abbas Jamalipour
The rapid growth of multimodal agentic AI and LLMs enables richer perception and decision making, yet bandwidth-limited multi-agent links hinder timely exchange of task-critical semantics. Existing approaches based on static compression or heuristic selection do not adapt to the receiver's current query or to contention on a shared channel, which leads to redundant transmission or missing task-critical details. We propose Retrieval-Augmented Multimodal Semantic Communication (RAMSemCom), a receiver-driven framework that iteratively retrieves only high-resolution image patches most relevant to the active task when the initial downsampled summary is insufficient. A centralized Deep Reinforcement Learning (DRL) scheduler at the roadside unit coordinates per-agent patch budgets and timing under shared bandwidth and round deadlines, using semantic relevance and channel state to orchestrate concurrent requests. Retrieved patches are overlaid via a spatial-awareness layout to preserve scene geometry without full-frame transfer. In a multi-vehicle urban driving case study, RAMSemCom achieves earlier task completion and better performance than strong non-learning and token-level baselines under identical bandwidth and latency budgets.