Computer Science editorial
Open AccessOA2026
TRUAV: Distributed Multi-Agent Reinforcement Learning for Trajectory Planning and Routing Enhancement in UAV-Aided IoT-Enabled VANETs
Independent tabular Q-learning enables scalable UAV trajectory planning and routing in dense urban VANETs without global state aggregation
Muhammad Umar Farooq Qaisar; Lin Zhang; Zhen Chen; Wajdy Othman; Shehzad Ashraf Chaudhry; Chang Liuยท 2026ยท DOI 10.48550/arXiv.2607.23734
The core problem
Unmanned aerial vehicles (UAVs) have emerged as a key enabler of next-generation Internet of Things (IoT) ecosystems, offering flexible aerial relaying to extend connectivity across dynamic vehicular ad hoc networks (VANETs) in smart city environments. Conventional centralized approaches for UAV trajectory planning require continuous global network state aggregation, making them impractical under bandwidth and energy constraints typical of dense urban deployments. This article presents TRUAV, a distributed multi-agent reinforcement learning framework based on independent tabular Q-learning for joint UAV trajectory planning and routing enhancement in UAV-aided VANETs. The core problem addressed is the tension between the need for coordinated, routing-aware UAV positioning and the prohibitive cost of centralized global state exchange in bandwidth- and energy-constrained urban settings.
Innovation
Numerical simulations over a large urban area with 200 mobile vehicles show that the proposed TRUAV framework achieves network coverage and packet delivery ratios comparable to centralized deep reinforcement learning methods, while also improving relay delay and energy efficiency. The evaluation considered a dense urban deployment scenario with dynamic vehicular mobility and IoT traffic. Key performance indicators included network coverage, packet delivery ratio (PDR), relay delay, and energy efficiency. TRUAV matched centralized deep reinforcement learning baselines on coverage and PDR, and outperformed them on relay delay and energy efficiency, demonstrating that distributed independent tabular Q-learning can achieve competitive performance without global state aggregation.
Unmanned aerial vehicles (UAVs) have emerged as a key enabler of next-generation Internet of Things (IoT) ecosystems, offering flexible aerial relaying to extend connectivity across dynamic vehicular ad hoc networks (VANETs) in smart city environments. Conventional centralized approaches for UAV trajectory planning require continuous global network state aggregation, making them impractical under bandwidth and energy constraints typical of dense urban deployments. This article presents TRUAV, a distributed multi-agent reinforcement learning framework based on independent tabular Q-learning for joint UAV trajectory planning and routing enhancement in UAV-aided VANETs. The core problem addressed is the tension between the need for coordinated, routing-aware UAV positioning and the prohibitive cost of centralized global state exchange in bandwidth- and energy-constrained urban settings.
TRUAV equips each UAV with a local Q-learning agent that operates purely on locally observable information, including vehicle density, packet queue states, and neighbor UAV positions, thereby eliminating the need for global state exchange. A potential-game-inspired reward design encourages spatial diversity and routing-aware UAV positioning among interacting agents while accounting for energy consumption.
Why it matters
The results indicate that independent tabular Q-learning, when combined with a potential-game-inspired reward, can effectively coordinate multiple UAVs for joint trajectory planning and routing enhancement in UAV-aided IoT-enabled VANETs. The elimination of global state exchange reduces communication overhead and energy consumption, addressing a key impracticality of centralized approaches under bandwidth and energy constraints. The potential-game-inspired reward encourages spatial diversity and routing-aware positioning, mitigating conflicts among interacting agents. However, the use of tabular Q-learning may face scalability challenges as the state and action spaces grow with the number of UAVs and vehicles. The article also discusses emerging challenges and future research directions for distributed multi-agent UAV-assisted IoT systems, including scalability, partial observability, and integration with emerging IoT and vehicular standards.
Who should read this
CS practitioners and researchers
Opening member contentโฆ