Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Computer Science editorial

Open AccessOA2026

Temporally Consistent Graph Q-Networks for Intelligent Network Control

A multi-agent reinforcement learning approach for energy-efficient orchestration in mobile networks
Zacharias Veiksaar; Maxime Boutonยท 2026ยท DOI 10.48550/arXiv.2606.13848

The core problem

Mobile networks are growing in complexity, with next-generation networks expected to support increasing traffic loads and diverse services. Optimizing antenna parameters under dynamic or changing objectives is increasingly challenging. The paper proposes a novel multi-agent reinforcement learning (MARL) algorithm for high-level control and orchestration of mobile networks. The Temporally Consistent Graph Q-Network (TC-GQN) learns a self-predicting representation of the whole network that is task-independent and aggregates information from all base-stations. A graph neural network is trained using a global reward function to assign coordinated local actions based on the learned encoding of the global network state. The algorithm is evaluated in a simulated environment to orchestrate an energy-saving feature across multiple sectors and multiple carriers under different quality of service (QoS) constraints.

Innovation

The algorithm is evaluated in a simulated environment to orchestrate an energy-saving feature across multiple sectors and multiple carriers under different quality of service (QoS) constraints. The proposed algorithm outperforms state-of-the-art graph-based baselines and a competitive rule-based controller by improving hardware sleep time while maintaining QoS. Moreover, the learned representation enables rapid adaptation to changing intents. Quantitative results are not provided in the abstract, but the improvements are reported in terms of hardware sleep time and QoS maintenance.
Mobile networks are growing in complexity, with next-generation networks expected to support increasing traffic loads and diverse services. Optimizing antenna parameters under dynamic or changing objectives is increasingly challenging. The paper proposes a novel multi-agent reinforcement learning (MARL) algorithm for high-level control and orchestration of mobile networks. The Temporally Consistent Graph Q-Network (TC-GQN) learns a self-predicting representation of the whole network that is task-independent and aggregates information from all base-stations. A graph neural network is trained using a global reward function to assign coordinated local actions based on the learned encoding of the global network state. The algorithm is evaluated in a simulated environment to orchestrate an energy-saving feature across multiple sectors and multiple carriers under different quality of service (QoS) constraints.
The TC-GQN algorithm is a multi-agent reinforcement learning approach designed for high-level control and orchestration of mobile networks. It learns a self-predicting representation of the entire network that is task-independent and aggregates information from all base-stations. A graph neural network (GNN) is trained using a global reward function to assign coordinated local actions based on the learned encoding of the global network state. The architecture can be represented as follows:

Why it matters

The TC-GQN algorithm addresses the challenge of optimizing antenna parameters under dynamic objectives by learning a task-independent, self-predicting representation of the network. This representation aggregates information from all base stations, enabling coordinated local actions that improve energy efficiency without sacrificing QoS. The use of a global reward function allows the GNN to learn a holistic view of the network, which is crucial for multi-agent coordination. The temporal consistency of the representation facilitates rapid adaptation to changing intents, making the approach suitable for next-generation networks with diverse services. The results demonstrate the potential of MARL with graph neural networks for intelligent network control, outperforming both graph-based baselines and rule-based controllers. Future work may explore scalability and real-world deployment.

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ