EDBT 2026 Demo / reviewers in the wild / expert
Pedro Enrique Iturria-Rivera
dblp:254/0175
· DBLP profile ↗
15ranked-venue papers
6as first author
14since 2021 · last 2025
0000-0002-8757-2639ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 11 · 3 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Prioritized Value-Decomposition Network for Explainable AI-Enabled Network Slicing
Shavbo Salehi, Pedro Enrique Iturria-Rivera, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Yigit Ozcan, Melike Erol-Kantarci |
ICC | 2 |
| 2025 | LLM-Based Intent Processing and Network Optimization Using Attention-Based Hierarchical Reinforcement LearningabstractIntent-based network automation is a promising tool that enables easier network management; however, certain challenges must be addressed effectively. These are: 1) processing intents, i.e., identification of logic and necessary parameters to fulfill an intent, 2) validating an intent to align it with current network status, and 3) satisfying intents via network optimizing applications. This paper addresses these points via a three-fold strategy to introduce intent-based automation for modern 5G architectures. First, intents are processed via a lightweight Large Language Model (LLM). Secondly, once an intent is processed, it is validated against future incoming traffic volume profiles (high or low). Finally, a series of network optimization applications has been developed. With their machine learning-based functionalities, they can improve certain key performance indicators such as throughput, delay, and energy efficiency. In the final stage, using an attention-based hierarchical reinforcement learning algorithm, these applications are optimally initiated to satisfy the intent of an operator. Our simulations show that the proposed method can achieve at least a 12% increase in throughput, a 17.1% increase in energy efficiency, and a 26.5% decrease in network delay compared to the baseline algorithms. Md Arafat Habib, Pedro Enrique Iturria-Rivera, Yigit Ozcan, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Melike Erol-Kantarci |
WCNC | 2 |
| 2024 | Self-Play Ensemble Q-learning enabled Resource Allocation for Network SlicingabstractIn 5G networks, network slicing has emerged as a pivotal paradigm to address diverse user demands and service requirements. To meet the requirements, reinforcement learning (RL) algorithms have been utilized widely, but this method has the problem of overestimation and exploration-exploitation trade-offs. To tackle these problems, this paper explores the application of self-play ensemble Q-learning, an extended version of the RL-based technique. Self-play ensemble Q-learning utilizes multiple Q-tables with various exploration-exploitation rates leading to different observations for choosing the most suitable action for each state. Moreover, through self-play, each model endeavors to enhance its performance compared to its previous iterations, boosting system efficiency, and decreasing the effect of overestimation. For performance evaluation, we consider three RL-based algorithms; self-play ensemble Q-learning, double Q-learning, and Q-learning, and compare their performance under different network traffic. Through simulations, we demonstrate the effectiveness of self-play ensemble Q-learning in meeting the diverse demands within 21.92% in latency, 24.22% in throughput, and 23.63% in packet drop rate in comparison with the baseline methods. Furthermore, we evaluate the robustness of self-play ensemble Q-learning and double Q-learning in situations where one of the Q-tables is affected by a malicious user. Our results depicted that the self-play ensemble Q-learning method is more robust against adversarial users and prevents a noticeable drop in system performance, mitigating the impact of users manipulating policies. Shavbo Salehi, Pedro Enrique Iturria-Rivera, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Yigit Ozcan, Melike Erol-Kantarci |
GLOBECOM | 2 |
| 2024 | Extended Reality (XR) Codec Adaptation in 5G using Multi-Agent Reinforcement Learning with Attention Action SelectionabstractExtended Reality (XR) services will revolutionize applications over $5^{\text {th }}$ and $\mathbf{6}^{\text {th }}$ generation wireless networks by providing seamless virtual and augmented reality experiences. These applications impose significant challenges on network infrastructure, which can be addressed by machine learning algorithms due to their adaptability. This paper presents a Multi-Agent Reinforcement Learning (MARL) solution for optimizing codec parameters of XR traffic, comparing it to the Adjust Packet Size (APS) algorithm. Our cooperative multi-agent system uses an Optimistic Mixture of Q-Values ($\mathbf{O Q M I X}$) approach for handling Cloud Gaming (CG), Augmented Reality (AR), and Virtual Reality (VR) traffic. Enhancements include an attention mechanism and slate-Markov Decision Process (MDP) for improved action selection. Simulations show our solution outperforms APS with average gains of $30.1 \%, 15.6 \%, 16.5 \% 50.3 \%$ in XR index, jitter, delay, and Packet Loss Ratio (PLR), respectively. APS tends to increase throughput but also packet losses, whereas oQMIX reduces PLR, delay, and jitter while maintaining goodput. Pedro Enrique Iturria-Rivera, Raimundas Gaigalas, Medhat H. M. Elsayed, Majid Bavand, Yigit Ozcan, Melike Erol-Kantarci |
PIMRC | 1 |
| 2023 | Traffic Steering for 5G Multi-RAT Deployments using Deep Reinforcement LearningabstractIn 5G non-standalone mode, traffic steering is a critical technique to take full advantage of 5G new radio while optimizing dual connectivity of 5G and LTE networks in multiple radio access technology (RAT). An intelligent traffic steering mechanism can play an important role to maintain seamless user experience by choosing appropriate RAT (5G or LTE) dynamically for a specific user traffic flow with certain QoS requirements. In this paper, we propose a novel traffic steering mechanism based on Deep Q-learning that can automate traffic steering decisions in a dynamic environment having multiple RATs, and maintain diverse QoS requirements for different traffic classes. The proposed method is compared with two baseline algorithms: a heuristic-based algorithm and Q-learning-based traffic steering. Compared to the Q-learning and heuristic baselines, our results show that the proposed algorithm achieves better performance in terms of 6% and 10% higher average system throughput, and 23% and 33% lower network delay, respectively. Md Arafat Habib, Hao Zhou 0013, Pedro Enrique Iturria-Rivera, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Steve Furr, Melike Erol-Kantarci |
CCNC | 3 |
| 2023 | Split Learning for Sensing-Aided Single and Multi-Level Beam Selection in Multi-Vendor RANabstractProper and efficient beam selection is of great importance to unleash the full potential of mmWave communications. Traditionally, each candidate beam is evaluated using reference signals (beam sweeping), however, the exhaustive search method can be time-consuming with high signaling overhead. To avoid such problems, in B5G and 6G, sensing information is considered to be used, as in Integrated Sensing and Communication (ISAC) solutions, and Machine Learning (ML) methods can be applied to map sensing data inputs to an optimal beam index. When using sensing information sources external to the Radio Access Network (RAN) in a multi-vendor disaggregated environment, those methods need to account for issues such as privacy and data ownership. In this work, we apply multi-modal sensing information to the beam selection task. Specifically, we propose a multi-modal sensing-aided ML strategy based on Split Learning (SL) that can cope with deployment challenges in novel RAN architectures. Moreover, the method is applied to single and multi-level beam selection decisions, where the latter considers the case of hierarchical codebook structures. With the proposed approach, accuracy levels above 90% can be achieved while overhead diminishes by 85% or more. SL achieves comparable performance with centralized learning-based strategies, with the added value of accounting for privacy and data ownership issues. We also show that sensing-aided ML-based beam selection decisions in multi-level codebooks are more effective when applied to their first level. Ycaro Dantas, Pedro Enrique Iturria-Rivera, Hao Zhou 0013, Yigit Ozcan, Majid Bavand, Medhat H. M. Elsayed, Raimundas Gaigalas, Melike Erol-Kantarci |
GLOBECOM | 2 |
| 2023 | Beam Selection for Energy-Efficient mmWave Network Using Advantage Actor Critic LearningabstractThe growing adoption of mmWave frequency bands to realize the full potential of 5G, turns beamforming into a key enabler for current and next-generation wireless technologies. Many mmWave networks rely on beam selection with Grid-of-Beams (GoB) approach to handle user-beam association. In beam selection with GoB, users select the appropriate beam from a set of pre-defined beams and the overhead during the beam selection process is a common challenge in this area. In this paper, we propose an Advantage Actor Critic (A2C) learning-based framework to improve the GoB and the beam selection process, as well as optimize transmission power in a mmWave network. The proposed beam selection technique allows performance improvement while considering transmission power improves Energy Efficiency (EE) and ensures the coverage is maintained in the network. We further investigate how the proposed algorithm can be deployed in a Service Management and Orchestration (SMO) platform. Our simulations show that A2C-based joint optimization of beam selection and transmission power is more effective than using Equally Spaced Beams (ESB) and fixed power strategy, or optimization of beam selection and transmission power disjointly. Compared to the ESB and fixed transmission power strategy, the proposed approach achieves more than twice the average EE in the scenarios under test and is closer to the maximum theoretical EE. Ycaro Dantas, Pedro Enrique Iturria-Rivera, Hao Zhou 0013, Majid Bavand, Medhat H. M. Elsayed, Raimundas Gaigalas, Melike Erol-Kantarci |
ICC | 2 |
| 2023 | Hierarchical Reinforcement Learning Based Traffic Steering in Multi-RAT 5G DeploymentsabstractIn 5G non-standalone mode, an intelligent traffic steering mechanism can vastly aid in ensuring a smooth user experience by selecting the best radio access technology (RAT) from a multi-RAT environment for a specific traffic flow. In this paper, we propose a novel load-aware traffic steering algorithm based on hierarchical reinforcement learning (HRL) while satisfying the diverse quality of service requirements of different traffic types. HRL can significantly increase system performance using a bi-level architecture having a meta-controller and a controller. In our proposed method, the meta-controller provides an appropriate threshold for load balancing, while the controller performs traffic admission to an appropriate RAT in the lower level. Simulation results show that HRL outperforms a Deep Q-Learning (DQN) and a threshold-based heuristic baseline with 8.49%, 12.52% higher average system throughput and 27.74%, 39.13% lower network delay, respectively. Md Arafat Habib, Hao Zhou 0013, Pedro Enrique Iturria-Rivera, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Yigit Ozcan, Melike Erol-Kantarci |
ICC | 3 |
| 2023 | RL meets Multi-Link Operation in IEEE 802.11be: Multi-Headed Recurrent Soft-Actor Critic-based Traffic AllocationabstractIEEE 802.11be -Extremely High Throughput-, commercially known as Wireless-Fidelity (Wi-Fi) 7 is the newest IEEE 802.11 amendment that comes to address the increasingly throughput hungry services such as Ultra High Definition (4K/8K) Video and Virtual/Augmented Reality (VR/AR). To do so, IEEE 802.11be presents a set of novel features that will boost the Wi-Fi technology to its edge. Among them, Multi-Link Operation (MLO) devices are anticipated to become a reality, leaving Single-Link Operation (SLO) Wi-Fi in the past. To achieve superior throughput and very low latency, a careful design approach must be taken, on how the incoming traffic is distributed in MLO capable devices. In this paper, we present a Reinforcement Learning (RL) algorithm named Multi-Headed Recurrent Soft-Actor Critic (MH-RSAC) to distribute incoming traffic in 802.11be MLO capable networks. Moreover, we compare our results with two non-RL baselines previously proposed in the literature named: Single Link Less Congested Interface (SLCI) and Multi-Link Congestion-aware Load balancing at flow arrivals (MCAA). Simulation results reveal that the MH-RSAC algorithm is able to obtain gains in terms of Throughput Drop Ratio (TDR) up to 35.2% and 6% when compared with the SLCI and MCAA algorithms, respectively. Finally, we observed that our scheme is able to respond more efficiently to high throughput and dynamic traffic such as VR and Web Browsing (WB) when compared with the baselines. Results showed an improvement of the MH-RSAC scheme in terms of Flow Satisfaction (FS) of up to 25.6% and 6% over the the SCLI and MCAA algorithms. Pedro Enrique Iturria-Rivera, Marcel Chenier, Bernard Herscovici, Burak Kantarci, Melike Erol-Kantarci |
ICC | 1 |
| 2023 | Channel Selection for Wi-Fi 7 Multi-Link Operation via Optimistic-Weighted VDN and Parallel Transfer Reinforcement LearningabstractDense and unplanned IEEE 802.11 Wireless Fidelity (Wi-Fi) deployments and the continuous increase of throughput and latency stringent services for users have led to machine learning algorithms to be considered as promising techniques in the industry and the academia. Specifically, the ongoing IEEE 802.11be EHT —Extremely High Throughput, known as Wi-Fi 7— amendment propose, for the first time, Multi-Link Operation (MLO). Among others, this new feature will increase the complexity of channel selection due the novel multiple interfaces proposal. In this paper, we present a Parallel Transfer Reinforcement Learning (PTRL)-based cooperative Multi-Agent Reinforcement Learning (MARL) algorithm named Parallel Transfer Reinforcement Learning Optimistic-Weighted Value Decomposition Networks (oVDN) to improve intelligent channel selection in IEEE 802.11be MLO-capable networks. Additionally, we compare the impact of different parallel transfer learning alternatives and a centralized non-transfer MARL baseline. Two PTRL methods are presented: Multi-Agent System (MAS) Joint Q-function Transfer, where the joint Q-function is transferred and MAS Best/Worst Experience Transfer where the best and worst experiences are transferred among MASs. Simulation results show that oVDNg–only the best experiences are utilized– is the best algorithm variant. Moreover, oVDNgoffers a gain up to 3%, 7.2% and 11% when compared with VDN, VDN-nonQ and non-PTRL baselines. Furthermore, oVDNgexperienced a reward convergence gain in the 5 GHz interface of 33.3% over oVDNband oVDN where only worst and both types of experiences are considered, respectively. Finally, our best PTRL alternative showed an improvement over the non-PTRL baseline in terms of speed of convergence up to 40 episodes and reward up to 135%. Pedro Enrique Iturria-Rivera, Marcel Chenier, Bernard Herscovici, Burak Kantarci, Melike Erol-Kantarci |
PIMRC | 1 |
| 2022 | Competitive Multi-Agent Load Balancing with Adaptive Policies in Wireless NetworksabstractUsing Machine Learning (ML) techniques for the next generation wireless networks have shown promising results in the recent years, due to high learning and adaptation capability of ML algorithms. More specifically, ML techniques have been used for load balancing in Self-Organizing Networks (SON). In the context of load balancing and ML, several studies propose network management automation (NMA) from the perspective of a single and centralized agent. However, a single agent domain does not consider the interaction among the agents. In this paper, we propose a more realistic load balancing approach using novel Multi-Agent Deep Deterministic Policy Gradient with Adaptive Policies (MADDPG-AP) scheme that considers throughput, resource block utilization and latency in the network. We compare our proposal with a single-agent RL algorithm named Clipped Double Q-Learning (CDQL) . Simulation results reveal a significant improvement in latency, packet loss ratio and convergence time. Pedro Enrique Iturria-Rivera, Melike Erol-Kantarci |
CCNC | 1 |
| 2022 | Hierarchical Deep Q-Learning Based Handover in Wireless Networks with Dual Connectivityabstract5G New Radio proposes the usage of frequencies above 10 GHz to speed up LTE's existent maximum data rates. However, the effective size of 5G antennas and consequently its repercussions in the signal degradation in urban scenarios makes it a challenge to maintain stable coverage and connectivity. In order to obtain the best from both technologies, recent dual connectivity solutions have proved their capabilities to improve performance when compared with coexistent standalone 5G and 4G technologies. Reinforcement learning (RL) has shown its huge potential in wireless scenarios where parameter learning is required given the dynamic nature of such context. In this paper, we propose two reinforcement learning algorithms: a single agent RL algorithm named Clipped Double Q-Learning (CDQL) and a hierarchical Deep Q-Learning (HiDQL) to improve Multiple Radio Access Technology (multi-RAT) dual-connectivity handover. We compare our proposal with two baselines: a fixed parameter and a dynamic parameter solution. Simulation results reveal significant improvements in terms of latency with a gain of 47.6% and 26.1% for Digital-Analog beamforming (BF), 17.1% and 21.6% for Hybrid-Analog BF, and 24.7% and 39% for Analog-Analog BF when comparing the RL-schemes HiDQL and CDQL with the with the existent solutions, HiDQL presented a slower convergence time, however obtained a more optimal solution than CDQL. Additionally, we foresee the advantages of utilizing context-information as geo-location of the UEs to reduce the beam exploration sector, and thus improving further multi-RAT handover latency results. Pedro Enrique Iturria-Rivera, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Steve Furr, Melike Erol-Kantarci |
GLOBECOM | 1 |
| 2022 | Deep Reinforcement Learning-Based Joint User Association and CU-DU Placement in O-RANabstractOpen Radio Access Networks (O-RAN) architecture is based on disaggregation, virtualization, openness, and intelligence. These features allow the RAN network functions (NFs) to be split into Central Unit (CU), Distributed Unit (DU), and Radio Unit (RU); and deployed on open hardware and cloud nodes as Virtualized Network Functions (VNFs) or Containerized Network Functions (CNFs). In this paper, we propose strategies for the placement of CU and DU network functions in the regional and edge O-Cloud nodes while jointly associating the users to RUs. The aim is to minimize the end-to-end delay of users and minimize the cost of O-RAN deployment. Thus, we first formulate the end-to-end delay, the cost, and the constraints. We then model the problem as a multi-objective optimization problem The optimization formulation consists of a huge number of constraints and variables. To provide a solution to the problem, we develop the corresponding Markov Decision Problem (MDP) and propose a Deep Q-Network (DQN)-based algorithm. The simulation results demonstrate that our proposed scheme reduces the average user delay up to 40% and the deployment cost up to 20% with respect to our baselines. Roghayeh Joda, Turgay Pamuklu, Pedro Enrique Iturria-Rivera, Melike Erol-Kantarci |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2021 | QoS-Aware Load Balancing in Wireless Networks using Clipped Double Q-LearningabstractIn recent years, long-term evolution (LTE) and 5G NR (5thGeneration New Radio) technologies have showed great potential to utilize Machine Learning (ML) algorithms in optimizing their operations, both thanks to the availability of fine-grained data from the field, as well as the need arising from growing complexity of networks. The aforementioned complexity sparked mobile operators’ attention as a way to reduce the capital expenditures (CAPEX) and the operational (OPEX) expenditures of their networks through network management automation (NMA). NMA falls under the umbrella of Self-Organizing Networks (SON) in which 3GPP has identified some challenges and opportunities in load balancing mechanisms for the Radio Access Networks (RANs). In the context of machine learning and load balancing, several studies have focused on maximizing the overall network throughput or the resource block utilization (RBU). In this paper, we propose a novel Clipped Double Q-Learning (CDQL)-based load balancing approach considering resource block utilization, latency and the Channel Quality Indicator (CQI). We compare our proposal with a traditional handover algorithm and a resource block utilization based handover mechanism. Simulation results reveal that our scheme is able to improve throughput, latency, jitter and packet loss ratio in comparison to the baseline algorithms. Pedro Enrique Iturria-Rivera, Melike Erol-Kantarci |
MASS | 1 |
| 2020 | Efficient strategy to optimize key devices positions in large-scale RF mesh networks
Ahmad Mohamad Mezher, Pedro Enrique Iturria-Rivera, Julián L. Cárdenas-Barrera, Julian Meng, Eduardo Castillo Guerra |
Ad Hoc Networks | 2 |