EDBT 2026 Demo / reviewers in the wild / expert
Medhat H. M. Elsayed
dblp:156/0070
· DBLP profile ↗
22ranked-venue papers
6as first author
18since 2021 · last 2025
0000-0002-1106-6078ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 19 · 6 first-author · 15 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Prioritized Value-Decomposition Network for Explainable AI-Enabled Network Slicing
Shavbo Salehi, Pedro Enrique Iturria-Rivera, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Yigit Ozcan, Melike Erol-Kantarci |
ICC | 3 |
| 2025 | LLM-Based Intent Processing and Network Optimization Using Attention-Based Hierarchical Reinforcement LearningabstractIntent-based network automation is a promising tool that enables easier network management; however, certain challenges must be addressed effectively. These are: 1) processing intents, i.e., identification of logic and necessary parameters to fulfill an intent, 2) validating an intent to align it with current network status, and 3) satisfying intents via network optimizing applications. This paper addresses these points via a three-fold strategy to introduce intent-based automation for modern 5G architectures. First, intents are processed via a lightweight Large Language Model (LLM). Secondly, once an intent is processed, it is validated against future incoming traffic volume profiles (high or low). Finally, a series of network optimization applications has been developed. With their machine learning-based functionalities, they can improve certain key performance indicators such as throughput, delay, and energy efficiency. In the final stage, using an attention-based hierarchical reinforcement learning algorithm, these applications are optimally initiated to satisfy the intent of an operator. Our simulations show that the proposed method can achieve at least a 12% increase in throughput, a 17.1% increase in energy efficiency, and a 26.5% decrease in network delay compared to the baseline algorithms. Md Arafat Habib, Pedro Enrique Iturria-Rivera, Yigit Ozcan, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Melike Erol-Kantarci |
WCNC | 4 |
| 2025 | Intelligent Attacks and Defense Methods in Federated Learning-Enabled Energy-Efficient Wireless NetworksabstractFederated learning (FL) is a promising technique for learning-based functions in wireless networks, thanks to its distributed implementation capability. On the other hand, distributed learning may increase the risk of exposure to malicious attacks where attacks on a local model may spread to other models by parameter exchange. Meanwhile, such attacks can be hard to detect due to the dynamic wireless environment, especially considering local models can be heterogeneous with non-independent and identically distributed (non-IID) data. Therefore, it is critical to evaluate the effect of malicious attacks and develop advanced defense techniques for FL-enabled wireless networks. In this work, we introduce a federated deep reinforcement learning-based cell sleep control scenario that enhances the energy efficiency of the network. We propose multiple intelligent attacks targeting the learning-based approach and we propose defense methods to mitigate such attacks. In particular, we have designed two attack models, generative adversarial network (GAN)-enhanced model poisoning attack and regularization-based model poisoning attack. As a counteraction, we have proposed two defense schemes, autoencoder-based defense, and knowledge distillation (KD)-enabled defense. The autoencoder-based defense method leverages an autoencoder to identify the malicious participants and only aggregate the parameters of benign local models during the global aggregation, while KD-based defense protects the model from attacks by controlling the knowledge transferred between the global model and local models. The simulation results demonstrate that the proposed attacks can degrade the network performance by 34% and 77%, and lead to lower throughput and energy efficiency. On the other hand, our proposed defense schemes can effectively protect the system from attacks. The system performance can be recovered to approximately 95% of a secure system by using the proposed KD-based defense. Han Zhang 0055, Hao Zhou 0013, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Yigit Ozcan, Melike Erol-Kantarci |
IEEE Trans. Wirel. Commun. | 3 |
| 2024 | Self-Play Ensemble Q-learning enabled Resource Allocation for Network SlicingabstractIn 5G networks, network slicing has emerged as a pivotal paradigm to address diverse user demands and service requirements. To meet the requirements, reinforcement learning (RL) algorithms have been utilized widely, but this method has the problem of overestimation and exploration-exploitation trade-offs. To tackle these problems, this paper explores the application of self-play ensemble Q-learning, an extended version of the RL-based technique. Self-play ensemble Q-learning utilizes multiple Q-tables with various exploration-exploitation rates leading to different observations for choosing the most suitable action for each state. Moreover, through self-play, each model endeavors to enhance its performance compared to its previous iterations, boosting system efficiency, and decreasing the effect of overestimation. For performance evaluation, we consider three RL-based algorithms; self-play ensemble Q-learning, double Q-learning, and Q-learning, and compare their performance under different network traffic. Through simulations, we demonstrate the effectiveness of self-play ensemble Q-learning in meeting the diverse demands within 21.92% in latency, 24.22% in throughput, and 23.63% in packet drop rate in comparison with the baseline methods. Furthermore, we evaluate the robustness of self-play ensemble Q-learning and double Q-learning in situations where one of the Q-tables is affected by a malicious user. Our results depicted that the self-play ensemble Q-learning method is more robust against adversarial users and prevents a noticeable drop in system performance, mitigating the impact of users manipulating policies. Shavbo Salehi, Pedro Enrique Iturria-Rivera, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Yigit Ozcan, Melike Erol-Kantarci |
GLOBECOM | 3 |
| 2024 | Jamming Attacks and Mitigation in Transfer Learning Enabled 5G RAN SlicingabstractRadio access technology is crucial in both 5G and 6G cellular networks, providing differentiated services that demand reliability, low latency, and high throughput. To meet these requirements, machine learning (ML) has demonstrated considerable progress by facilitating resource allocation. However, these ML techniques can be susceptible to attacks, and the jamming attack is one of the most considered attacks in the literature, disrupting network functionality by sending interference signals. This paper, to the best of our knowledge for the first time, examines the vulnerability of radio access networks (RANs) to jamming attacks on resource allocation of a transfer reinforcement learning (TRL) based system and provides a mitigation approach to such attacks. A system model is presented for RAN slicing, followed by an introduction of the TRL algorithm for resource allocation. Afterward, we investigate covert patterned jamming attack (CPJA) on the TRL algorithm in downlink communication which decreases system throughput by 17% and 38.14% in the expert and learner agents and increases latency by 7.36% and 9.37% respectively. In addition, we propose a neural network (NN) solution to mitigate the CPJA trained on the network side and provide the trained NN model to the users' equipment (UEs) to eliminate interference from the signal by the filter. The trained NN is applied to predict the future activity of the interference generated by the attacker. Attack mitigation reduces the impact of the attack while the system's throughput suffers a 6% and 1.8% degradation, and its latency increases by 6.5% and 3.83% compared to the original system for expert and learner agents, respectively. Shavbo Salehi, Hao Zhou 0013, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Yigit Ozcan, Melike Erol-Kantarci |
ICC | 3 |
| 2024 | Federated Learning with Dual Attention for Robust Modulation Classification under AttacksabstractFederated learning (FL) allows distributed partic-ipants to train machine learning models in a decentralized manner. It can be used for radio signal classification with multiple receivers due to its benefits in terms of privacy and scalability. However, the existing FL algorithms usually suffer from slow and unstable convergence and are vulnerable to poisoning attacks from malicious participants. In this work, we aim to design a versatile FL framework that simultaneously promotes the performance of the model both in a secure system and under attack. To this end, we leverage attention mechanisms as a defense against attacks in FL and propose a robust FL algorithm by integrating the attention mechanisms into the global model aggregation step. To be more specific, two attention models are combined to calculate the amount of attention cast on each participant. It will then be used to determine the weights of local models during the global aggregation. The proposed algorithm is verified on a real-world dataset and it outperforms existing algorithms, both in secure systems and in systems under data poisoning attacks. Han Zhang 0055, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Yigit Ozcan, Melike Erol-Kantarci |
ICC | 2 |
| 2024 | Extended Reality (XR) Codec Adaptation in 5G using Multi-Agent Reinforcement Learning with Attention Action SelectionabstractExtended Reality (XR) services will revolutionize applications over $5^{\text {th }}$ and $\mathbf{6}^{\text {th }}$ generation wireless networks by providing seamless virtual and augmented reality experiences. These applications impose significant challenges on network infrastructure, which can be addressed by machine learning algorithms due to their adaptability. This paper presents a Multi-Agent Reinforcement Learning (MARL) solution for optimizing codec parameters of XR traffic, comparing it to the Adjust Packet Size (APS) algorithm. Our cooperative multi-agent system uses an Optimistic Mixture of Q-Values ($\mathbf{O Q M I X}$) approach for handling Cloud Gaming (CG), Augmented Reality (AR), and Virtual Reality (VR) traffic. Enhancements include an attention mechanism and slate-Markov Decision Process (MDP) for improved action selection. Simulations show our solution outperforms APS with average gains of $30.1 \%, 15.6 \%, 16.5 \% 50.3 \%$ in XR index, jitter, delay, and Packet Loss Ratio (PLR), respectively. APS tends to increase throughput but also packet losses, whereas oQMIX reduces PLR, delay, and jitter while maintaining goodput. Pedro Enrique Iturria-Rivera, Raimundas Gaigalas, Medhat H. M. Elsayed, Majid Bavand, Yigit Ozcan, Melike Erol-Kantarci |
PIMRC | 3 |
| 2023 | Traffic Steering for 5G Multi-RAT Deployments using Deep Reinforcement LearningabstractIn 5G non-standalone mode, traffic steering is a critical technique to take full advantage of 5G new radio while optimizing dual connectivity of 5G and LTE networks in multiple radio access technology (RAT). An intelligent traffic steering mechanism can play an important role to maintain seamless user experience by choosing appropriate RAT (5G or LTE) dynamically for a specific user traffic flow with certain QoS requirements. In this paper, we propose a novel traffic steering mechanism based on Deep Q-learning that can automate traffic steering decisions in a dynamic environment having multiple RATs, and maintain diverse QoS requirements for different traffic classes. The proposed method is compared with two baseline algorithms: a heuristic-based algorithm and Q-learning-based traffic steering. Compared to the Q-learning and heuristic baselines, our results show that the proposed algorithm achieves better performance in terms of 6% and 10% higher average system throughput, and 23% and 33% lower network delay, respectively. Md Arafat Habib, Hao Zhou 0013, Pedro Enrique Iturria-Rivera, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Steve Furr, Melike Erol-Kantarci |
CCNC | 4 |
| 2023 | Split Learning for Sensing-Aided Single and Multi-Level Beam Selection in Multi-Vendor RANabstractProper and efficient beam selection is of great importance to unleash the full potential of mmWave communications. Traditionally, each candidate beam is evaluated using reference signals (beam sweeping), however, the exhaustive search method can be time-consuming with high signaling overhead. To avoid such problems, in B5G and 6G, sensing information is considered to be used, as in Integrated Sensing and Communication (ISAC) solutions, and Machine Learning (ML) methods can be applied to map sensing data inputs to an optimal beam index. When using sensing information sources external to the Radio Access Network (RAN) in a multi-vendor disaggregated environment, those methods need to account for issues such as privacy and data ownership. In this work, we apply multi-modal sensing information to the beam selection task. Specifically, we propose a multi-modal sensing-aided ML strategy based on Split Learning (SL) that can cope with deployment challenges in novel RAN architectures. Moreover, the method is applied to single and multi-level beam selection decisions, where the latter considers the case of hierarchical codebook structures. With the proposed approach, accuracy levels above 90% can be achieved while overhead diminishes by 85% or more. SL achieves comparable performance with centralized learning-based strategies, with the added value of accounting for privacy and data ownership issues. We also show that sensing-aided ML-based beam selection decisions in multi-level codebooks are more effective when applied to their first level. Ycaro Dantas, Pedro Enrique Iturria-Rivera, Hao Zhou 0013, Yigit Ozcan, Majid Bavand, Medhat H. M. Elsayed, Raimundas Gaigalas, Melike Erol-Kantarci |
GLOBECOM | 6 |
| 2023 | Policy Poisoning Attacks on Transfer Learning Enabled Resource Allocation for Network SlicingabstractAs wireless networks continue to evolve, machine learning (ML) algorithms are used to address communication challenges and meet various service requirements. While ML methods are promising, they can be prone to malicious attacks, which may degrade user experience and network performance. Specifically, the security challenges of radio access networks (RANs) are highlighted due to frequent interactions with a large number of users, and evaluating these attacks is critical for securing wireless communications. In this paper, for the first time, we investigate the vulnerability of transfer reinforcement learning (TRL) algorithms for resource allocation in 5G RAN slicing. In particular, we first present the system model for RAN slicing, and then the TRL algorithm is introduced for resource allocation. Afterward, we investigate three types of attack methods on the TRL algorithm. The simulations indicate that the attack on an expert agent can affect the performance of a learner agent since the expert shares knowledge with the learner. We show that the effect of the black-box policy poisoning attack on the learner increases latency by 18.31% and reduces throughput by 13.80%, while white-box policy poisoning attacks result in a 48.96% increase in latency and an 87.02% reduction in throughput. Shavbo Salehi, Hao Zhou 0013, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Yigit Ozcan, Melike Erol-Kantarci |
GLOBECOM | 3 |
| 2023 | Beam Selection for Energy-Efficient mmWave Network Using Advantage Actor Critic LearningabstractThe growing adoption of mmWave frequency bands to realize the full potential of 5G, turns beamforming into a key enabler for current and next-generation wireless technologies. Many mmWave networks rely on beam selection with Grid-of-Beams (GoB) approach to handle user-beam association. In beam selection with GoB, users select the appropriate beam from a set of pre-defined beams and the overhead during the beam selection process is a common challenge in this area. In this paper, we propose an Advantage Actor Critic (A2C) learning-based framework to improve the GoB and the beam selection process, as well as optimize transmission power in a mmWave network. The proposed beam selection technique allows performance improvement while considering transmission power improves Energy Efficiency (EE) and ensures the coverage is maintained in the network. We further investigate how the proposed algorithm can be deployed in a Service Management and Orchestration (SMO) platform. Our simulations show that A2C-based joint optimization of beam selection and transmission power is more effective than using Equally Spaced Beams (ESB) and fixed power strategy, or optimization of beam selection and transmission power disjointly. Compared to the ESB and fixed transmission power strategy, the proposed approach achieves more than twice the average EE in the scenarios under test and is closer to the maximum theoretical EE. Ycaro Dantas, Pedro Enrique Iturria-Rivera, Hao Zhou 0013, Majid Bavand, Medhat H. M. Elsayed, Raimundas Gaigalas, Melike Erol-Kantarci |
ICC | 5 |
| 2023 | Hierarchical Reinforcement Learning Based Traffic Steering in Multi-RAT 5G DeploymentsabstractIn 5G non-standalone mode, an intelligent traffic steering mechanism can vastly aid in ensuring a smooth user experience by selecting the best radio access technology (RAT) from a multi-RAT environment for a specific traffic flow. In this paper, we propose a novel load-aware traffic steering algorithm based on hierarchical reinforcement learning (HRL) while satisfying the diverse quality of service requirements of different traffic types. HRL can significantly increase system performance using a bi-level architecture having a meta-controller and a controller. In our proposed method, the meta-controller provides an appropriate threshold for load balancing, while the controller performs traffic admission to an appropriate RAT in the lower level. Simulation results show that HRL outperforms a Deep Q-Learning (DQN) and a threshold-based heuristic baseline with 8.49%, 12.52% higher average system throughput and 27.74%, 39.13% lower network delay, respectively. Md Arafat Habib, Hao Zhou 0013, Pedro Enrique Iturria-Rivera, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Yigit Ozcan, Melike Erol-Kantarci |
ICC | 4 |
| 2022 | Hierarchical Deep Q-Learning Based Handover in Wireless Networks with Dual Connectivityabstract5G New Radio proposes the usage of frequencies above 10 GHz to speed up LTE's existent maximum data rates. However, the effective size of 5G antennas and consequently its repercussions in the signal degradation in urban scenarios makes it a challenge to maintain stable coverage and connectivity. In order to obtain the best from both technologies, recent dual connectivity solutions have proved their capabilities to improve performance when compared with coexistent standalone 5G and 4G technologies. Reinforcement learning (RL) has shown its huge potential in wireless scenarios where parameter learning is required given the dynamic nature of such context. In this paper, we propose two reinforcement learning algorithms: a single agent RL algorithm named Clipped Double Q-Learning (CDQL) and a hierarchical Deep Q-Learning (HiDQL) to improve Multiple Radio Access Technology (multi-RAT) dual-connectivity handover. We compare our proposal with two baselines: a fixed parameter and a dynamic parameter solution. Simulation results reveal significant improvements in terms of latency with a gain of 47.6% and 26.1% for Digital-Analog beamforming (BF), 17.1% and 21.6% for Hybrid-Analog BF, and 24.7% and 39% for Analog-Analog BF when comparing the RL-schemes HiDQL and CDQL with the with the existent solutions, HiDQL presented a slower convergence time, however obtained a more optimal solution than CDQL. Additionally, we foresee the advantages of utilizing context-information as geo-location of the UEs to reduce the beam exploration sector, and thus improving further multi-RAT handover latency results. Pedro Enrique Iturria-Rivera, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Steve Furr, Melike Erol-Kantarci |
GLOBECOM | 2 |
| 2022 | Hierarchical Reinforcement Learning for RIS-Assisted Energy-Efficient RANabstractReconfigurable intelligent surface (RIS) is emerging as a promising technology to boost the energy efficiency (EE) of 5G beyond and 6G networks. Inspired by this potential, in this paper, we investigate the RIS-assisted energy-efficient radio access networks (RAN). In particular, we combine RIS with sleep control techniques, and develop a hierarchical reinforcement learning (HRL) algorithm for network management. In HRL, the meta-controller decides the on/off status of the small base stations (SBSs) in heterogeneous networks, while the sub-controller can change the transmission power levels of SBSs to save energy. The simulations show that the RIS-assisted sleep control can achieve significantly lower energy consumption, higher throughput, and more than doubled energy efficiency than no- RIS conditions. Hao Zhou 0013, Long Kong, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Steve Furr, Melike Erol-Kantarci |
GLOBECOM | 3 |
| 2021 | Reinforcement Learning Based Energy-Efficient Component Carrier Activation-Deactivation in 5GabstractCarrier aggregation (CA) is considered a key enabler technology for delivering higher rates to users of LTE and 5G networks. However, the increased transmission rate comes with the price of higher energy consumption which stems from users continuously monitoring the control channel of the active component carriers (CCs) whether data transmission is ongoing or not. In order to reduce energy consumption, we exploit the activation-deactivation procedure at the medium access control (MAC) layer of LTE/5G network. In this paper, we propose a reinforcement learning-based algorithm to improve energy-efficiency by dynamically activating-deactivating secondary component carriers (SCCs) with awareness of the user traffic profiles. The proposed algorithm aims to predict the arrival of data and identify SCCs to activate for each user. In addition, a traffic splitting approach and an intelligent exploration strategy are proposed to balance users' load among CCs and improve the convergence of the algorithm, respectively. Results of the proposed algorithm are compared with three baseline algorithms. The first baseline always activates all CCs for each user, the second baseline activates one carrier only (i.e., the primary carrier) and the third baseline algorithm relies on a reactive method, where the activation-deactivation decision is performed after observing the arrival of data. Results show that Q-learning outperforms the baseline algorithms by achieving the highest sum throughput (and lowest average delay) with the lowest number of activated SCCs, which is obtained by learning to dynamically activate SCCs according to the traffic pattern. Hence, Q-learning is considered the most energy-efficient compared to the baseline algorithms. Medhat H. M. Elsayed, Roghayeh Joda, Hatem Abou-Zeid, Ramy Atawia, Akram Bin Sediq, Gary Boudreau, Melike Erol-Kantarci |
GLOBECOM | 1 |
| 2021 | QoS-Aware Joint Component Carrier Selection and Resource Allocation for Carrier Aggregation in 5GabstractCarrier Aggregation (CA) has been a breakthrough in LTE that led to increased throughput for users, and is still one of the key technologies in 5G that helps to enhance spectrum utilization. In CA, Component Carriers (CCs) are dynamically activated and deactivated depending on several performance factors. Optimal selection of CCs has been studied in the literature. However, the latency associated with activation and deactivation of CCs, control channel overhead for switching CCs, as well as the energy consumed for monitoring the active CCs have not been a part of the optimal CC selection problem. Nevertheless, those become stringent design constraints in practice. In this paper, we address optimal CC selection and resource allocation in 5G networks, where the above constraints are considered and the 5G network supports several service types with different 5G QoS Identifiers (5QI). The proposed optimum joint CC selection and Radio Resource Block (RB) allocation schemes maximize average throughput of users and satisfy QoS of users in terms of delay. In addition, the proposed schemes take CC activation and deactivation burden into consideration and aim to minimize the number of activations and deactivations. The simulation results demonstrate that our proposed solution outperforms the state of the art solution while satisfying the QoS requirements and creating close to 95.5% reduction on the number of CCs activations and deactivations. Roghayeh Joda, Medhat H. M. Elsayed, Hatem Abou-Zeid, Ramy Atawia, Akram Bin Sediq, Gary Boudreau, Melike Erol-Kantarci |
ICC | 2 |
| 2021 | RAN Resource Slicing in 5G Using Multi-Agent Correlated Q-Learningabstract5G is regarded as a revolutionary mobile network, which is expected to satisfy a vast number of novel services, ranging from remote health care to smart cities. However, heterogeneous Quality of Service (QoS) requirements of different services and limited spectrum make the radio resource allocation a challenging problem in 5G. In this paper, we propose a multi-agent reinforcement learning (MARL) method for radio resource slicing in 5G. We model each slice as an intelligent agent that competes for limited radio resources, and the correlated Q-learning is applied for inter-slice resource block (RB) allocation. The proposed correlated Q-learning based inter-slice RB allocation (COQRA) scheme is compared with Nash Q-learning (NQL), Latency-Reliability-Throughput Q-learning (LRTQ) methods, and the priority proportional fairness (PPF) algorithm. Our simulation results show that the proposed CO-QRA achieves 32.4% lower latency and 6.3% higher throughput when compared with LRTQ, and 5.8% lower latency and 5.9% higher throughput than NQL. Significantly higher throughput and lower packet drop rate (PDR) is observed in comparison to PPF. Hao Zhou 0013, Medhat H. M. Elsayed, Melike Erol-Kantarci |
PIMRC | 2 |
| 2021 | Transfer Reinforcement Learning for 5G New Radio mmWave NetworksabstractIn this paper, we aim at interference mitigation in 5G millimeter-Wave (mm-Wave) communications by employing beamforming and Non-Orthogonal Multiple Access (NOMA) techniques with the aim of improving network's aggregate rate. Despite the potential capacity gains of mm-Wave and NOMA, many technical challenges might hinder that performance gain. In particular, the performance of Successive Interference Cancellation (SIC) diminishes rapidly as the number of users increases per beam, which leads to higher intra-beam interference. Furthermore, intersection regions between adjacent cells give rise to inter-beam inter-cell interference. To mitigate both interference levels, optimal selection of the number of beams in addition to best allocation of users to those beams is essential. In this paper, we address the problem of joint user-cell association and selection of number of beams for the purpose of maximizing the aggregate network capacity. We propose three machine learning-based algorithms; transfer Q-learning (TQL), Q-learning, and Best SINR association with Density-based Spatial Clustering of Applications with Noise (BSDC) algorithms and compare their performance under different scenarios. Under mobility, TQL and Q-learning demonstrate 12% rate improvement over BSDC at the highest offered traffic load. For stationary scenarios, Q-learning and BSDC outperform TQL, however TQL achieves about 29% convergence speedup compared to Q-learning. Medhat H. M. Elsayed, Melike Erol-Kantarci, Halim Yanikomeroglu |
IEEE Trans. Wirel. Commun. | 1 |
| 2020 | Radio Resource and Beam Management in 5G mmWave Using Clustering and Deep Reinforcement LearningabstractTo optimally cover users in millimeter-Wave (mmWave) networks, clustering is needed to identify the number and direction of beams. The mobility of users motivates the need for an online clustering scheme to maintain up-to-date beams towards those clusters. Furthermore, mobility of users leads to varying patterns of clusters (i.e., users move from the coverage of one beam to another), causing dynamic traffic load per beam. As such, efficient radio resource allocation and beam management is needed to address the dynamicity that arises from mobility of users and their traffic. In this paper, we consider the coexistence of Ultra-Reliable Low-Latency Communication (URLLC) and enhanced Mobile BroadBand (eMBB) users in 5G mmWave networks and propose a Quality-of-Service (QoS) aware clustering and resource allocation scheme. Specifically, Density-Based Spatial Clustering of Applications with Noise (DBSCAN) is used for online clustering of users and the selection of the number of beams. In addition, Long Short Term Memory (LSTM)-based Deep Reinforcement Learning (DRL) scheme is used for resource block allocation. The performance of the proposed scheme is compared to a baseline that uses K-means and priority-based proportional fairness for clustering and resource allocation, respectively. Our simulation results show that the proposed scheme outperforms the baseline algorithm in terms of latency, reliability, and rate of URLLC users as well as rate of eMBB users. Medhat H. M. Elsayed, Melike Erol-Kantarci |
GLOBECOM | 1 |
| 2020 | Machine Learning-based Inter-Beam Inter-Cell Interference Mitigation in mmWaveabstractIn this paper, we address inter-beam inter-cell interference mitigation in 5G networks that employ millimetre-wave (mmWave), beamforming and non-orthogonal multiple access (NOMA) techniques. Those techniques play a key role in improving network capacity and spectral efficiency by multiplexing users on both spatial and power domains. In addition, the coverage area of multiple beams from different cells can intersect, allowing more flexibility in user-cell association. However, the intersection of coverage areas also implies increased inter-beam inter-cell interference, i.e. interference among beams formed by nearby cells. Therefore, joint user-cell association and inter-beam power allocation stand as a promising solution to mitigate inter-beam, inter-cell interference. In this paper, we consider a 5G mmWave network and propose a reinforcement learning algorithm to perform joint user-cell association and inter-beam power allocation to maximize the sum rate of the network. The proposed algorithm is compared to a uniform power allocation that equally divides power among beams per cell. Simulation results present a performance enhancement of 13 - 30% in network's sum-rate corresponding to the lowest and highest traffic loads, respectively. Medhat H. M. Elsayed, Kevin Shimotakahara, Melike Erol-Kantarci |
ICC | 1 |
| 2019 | Reinforcement Learning-Based Joint Power and Resource Allocation for URLLC in 5GabstractNext-generation wireless networks are moving rapidly towards supporting heterogeneous services that bring along several challenges in radio resource allocation. In this paper, we address the problem of multiplexing Ultra- Reliable Low- Latency Communication (URLLC) users and enhanced Mobile Broadband (eMBB) users on a shared channel of 5G New Radio (NR).We propose a joint power and resource allocation algorithm based on Q-learning. The proposed algorithm is crafted carefully to improve reliability and latency of URLLC users without hindering throughput of eMBB users. In particular, the algorithm rewards the actions that mitigate inter-cell interference as well as improve transmission and scheduling delays. We compare our results with a priority-based proportional fairness algorithm with fixed power allocation that relies on giving URLLC users priority in resource scheduling. Simulation results reveal that our algorithm is able to achieve 4% increase in reliability as well as lower latency results in high traffic load scenarios. Medhat H. M. Elsayed, Melike Erol-Kantarci |
GLOBECOM | 1 |
| 2018 | Deep Reinforcement Learning for Reducing Latency in Mission Critical ServicesabstractNext-generation wireless networks will be supporting mission critical services such as safety related applications of connected autonomous vehicles, and real-time control of medical and industrial systems, as well as serving traditional mobile users. In mission critical services, high-reliability and low-latency requirements should be satisfied. In this paper, we aim to reduce the latency of uplink scheduling of a network of Mission Critical Devices (MCDs) while maintaining fairness among other users, served by a dense small cell network. We propose a Deep Reinforcement Learning algorithm, namely Delay Minimizing Deep Q-Learning (DMDQ), that combines Long Short-term Memory with Q-learning. The problem is cast as a resource block allocation for delay minimization. The proposed algorithm is compared to a tabular Q-learning approach and a simple Round Robin (RR) algorithm in terms of latency, throughput, fairness and convergence. Our performance results show that DMDQ outperforms both schemes in terms of latency and offers high fairness. The Q-learning approach achieves slightly higher throughput than DMDQ however DMDQ convergences faster. Medhat H. M. Elsayed, Melike Erol-Kantarci |
GLOBECOM | 1 |