VLDB 2026 Research / reviewers in the wild / expert
Majid Bavand
dblp:176/8136
· DBLP profile ↗
16ranked-venue papers
0as first author
16since 2021 · last 2025
0000-0002-8331-0033ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 14 · 14 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Prioritized Value-Decomposition Network for Explainable AI-Enabled Network Slicing
Shavbo Salehi, Pedro Enrique Iturria-Rivera, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Yigit Ozcan, Melike Erol-Kantarci |
ICC | 4 |
| 2025 | LLM-Based Intent Processing and Network Optimization Using Attention-Based Hierarchical Reinforcement LearningabstractIntent-based network automation is a promising tool that enables easier network management; however, certain challenges must be addressed effectively. These are: 1) processing intents, i.e., identification of logic and necessary parameters to fulfill an intent, 2) validating an intent to align it with current network status, and 3) satisfying intents via network optimizing applications. This paper addresses these points via a three-fold strategy to introduce intent-based automation for modern 5G architectures. First, intents are processed via a lightweight Large Language Model (LLM). Secondly, once an intent is processed, it is validated against future incoming traffic volume profiles (high or low). Finally, a series of network optimization applications has been developed. With their machine learning-based functionalities, they can improve certain key performance indicators such as throughput, delay, and energy efficiency. In the final stage, using an attention-based hierarchical reinforcement learning algorithm, these applications are optimally initiated to satisfy the intent of an operator. Our simulations show that the proposed method can achieve at least a 12% increase in throughput, a 17.1% increase in energy efficiency, and a 26.5% decrease in network delay compared to the baseline algorithms. Md Arafat Habib, Pedro Enrique Iturria-Rivera, Yigit Ozcan, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Melike Erol-Kantarci |
WCNC | 5 |
| 2025 | Intelligent Attacks and Defense Methods in Federated Learning-Enabled Energy-Efficient Wireless NetworksabstractFederated learning (FL) is a promising technique for learning-based functions in wireless networks, thanks to its distributed implementation capability. On the other hand, distributed learning may increase the risk of exposure to malicious attacks where attacks on a local model may spread to other models by parameter exchange. Meanwhile, such attacks can be hard to detect due to the dynamic wireless environment, especially considering local models can be heterogeneous with non-independent and identically distributed (non-IID) data. Therefore, it is critical to evaluate the effect of malicious attacks and develop advanced defense techniques for FL-enabled wireless networks. In this work, we introduce a federated deep reinforcement learning-based cell sleep control scenario that enhances the energy efficiency of the network. We propose multiple intelligent attacks targeting the learning-based approach and we propose defense methods to mitigate such attacks. In particular, we have designed two attack models, generative adversarial network (GAN)-enhanced model poisoning attack and regularization-based model poisoning attack. As a counteraction, we have proposed two defense schemes, autoencoder-based defense, and knowledge distillation (KD)-enabled defense. The autoencoder-based defense method leverages an autoencoder to identify the malicious participants and only aggregate the parameters of benign local models during the global aggregation, while KD-based defense protects the model from attacks by controlling the knowledge transferred between the global model and local models. The simulation results demonstrate that the proposed attacks can degrade the network performance by 34% and 77%, and lead to lower throughput and energy efficiency. On the other hand, our proposed defense schemes can effectively protect the system from attacks. The system performance can be recovered to approximately 95% of a secure system by using the proposed KD-based defense. Han Zhang 0055, Hao Zhou 0013, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Yigit Ozcan, Melike Erol-Kantarci |
IEEE Trans. Wirel. Commun. | 4 |
| 2024 | Self-Play Ensemble Q-learning enabled Resource Allocation for Network SlicingabstractIn 5G networks, network slicing has emerged as a pivotal paradigm to address diverse user demands and service requirements. To meet the requirements, reinforcement learning (RL) algorithms have been utilized widely, but this method has the problem of overestimation and exploration-exploitation trade-offs. To tackle these problems, this paper explores the application of self-play ensemble Q-learning, an extended version of the RL-based technique. Self-play ensemble Q-learning utilizes multiple Q-tables with various exploration-exploitation rates leading to different observations for choosing the most suitable action for each state. Moreover, through self-play, each model endeavors to enhance its performance compared to its previous iterations, boosting system efficiency, and decreasing the effect of overestimation. For performance evaluation, we consider three RL-based algorithms; self-play ensemble Q-learning, double Q-learning, and Q-learning, and compare their performance under different network traffic. Through simulations, we demonstrate the effectiveness of self-play ensemble Q-learning in meeting the diverse demands within 21.92% in latency, 24.22% in throughput, and 23.63% in packet drop rate in comparison with the baseline methods. Furthermore, we evaluate the robustness of self-play ensemble Q-learning and double Q-learning in situations where one of the Q-tables is affected by a malicious user. Our results depicted that the self-play ensemble Q-learning method is more robust against adversarial users and prevents a noticeable drop in system performance, mitigating the impact of users manipulating policies. Shavbo Salehi, Pedro Enrique Iturria-Rivera, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Yigit Ozcan, Melike Erol-Kantarci |
GLOBECOM | 4 |
| 2024 | Jamming Attacks and Mitigation in Transfer Learning Enabled 5G RAN SlicingabstractRadio access technology is crucial in both 5G and 6G cellular networks, providing differentiated services that demand reliability, low latency, and high throughput. To meet these requirements, machine learning (ML) has demonstrated considerable progress by facilitating resource allocation. However, these ML techniques can be susceptible to attacks, and the jamming attack is one of the most considered attacks in the literature, disrupting network functionality by sending interference signals. This paper, to the best of our knowledge for the first time, examines the vulnerability of radio access networks (RANs) to jamming attacks on resource allocation of a transfer reinforcement learning (TRL) based system and provides a mitigation approach to such attacks. A system model is presented for RAN slicing, followed by an introduction of the TRL algorithm for resource allocation. Afterward, we investigate covert patterned jamming attack (CPJA) on the TRL algorithm in downlink communication which decreases system throughput by 17% and 38.14% in the expert and learner agents and increases latency by 7.36% and 9.37% respectively. In addition, we propose a neural network (NN) solution to mitigate the CPJA trained on the network side and provide the trained NN model to the users' equipment (UEs) to eliminate interference from the signal by the filter. The trained NN is applied to predict the future activity of the interference generated by the attacker. Attack mitigation reduces the impact of the attack while the system's throughput suffers a 6% and 1.8% degradation, and its latency increases by 6.5% and 3.83% compared to the original system for expert and learner agents, respectively. Shavbo Salehi, Hao Zhou 0013, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Yigit Ozcan, Melike Erol-Kantarci |
ICC | 4 |
| 2024 | Federated Learning with Dual Attention for Robust Modulation Classification under AttacksabstractFederated learning (FL) allows distributed partic-ipants to train machine learning models in a decentralized manner. It can be used for radio signal classification with multiple receivers due to its benefits in terms of privacy and scalability. However, the existing FL algorithms usually suffer from slow and unstable convergence and are vulnerable to poisoning attacks from malicious participants. In this work, we aim to design a versatile FL framework that simultaneously promotes the performance of the model both in a secure system and under attack. To this end, we leverage attention mechanisms as a defense against attacks in FL and propose a robust FL algorithm by integrating the attention mechanisms into the global model aggregation step. To be more specific, two attention models are combined to calculate the amount of attention cast on each participant. It will then be used to determine the weights of local models during the global aggregation. The proposed algorithm is verified on a real-world dataset and it outperforms existing algorithms, both in secure systems and in systems under data poisoning attacks. Han Zhang 0055, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Yigit Ozcan, Melike Erol-Kantarci |
ICC | 3 |
| 2024 | Extended Reality (XR) Codec Adaptation in 5G using Multi-Agent Reinforcement Learning with Attention Action SelectionabstractExtended Reality (XR) services will revolutionize applications over $5^{\text {th }}$ and $\mathbf{6}^{\text {th }}$ generation wireless networks by providing seamless virtual and augmented reality experiences. These applications impose significant challenges on network infrastructure, which can be addressed by machine learning algorithms due to their adaptability. This paper presents a Multi-Agent Reinforcement Learning (MARL) solution for optimizing codec parameters of XR traffic, comparing it to the Adjust Packet Size (APS) algorithm. Our cooperative multi-agent system uses an Optimistic Mixture of Q-Values ($\mathbf{O Q M I X}$) approach for handling Cloud Gaming (CG), Augmented Reality (AR), and Virtual Reality (VR) traffic. Enhancements include an attention mechanism and slate-Markov Decision Process (MDP) for improved action selection. Simulations show our solution outperforms APS with average gains of $30.1 \%, 15.6 \%, 16.5 \% 50.3 \%$ in XR index, jitter, delay, and Packet Loss Ratio (PLR), respectively. APS tends to increase throughput but also packet losses, whereas oQMIX reduces PLR, delay, and jitter while maintaining goodput. Pedro Enrique Iturria-Rivera, Raimundas Gaigalas, Medhat H. M. Elsayed, Majid Bavand, Yigit Ozcan, Melike Erol-Kantarci |
PIMRC | 4 |
| 2024 | Asynchronous Bidirectional Communication in Cell-Free NetworksabstractWe consider a bidirectional communication between two single-antenna transceivers using multiple multi-antenna access points (APs) in a cell-free network architecture. In such a network, because of different propagation delays associated with different APs, the end-to-end link is a multi-path channel that results in inter-symbol-interference (ISI) in the signals received at the transceivers. To tackle ISI, we resort to cyclic prefix (CP) assisted block transmission of the information symbols and employ joint pre- and post-channel equalizers at both the transceivers to mitigate the impact of intra-block interference. Considering the amplify-and-forward technique at the APs, we cast the joint design of equalizers, beamforming matrices, and transceivers’ transmit powers as a power minimization problem while guaranteeing predefined data rates at the transceivers. Assuming symmetric beamforming matrices at the APs, we devise a semi-closed-form solution for this problem. We prove rigorously that at the optimum only a synchronous subset of the APs should participate in the information exchange between the two transceivers. This is achieved by proving that at the optimum, the pre-equalizer matrices should be unitary and the post-equalizer matrices should be invertible. Roozbeh Mohammadian, Zahra Pourgharehkhan, Shahram Shahbazpanahi, Majid Bavand, Gary Boudreau |
IEEE Trans. Wirel. Commun. | 4 |
| 2023 | Traffic Steering for 5G Multi-RAT Deployments using Deep Reinforcement LearningabstractIn 5G non-standalone mode, traffic steering is a critical technique to take full advantage of 5G new radio while optimizing dual connectivity of 5G and LTE networks in multiple radio access technology (RAT). An intelligent traffic steering mechanism can play an important role to maintain seamless user experience by choosing appropriate RAT (5G or LTE) dynamically for a specific user traffic flow with certain QoS requirements. In this paper, we propose a novel traffic steering mechanism based on Deep Q-learning that can automate traffic steering decisions in a dynamic environment having multiple RATs, and maintain diverse QoS requirements for different traffic classes. The proposed method is compared with two baseline algorithms: a heuristic-based algorithm and Q-learning-based traffic steering. Compared to the Q-learning and heuristic baselines, our results show that the proposed algorithm achieves better performance in terms of 6% and 10% higher average system throughput, and 23% and 33% lower network delay, respectively. Md Arafat Habib, Hao Zhou 0013, Pedro Enrique Iturria-Rivera, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Steve Furr, Melike Erol-Kantarci |
CCNC | 5 |
| 2023 | Split Learning for Sensing-Aided Single and Multi-Level Beam Selection in Multi-Vendor RANabstractProper and efficient beam selection is of great importance to unleash the full potential of mmWave communications. Traditionally, each candidate beam is evaluated using reference signals (beam sweeping), however, the exhaustive search method can be time-consuming with high signaling overhead. To avoid such problems, in B5G and 6G, sensing information is considered to be used, as in Integrated Sensing and Communication (ISAC) solutions, and Machine Learning (ML) methods can be applied to map sensing data inputs to an optimal beam index. When using sensing information sources external to the Radio Access Network (RAN) in a multi-vendor disaggregated environment, those methods need to account for issues such as privacy and data ownership. In this work, we apply multi-modal sensing information to the beam selection task. Specifically, we propose a multi-modal sensing-aided ML strategy based on Split Learning (SL) that can cope with deployment challenges in novel RAN architectures. Moreover, the method is applied to single and multi-level beam selection decisions, where the latter considers the case of hierarchical codebook structures. With the proposed approach, accuracy levels above 90% can be achieved while overhead diminishes by 85% or more. SL achieves comparable performance with centralized learning-based strategies, with the added value of accounting for privacy and data ownership issues. We also show that sensing-aided ML-based beam selection decisions in multi-level codebooks are more effective when applied to their first level. Ycaro Dantas, Pedro Enrique Iturria-Rivera, Hao Zhou 0013, Yigit Ozcan, Majid Bavand, Medhat H. M. Elsayed, Raimundas Gaigalas, Melike Erol-Kantarci |
GLOBECOM | 5 |
| 2023 | Policy Poisoning Attacks on Transfer Learning Enabled Resource Allocation for Network SlicingabstractAs wireless networks continue to evolve, machine learning (ML) algorithms are used to address communication challenges and meet various service requirements. While ML methods are promising, they can be prone to malicious attacks, which may degrade user experience and network performance. Specifically, the security challenges of radio access networks (RANs) are highlighted due to frequent interactions with a large number of users, and evaluating these attacks is critical for securing wireless communications. In this paper, for the first time, we investigate the vulnerability of transfer reinforcement learning (TRL) algorithms for resource allocation in 5G RAN slicing. In particular, we first present the system model for RAN slicing, and then the TRL algorithm is introduced for resource allocation. Afterward, we investigate three types of attack methods on the TRL algorithm. The simulations indicate that the attack on an expert agent can affect the performance of a learner agent since the expert shares knowledge with the learner. We show that the effect of the black-box policy poisoning attack on the learner increases latency by 18.31% and reduces throughput by 13.80%, while white-box policy poisoning attacks result in a 48.96% increase in latency and an 87.02% reduction in throughput. Shavbo Salehi, Hao Zhou 0013, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Yigit Ozcan, Melike Erol-Kantarci |
GLOBECOM | 4 |
| 2023 | Beam Selection for Energy-Efficient mmWave Network Using Advantage Actor Critic LearningabstractThe growing adoption of mmWave frequency bands to realize the full potential of 5G, turns beamforming into a key enabler for current and next-generation wireless technologies. Many mmWave networks rely on beam selection with Grid-of-Beams (GoB) approach to handle user-beam association. In beam selection with GoB, users select the appropriate beam from a set of pre-defined beams and the overhead during the beam selection process is a common challenge in this area. In this paper, we propose an Advantage Actor Critic (A2C) learning-based framework to improve the GoB and the beam selection process, as well as optimize transmission power in a mmWave network. The proposed beam selection technique allows performance improvement while considering transmission power improves Energy Efficiency (EE) and ensures the coverage is maintained in the network. We further investigate how the proposed algorithm can be deployed in a Service Management and Orchestration (SMO) platform. Our simulations show that A2C-based joint optimization of beam selection and transmission power is more effective than using Equally Spaced Beams (ESB) and fixed power strategy, or optimization of beam selection and transmission power disjointly. Compared to the ESB and fixed transmission power strategy, the proposed approach achieves more than twice the average EE in the scenarios under test and is closer to the maximum theoretical EE. Ycaro Dantas, Pedro Enrique Iturria-Rivera, Hao Zhou 0013, Majid Bavand, Medhat H. M. Elsayed, Raimundas Gaigalas, Melike Erol-Kantarci |
ICC | 4 |
| 2023 | Hierarchical Reinforcement Learning Based Traffic Steering in Multi-RAT 5G DeploymentsabstractIn 5G non-standalone mode, an intelligent traffic steering mechanism can vastly aid in ensuring a smooth user experience by selecting the best radio access technology (RAT) from a multi-RAT environment for a specific traffic flow. In this paper, we propose a novel load-aware traffic steering algorithm based on hierarchical reinforcement learning (HRL) while satisfying the diverse quality of service requirements of different traffic types. HRL can significantly increase system performance using a bi-level architecture having a meta-controller and a controller. In our proposed method, the meta-controller provides an appropriate threshold for load balancing, while the controller performs traffic admission to an appropriate RAT in the lower level. Simulation results show that HRL outperforms a Deep Q-Learning (DQN) and a threshold-based heuristic baseline with 8.49%, 12.52% higher average system throughput and 27.74%, 39.13% lower network delay, respectively. Md Arafat Habib, Hao Zhou 0013, Pedro Enrique Iturria-Rivera, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Yigit Ozcan, Melike Erol-Kantarci |
ICC | 5 |
| 2022 | Hierarchical Deep Q-Learning Based Handover in Wireless Networks with Dual Connectivityabstract5G New Radio proposes the usage of frequencies above 10 GHz to speed up LTE's existent maximum data rates. However, the effective size of 5G antennas and consequently its repercussions in the signal degradation in urban scenarios makes it a challenge to maintain stable coverage and connectivity. In order to obtain the best from both technologies, recent dual connectivity solutions have proved their capabilities to improve performance when compared with coexistent standalone 5G and 4G technologies. Reinforcement learning (RL) has shown its huge potential in wireless scenarios where parameter learning is required given the dynamic nature of such context. In this paper, we propose two reinforcement learning algorithms: a single agent RL algorithm named Clipped Double Q-Learning (CDQL) and a hierarchical Deep Q-Learning (HiDQL) to improve Multiple Radio Access Technology (multi-RAT) dual-connectivity handover. We compare our proposal with two baselines: a fixed parameter and a dynamic parameter solution. Simulation results reveal significant improvements in terms of latency with a gain of 47.6% and 26.1% for Digital-Analog beamforming (BF), 17.1% and 21.6% for Hybrid-Analog BF, and 24.7% and 39% for Analog-Analog BF when comparing the RL-schemes HiDQL and CDQL with the with the existent solutions, HiDQL presented a slower convergence time, however obtained a more optimal solution than CDQL. Additionally, we foresee the advantages of utilizing context-information as geo-location of the UEs to reduce the beam exploration sector, and thus improving further multi-RAT handover latency results. Pedro Enrique Iturria-Rivera, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Steve Furr, Melike Erol-Kantarci |
GLOBECOM | 3 |
| 2022 | Hierarchical Reinforcement Learning for RIS-Assisted Energy-Efficient RANabstractReconfigurable intelligent surface (RIS) is emerging as a promising technology to boost the energy efficiency (EE) of 5G beyond and 6G networks. Inspired by this potential, in this paper, we investigate the RIS-assisted energy-efficient radio access networks (RAN). In particular, we combine RIS with sleep control techniques, and develop a hierarchical reinforcement learning (HRL) algorithm for network management. In HRL, the meta-controller decides the on/off status of the small base stations (SBSs) in heterogeneous networks, while the sub-controller can change the transmission power levels of SBSs to save energy. The simulations show that the RIS-assisted sleep control can achieve significantly lower energy consumption, higher throughput, and more than doubled energy efficiency than no- RIS conditions. Hao Zhou 0013, Long Kong, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Steve Furr, Melike Erol-Kantarci |
GLOBECOM | 4 |
| 2022 | Learning-Based User Clustering in NOMA-Aided MIMO Networks With Spatially Correlated ChannelsabstractThis paper considers the integration of non-orthogonal multiple access (NOMA) into massive multi-input multi-output (MIMO) systems for downlink transmission. We consider the joint design of user clustering, transmit beamforming, and power allocation to minimize the total transmit power while meeting the signal-to-interference-and-noise ratio targets. We decompose this challenging mixed-integer programming problem into three separate subproblems to solve. We propose a low-complexity learning-based user clustering algorithm, which is a modified version of mean shift clustering with a new channel correlation based clustering metric. The proposed clustering algorithm determines the clusters to trade-off between spatial dimension and power dimension offered by respective MIMO and NOMA for user multiplexing. We then design zero-forcing transmit beamformers to eliminate inter-cluster interference and optimize power allocation to minimize the total transmit power. We provide two case studies for both co-located and distributed massive MIMO systems in spatially highly correlated prorogation environments. Simulation results show that our proposed algorithm forms NOMA clusters based on the available degrees of freedom in the system to effectively use both spatial and power dimensions, which results in a substantial performance improvement over MIMO-only methods or other existing clustering methods in such environments. Sharareh KianiHarchehgani, Min Dong 0001, Shahram Shahbazpanahi, Gary Boudreau, Majid Bavand |
IEEE Trans. Commun. | 5 |