VLDB 2026 Research / reviewers in the wild / expert
Hao Zhou 0013
dblp:63/778-13
· DBLP profile ↗
25ranked-venue papers
4as first author
25since 2021 · last 2026
0000-0002-5511-4609ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 20 · 2 first-author · 20 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond DRL: LLM-enabled In-Context Learning for Aerial Data Collection in Public Safety UAV
Yousef Emami, Hao Zhou 0013, Miguel Gutiérrez-Gaitán, Kai Li 0002, Jin Zhao 0001, Luís Almeida 0001 |
IWCMC | 2 |
| 2026 | FRSICL: LLM-Enabled In-Context Learning Flight Resource Allocation for Fresh Data Collection in UAV-Assisted Wildfire MonitoringabstractUncrewed Aerial Vehicles (UAVs) play a vital role in public safety, especially in monitoring wildfires, where early detection reduces environmental impact. In UAV-Assisted Wildfire Monitoring (UAWM) systems, jointly optimizing the data collection schedule and UAV velocity is essential to minimize the average Age of Information (AoI) for sensory data. Deep Reinforcement Learning (DRL) has been used for this optimization, but its limitations – including low sampling efficiency, discrepancies between simulation and real-world conditions, and complex training – make it unsuitable for time-critical applications such as wildfire monitoring. Recent advances in Large Language Models (LLMs) provide a promising alternative. With strong reasoning and generalization capabilities, LLMs can adapt to new tasks through In-Context Learning (ICL), which enables task adaptation using natural language prompts and example-based guidance without retraining. This paper proposes a novel online Flight Resource Allocation scheme based on LLM-Enabled In-Context Learning (FRSICL) to jointly optimize the data collection schedule and UAV velocity along the trajectory in real time, thereby asymptotically minimizing the average AoI across all ground sensors. Unlike DRL, FRSICL generates data collection schedules and velocities using natural language task descriptions and feedback from the environment, enabling dynamic decision-making without extensive retraining. Simulation results confirm the effectiveness of FRSICL compared to state-of-the-art baselines, namely Proximal Policy Optimization, Block Coordinate Descent, and Nearest Neighbor. Yousef Emami, Hao Zhou 0013, Miguel Gutiérrez-Gaitán, Kai Li 0002, Luís Almeida 0001 |
IEEE Internet Things J. | 2 |
| 2025 | Understanding 6G through Language Models: A Case Study on LLM-aided Structured Entity Extraction in Telecom DomainabstractKnowledge understanding is a foundational part of envisioned 6G networks to advance network intelligence and AI-native network architectures. In this paradigm, information extraction plays a pivotal role in transforming fragmented telecom knowledge into well-structured formats, empowering diverse AI models to better understand network terminologies. This work proposes a novel language model-based information extraction technique, aiming to extract structured entities from the telecom context. The proposed telecom structured entity extraction (TeleSEE) technique applies a token-efficient representation method to predict entity types and attribute keys, aiming to save the number of output tokens and improve prediction accuracy. Meanwhile, TeleSEE involves a hierarchical parallel decoding method, improving the standard encoder-decoder architecture by integrating additional prompting and decoding strategies into entity extraction tasks. In addition, to better evaluate the performance of the proposed technique in the telecom domain, we further designed a dataset named 6GTech, including 2390 sentences and 23747 words from more than 100 6G-related technical publications. Finally, the experiment shows that the proposed TeleSEE method achieves higher accuracy than other baseline techniques, and also presents 5 to 9 times higher sample processing speed. Ye Yuan 0017, Haolun Wu, Hao Zhou 0013, Xue (Steve) Liu, Hao Chen 0010, Jianzhong Zhang 0002 |
GLOBECOM | 3 |
| 2025 | CVaR-Based Variational Quantum Optimization for User Association in Handoff-Aware Vehicular NetworksabstractEfficient resource allocation is essential for optimizing various tasks in wireless networks, which are usually formulated as generalized assignment problems (GAP). GAP, as a generalized version of the linear sum assignment problem, involves both equality and inequality constraints that add computational challenges. In this work, we present a novel Conditional Value at Risk (CVaR)-based Variational Quantum Eigensolver (VQE) framework to address GAP in vehicular networks (VNets). Our approach leverages a hybrid quantum-classical structure, integrating a tailored cost function that balances both objective and constraint-specific penalties to improve solution quality and stability. Using the CVaR-VQE model, we handle the GAP efficiently by focusing optimization on the lower tail of the solution space, enhancing both convergence and resilience on noisy intermediate-scale quantum (NISQ) devices. We apply this framework to a user-association problem in VNets, where our method achieves 23.5% improvement compared to the deep neural network (DNN) approach. Zijiang Yan, Hao Zhou 0013, Jianhua Pei, Aryan Kaushik, Hina Tabassum, Ping Wang 0001 |
ICC | 2 |
| 2025 | Generative AI for Immersive Video: Recent Advances and Future OpportunitiesabstractImmersive video serves as a key component of eXtended Reality (XR) that aims to create and interact with simulated virtual or hybrid environments. Such a technology allows users to experience immersive sensations that transcend time and space, and meanwhile continuously providing training data for emerging technologies like Embodied AI. Thanks to the advancements in sensing, computing, and display, recent years have witnessed many excellent works for XR and related hardware or software systems. However, challenges like high creation cost, lack of immersion, and limited scalability hinder the practical application of immersive video services. Whilst recently emerged generative artificial intelligence (GenAI) provides us with new insights in tackling existing challenges. In this paper, we conduct a comprehensive survey into the recent advances and future opportunities on how GenAI can benefit immersive video services. By introducing a systematic taxonomy, we meticulously classify the pertinent techniques and applications into three well-defined categories aligned with the pipeline of immersive video service: content creation, network delivery, and client-side display. This categorization enables a structured exploration of the diverse roles on how GenAI can benefit immersive video service, providing a framework for a more comprehensive understanding and evaluation of these technologies. To the best of our knowledge, this work is the first systematic survey of GenAI in XR settings, laying a foundation for future research in this interdisciplinary domain. Kaiyuan Hu, Yili Jin 0001, Hao Zhou 0013, Linfeng Du, Jiangchuan Liu |
IJCAI | 3 |
| 2025 | LLM-Enabled In-Context Learning for Data Collection Scheduling in UAV-Assisted Sensor NetworksabstractUnmanned Aerial Vehicles (UAVs) are increasingly being utilized in various private and commercial applications, e.g., traffic control, parcel delivery, and Search and Rescue (SAR) missions. Machine Learning (ML) methods used in UAV-Assisted Sensor Networks (UASNETs) and, especially, in Deep Reinforcement Learning (DRL) face challenges such as complex and lengthy model training, gaps between simulation and reality, and low sampling efficiency, which conflict with the urgency of emergencies, such as SAR missions. In this paper, an In-Context Learning (ICL)-Data Collection Scheduling (ICLDC) system is proposed as an alternative to DRL in emergencies. The UAV collects sensory data and transmits it to a Large Language Model (LLM), which creates a task description in natural language. From this description, the UAV receives a data collection schedule that must be executed. A verifier ensures safe UAV operations by evaluating the schedules generated by the LLM and overriding unsafe schedules based on predefined rules. The system continuously adapts by incorporating feedback into the task descriptions and using this for future decisions. This method is tested against jailbreaking attacks, where the task description is manipulated to undermine network performance, highlighting the vulnerability of LLMs to such attacks. The proposed ICLDC significantly reduces cumulative packet loss compared to both the DQN and Maximum Channel Gain baselines. ICLDC presents a promising direction for intelligent scheduling and control in UASNETs. Yousef Emami, Hao Zhou 0013, Seyedsina Nabavirazavi, Luís Almeida 0001 |
IEEE Internet Things J. | 2 |
| 2025 | A Hierarchical DRL Approach for Resource Optimization in Multi-RIS Multi-Operator NetworksabstractAs reconfigurable intelligent surfaces (RIS) emerge as a pivotal technology in the upcoming sixth-generation (6G) networks, its deployment within practical multiple operator (OP) networks presents significant challenges, including the coordination of RIS configurations among OPs, interference management, and privacy maintenance. A promising strategy is to treat RIS as a public resource managed by an RIS provider (RP), which can enhance resource allocation efficiency by allowing dynamic access for multiple OPs. However, the intricate nature of coordinating management and optimizing RIS configurations significantly complicates the implementation process. In this paper, we propose a hierarchical deep reinforcement learning (HDRL) approach that decomposes the complicated RIS resource optimization problem into several subtasks. Specifically, a top-level RP-agent is responsible for RIS allocation, while low-level OP-agents control their assigned RISs and handle beamforming, RIS phase-shifts, and user association. By utilizing the semi-Markov decision process (SMDP) theory, we establish a sophisticated interaction mechanism between the RP and OPs, and introduce an advanced hierarchical proximal policy optimization (HPPO) algorithm. Furthermore, we propose an improved sequential-HPPO (S-HPPO) algorithm to address the curse of dimensionality encountered with a single RP-agent. Experimental results validate the stability of the HPPO algorithm across various environmental parameters, demonstrating its superiority over other benchmarks for joint resource optimization. Finally, we conduct a detailed comparative analysis between the proposed S-HPPO and HPPO algorithms, showcasing that the S-HPPO algorithm achieves faster convergence and improved performance in large-scale RIS allocation scenarios. Haocheng Zhang, Wei Wang 0381, Hao Zhou 0013, Zhiping Lu, Ming Li 0011 |
IEEE Trans. Wirel. Commun. | 3 |
| 2025 | Intelligent Attacks and Defense Methods in Federated Learning-Enabled Energy-Efficient Wireless NetworksabstractFederated learning (FL) is a promising technique for learning-based functions in wireless networks, thanks to its distributed implementation capability. On the other hand, distributed learning may increase the risk of exposure to malicious attacks where attacks on a local model may spread to other models by parameter exchange. Meanwhile, such attacks can be hard to detect due to the dynamic wireless environment, especially considering local models can be heterogeneous with non-independent and identically distributed (non-IID) data. Therefore, it is critical to evaluate the effect of malicious attacks and develop advanced defense techniques for FL-enabled wireless networks. In this work, we introduce a federated deep reinforcement learning-based cell sleep control scenario that enhances the energy efficiency of the network. We propose multiple intelligent attacks targeting the learning-based approach and we propose defense methods to mitigate such attacks. In particular, we have designed two attack models, generative adversarial network (GAN)-enhanced model poisoning attack and regularization-based model poisoning attack. As a counteraction, we have proposed two defense schemes, autoencoder-based defense, and knowledge distillation (KD)-enabled defense. The autoencoder-based defense method leverages an autoencoder to identify the malicious participants and only aggregate the parameters of benign local models during the global aggregation, while KD-based defense protects the model from attacks by controlling the knowledge transferred between the global model and local models. The simulation results demonstrate that the proposed attacks can degrade the network performance by 34% and 77%, and lead to lower throughput and energy efficiency. On the other hand, our proposed defense schemes can effectively protect the system from attacks. The system performance can be recovered to approximately 95% of a secure system by using the proposed KD-based defense. Han Zhang 0055, Hao Zhou 0013, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Yigit Ozcan, Melike Erol-Kantarci |
IEEE Trans. Wirel. Commun. | 2 |
| 2024 | Jamming Attacks and Mitigation in Transfer Learning Enabled 5G RAN SlicingabstractRadio access technology is crucial in both 5G and 6G cellular networks, providing differentiated services that demand reliability, low latency, and high throughput. To meet these requirements, machine learning (ML) has demonstrated considerable progress by facilitating resource allocation. However, these ML techniques can be susceptible to attacks, and the jamming attack is one of the most considered attacks in the literature, disrupting network functionality by sending interference signals. This paper, to the best of our knowledge for the first time, examines the vulnerability of radio access networks (RANs) to jamming attacks on resource allocation of a transfer reinforcement learning (TRL) based system and provides a mitigation approach to such attacks. A system model is presented for RAN slicing, followed by an introduction of the TRL algorithm for resource allocation. Afterward, we investigate covert patterned jamming attack (CPJA) on the TRL algorithm in downlink communication which decreases system throughput by 17% and 38.14% in the expert and learner agents and increases latency by 7.36% and 9.37% respectively. In addition, we propose a neural network (NN) solution to mitigate the CPJA trained on the network side and provide the trained NN model to the users' equipment (UEs) to eliminate interference from the signal by the filter. The trained NN is applied to predict the future activity of the interference generated by the attacker. Attack mitigation reduces the impact of the attack while the system's throughput suffers a 6% and 1.8% degradation, and its latency increases by 6.5% and 3.83% compared to the original system for expert and learner agents, respectively. Shavbo Salehi, Hao Zhou 0013, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Yigit Ozcan, Melike Erol-Kantarci |
ICC | 2 |
| 2023 | Traffic Steering for 5G Multi-RAT Deployments using Deep Reinforcement LearningabstractIn 5G non-standalone mode, traffic steering is a critical technique to take full advantage of 5G new radio while optimizing dual connectivity of 5G and LTE networks in multiple radio access technology (RAT). An intelligent traffic steering mechanism can play an important role to maintain seamless user experience by choosing appropriate RAT (5G or LTE) dynamically for a specific user traffic flow with certain QoS requirements. In this paper, we propose a novel traffic steering mechanism based on Deep Q-learning that can automate traffic steering decisions in a dynamic environment having multiple RATs, and maintain diverse QoS requirements for different traffic classes. The proposed method is compared with two baseline algorithms: a heuristic-based algorithm and Q-learning-based traffic steering. Compared to the Q-learning and heuristic baselines, our results show that the proposed algorithm achieves better performance in terms of 6% and 10% higher average system throughput, and 23% and 33% lower network delay, respectively. Md Arafat Habib, Hao Zhou 0013, Pedro Enrique Iturria-Rivera, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Steve Furr, Melike Erol-Kantarci |
CCNC | 2 |
| 2023 | Split Learning for Sensing-Aided Single and Multi-Level Beam Selection in Multi-Vendor RANabstractProper and efficient beam selection is of great importance to unleash the full potential of mmWave communications. Traditionally, each candidate beam is evaluated using reference signals (beam sweeping), however, the exhaustive search method can be time-consuming with high signaling overhead. To avoid such problems, in B5G and 6G, sensing information is considered to be used, as in Integrated Sensing and Communication (ISAC) solutions, and Machine Learning (ML) methods can be applied to map sensing data inputs to an optimal beam index. When using sensing information sources external to the Radio Access Network (RAN) in a multi-vendor disaggregated environment, those methods need to account for issues such as privacy and data ownership. In this work, we apply multi-modal sensing information to the beam selection task. Specifically, we propose a multi-modal sensing-aided ML strategy based on Split Learning (SL) that can cope with deployment challenges in novel RAN architectures. Moreover, the method is applied to single and multi-level beam selection decisions, where the latter considers the case of hierarchical codebook structures. With the proposed approach, accuracy levels above 90% can be achieved while overhead diminishes by 85% or more. SL achieves comparable performance with centralized learning-based strategies, with the added value of accounting for privacy and data ownership issues. We also show that sensing-aided ML-based beam selection decisions in multi-level codebooks are more effective when applied to their first level. Ycaro Dantas, Pedro Enrique Iturria-Rivera, Hao Zhou 0013, Yigit Ozcan, Majid Bavand, Medhat H. M. Elsayed, Raimundas Gaigalas, Melike Erol-Kantarci |
GLOBECOM | 3 |
| 2023 | Policy Poisoning Attacks on Transfer Learning Enabled Resource Allocation for Network SlicingabstractAs wireless networks continue to evolve, machine learning (ML) algorithms are used to address communication challenges and meet various service requirements. While ML methods are promising, they can be prone to malicious attacks, which may degrade user experience and network performance. Specifically, the security challenges of radio access networks (RANs) are highlighted due to frequent interactions with a large number of users, and evaluating these attacks is critical for securing wireless communications. In this paper, for the first time, we investigate the vulnerability of transfer reinforcement learning (TRL) algorithms for resource allocation in 5G RAN slicing. In particular, we first present the system model for RAN slicing, and then the TRL algorithm is introduced for resource allocation. Afterward, we investigate three types of attack methods on the TRL algorithm. The simulations indicate that the attack on an expert agent can affect the performance of a learner agent since the expert shares knowledge with the learner. We show that the effect of the black-box policy poisoning attack on the learner increases latency by 18.31% and reduces throughput by 13.80%, while white-box policy poisoning attacks result in a 48.96% increase in latency and an 87.02% reduction in throughput. Shavbo Salehi, Hao Zhou 0013, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Yigit Ozcan, Melike Erol-Kantarci |
GLOBECOM | 2 |
| 2023 | Beam Selection for Energy-Efficient mmWave Network Using Advantage Actor Critic LearningabstractThe growing adoption of mmWave frequency bands to realize the full potential of 5G, turns beamforming into a key enabler for current and next-generation wireless technologies. Many mmWave networks rely on beam selection with Grid-of-Beams (GoB) approach to handle user-beam association. In beam selection with GoB, users select the appropriate beam from a set of pre-defined beams and the overhead during the beam selection process is a common challenge in this area. In this paper, we propose an Advantage Actor Critic (A2C) learning-based framework to improve the GoB and the beam selection process, as well as optimize transmission power in a mmWave network. The proposed beam selection technique allows performance improvement while considering transmission power improves Energy Efficiency (EE) and ensures the coverage is maintained in the network. We further investigate how the proposed algorithm can be deployed in a Service Management and Orchestration (SMO) platform. Our simulations show that A2C-based joint optimization of beam selection and transmission power is more effective than using Equally Spaced Beams (ESB) and fixed power strategy, or optimization of beam selection and transmission power disjointly. Compared to the ESB and fixed transmission power strategy, the proposed approach achieves more than twice the average EE in the scenarios under test and is closer to the maximum theoretical EE. Ycaro Dantas, Pedro Enrique Iturria-Rivera, Hao Zhou 0013, Majid Bavand, Medhat H. M. Elsayed, Raimundas Gaigalas, Melike Erol-Kantarci |
ICC | 3 |
| 2023 | Hierarchical Reinforcement Learning Based Traffic Steering in Multi-RAT 5G DeploymentsabstractIn 5G non-standalone mode, an intelligent traffic steering mechanism can vastly aid in ensuring a smooth user experience by selecting the best radio access technology (RAT) from a multi-RAT environment for a specific traffic flow. In this paper, we propose a novel load-aware traffic steering algorithm based on hierarchical reinforcement learning (HRL) while satisfying the diverse quality of service requirements of different traffic types. HRL can significantly increase system performance using a bi-level architecture having a meta-controller and a controller. In our proposed method, the meta-controller provides an appropriate threshold for load balancing, while the controller performs traffic admission to an appropriate RAT in the lower level. Simulation results show that HRL outperforms a Deep Q-Learning (DQN) and a threshold-based heuristic baseline with 8.49%, 12.52% higher average system throughput and 27.74%, 39.13% lower network delay, respectively. Md Arafat Habib, Hao Zhou 0013, Pedro Enrique Iturria-Rivera, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Yigit Ozcan, Melike Erol-Kantarci |
ICC | 2 |
| 2022 | Knowledge Transfer based Radio and Computation Resource Allocation for 5G RAN SlicingabstractTo implement network slicing in 5G, resource allocation is a key function to allocate limited network resources such as radio and computation resources to multiple slices. However, the joint resource allocation also leads to a higher complexity in the network management. In this work, we propose a knowledge transfer based resource allocation (KTRA) method to jointly allocate radio and computation resources for 5G RAN slicing. Compared with existing works, the main difference is that the proposed KTRA method has a knowledge transfer capability. It is designed to use the prior knowledge of similar tasks to improve performance of the target task, e.g., faster convergence speed or higher average reward. The proposed KTRA is compared with Q-learning based resource allocation (QLRA), and KTRA method presents a 18.4% lower URLLC delay and a 30.1% higher eMBB throughput as well as a faster convergence speed. Hao Zhou 0013, Melike Erol-Kantarci |
CCNC | 1 |
| 2022 | Joint Sensing and Communications for Deep Reinforcement Learning-based Beam Management in 6GabstractUser location is a piece of critical information for network management and control. However, location uncertainty is unavoidable in certain settings leading to localization errors. In this paper, we consider the user location uncertainty in the mmWave networks, and investigate joint vision-aided sensing and communications using deep reinforcement learning-based beam management for future 6G networks. In particular, we first extract pixel characteristic-based features from satellite images to improve localization accuracy. Then we propose a UK-medoids based method for user clustering with location uncertainty, and the clustering results are consequently used for the beam management. Finally, we apply the DRL algorithm for intra-beam radio resource allocation. The simulations first show that our proposed vision-aided method can substantially reduce the localization error. The proposed UK-medoids and DRL based scheme (UKM-DRL) is compared with two other schemes: K-means based clustering and DRL based resource allocation (K-DRL) and UK-means based clustering and DRL based resource allocation (UK-DRL). The proposed method has 17.2% higher throughput and 7.7% lower delay than UK-DRL, and more than doubled throughput and 55.8% lower delay than K-DRL. Yujie Yao, Hao Zhou 0013, Melike Erol-Kantarci |
GLOBECOM | 2 |
| 2022 | Federated Deep Reinforcement Learning for Resource Allocation in O-RAN SlicingabstractRecently, open radio access network (O-RAN) has become a promising technology to provide an open environment for network vendors and operators. Coordinating the x-applications (xAPPs) is critical to increase flexibility and guarantee high overall network performance in O-RAN. Meanwhile, federated reinforcement learning has been proposed as a promising technique to enhance the collaboration among distributed reinforcement learning agents and improve learning efficiency. In this paper, we propose a federated deep reinforcement learning algorithm to coordinate multiple independent xAPPs in O-RAN for network slicing. We design two xAPPs, namely a power control xAPP and a slice-based resource allocation xAPP, and we use a federated learning model to coordinate two xAPP agents to enhance learning efficiency and improve network performance. Compared with conventional deep reinforcement learning, our proposed algorithm can achieve 11% higher throughput for enhanced mobile broadband (eMBB) slices and 33% lower delay for ultra-reliable low-latency communication (URLLC) slices. Han Zhang 0055, Hao Zhou 0013, Melike Erol-Kantarci |
GLOBECOM | 2 |
| 2022 | Hierarchical Reinforcement Learning for RIS-Assisted Energy-Efficient RANabstractReconfigurable intelligent surface (RIS) is emerging as a promising technology to boost the energy efficiency (EE) of 5G beyond and 6G networks. Inspired by this potential, in this paper, we investigate the RIS-assisted energy-efficient radio access networks (RAN). In particular, we combine RIS with sleep control techniques, and develop a hierarchical reinforcement learning (HRL) algorithm for network management. In HRL, the meta-controller decides the on/off status of the small base stations (SBSs) in heterogeneous networks, while the sub-controller can change the transmission power levels of SBSs to save energy. The simulations show that the RIS-assisted sleep control can achieve significantly lower energy consumption, higher throughput, and more than doubled energy efficiency than no- RIS conditions. Hao Zhou 0013, Long Kong, Medhat H. M. Elsayed, Majid Bavand, Raimundas Gaigalas, Steve Furr, Melike Erol-Kantarci |
GLOBECOM | 1 |
| 2022 | Variational Autoencoder Generative Adversarial Network for Synthetic Data Generation in Smart HomeabstractData is the fuel of data science and machine learning techniques for smart grid applications, similar to many other fields. However, the availability of data can be an issue due to privacy concerns, data size, data quality, and so on. To this end, in this paper, we propose a Variational AutoEncoder Generative Adversarial Network (VAE-GAN) as a smart grid data generative model which is capable of learning various types of data distributions and generating plausible samples from the same distribution without performing any prior analysis on the data before the training phase. We compared the Kullback–Leibler (KL) divergence, maximum mean discrepancy (MMD), and Wasserstein distance between the synthetic data (electrical load and PV production) distribution generated by the proposed model, vanilla GAN network, and the real data distribution, to evaluate the performance of our model. Furthermore, we used five key statistical parameters to describe the smart grid data distribution and compared them between synthetic data generated by both models and real data. Experiments indicate that the proposed synthetic data generative model outperforms the vanilla GAN network. The distribution of VAE-GAN synthetic data is the most comparable to that of real data. Mina Razghandi, Hao Zhou 0013, Melike Erol-Kantarci, Damla Turgut |
ICC | 2 |
| 2022 | Team Learning-Based Resource Allocation for Open Radio Access Network (O-RAN)abstractRecently, the concept of open radio access network (O-RAN) has been proposed, which aims to adopt intelligence and openness in the next generation radio access networks (RAN). It provides standardized interfaces and the ability to host network applications from third-party vendors by x-applications (xAPPs), which enables higher flexibility for network management. However, this may lead to conflicts in network function implementations, especially when these functions are implemented by different vendors. In this paper, we aim to mitigate the conflicts between xAPPs for near-real-time (near-RT) radio intelligent controller (RIC) of O-RAN. In particular, we propose a team learning algorithm to enhance the performance of the network by increasing cooperation between xAPPs. We compare the team learning approach with independent deep Q-learning where network functions individually optimize resources. Our simulations show that team learning has better network performance under various user mobility and traffic loads. With 6 Mbps traffic load and 20 m/s user movement speed, team learning achieves 8% higher throughput and 64.8% lower PDR. Han Zhang 0055, Hao Zhou 0013, Melike Erol-Kantarci |
ICC | 2 |
| 2022 | Deep Reinforcement Learning-based Radio Resource Allocation and Beam Management under Location Uncertainty in 5G mm Wave NetworksabstractMillimeter Wave (mmWave) is an important part of 5G new radio (NR), in which highly directional beams are adapted to compensate for the substantial propagation loss based on UE locations. However, the location information may have some errors such as GPS errors. In any case, some uncertainty, and localization error is unavoidable in most settings. Applying these distorted locations for clustering will increase the error of beam management. Meanwhile, the traffic demand may change dynamically in the wireless environment. Therefore, a scheme that can handle both the uncertainty of localization and dynamic radio resource allocation is needed. In this paper, we propose a UK-means-based clustering and deep reinforcement learning-based resource allocation algorithm (UK-DRL) for radio resource allocation and beam management in 5G mm Wave networks. We first apply UK-means as the clustering algorithm to mitigate the localization uncertainty, then deep reinforcement learning (DRL) is adopted to dynamically allocate radio resources. Finally, we compare the UK-DRL with K-means-based clustering and DRL-based resource allocation algorithm (K-DRL), the simulations show that our proposed UK-DRL-based method achieves 150% higher throughput and 61.5% lower delay compared with K-DRL when traffic load is 4Mbps. Yujie Yao, Hao Zhou 0013, Melike Erol-Kantarci |
ISCC | 2 |
| 2022 | Multiagent Bayesian Deep Reinforcement Learning for Microgrid Energy Management Under Communication FailuresabstractMicrogrids (MGs) are important players for the future transactive energy systems where a number of intelligent Internet of Things (IoT) devices interact for energy management in the smart grid. Although there have been many works on MG energy management, most studies assume a perfect communication environment, where communication failures are not considered. In this article, we consider the MG as a multiagent environment with IoT devices in which AI agents exchange information with their peers for collaboration. However, the collaboration information may be lost due to communication failures or packet loss. Such events may affect the operation of the whole MG. To this end, we propose a multiagent Bayesian deep reinforcement learning (BA-DRL) method for MG energy management under communication failures. We first define a multiagent partially observable Markov decision process (MA-POMDP) to describe agents under communication failures, in which each agent can update its beliefs on the actions of its peers. Then, we apply a double deep$Q$-learning (DDQN) architecture for$Q$-value estimation in BA-DRL, and propose a belief-based correlated equilibrium for the joint-action selection of multiagent BA-DRL. Finally, the simulation results show that BA-DRL is robust to both power supply uncertainty and communication failure uncertainty. BA-DRL has 4.1% and 10.3% higher reward than Nash deep$Q$-learning (Nash-DQN) and alternating direction method of multipliers (ADMM), respectively, under 1% communication failure probability. Hao Zhou 0013, Atakan Aral, Ivona Brandic, Melike Erol-Kantarci |
IEEE Internet Things J. | 1 |
| 2021 | Smart Home Energy Management: Sequence-to-Sequence Load Forecasting and Q-LearningabstractA smart home energy management system (HEMS) can contribute towards reducing the energy costs of customers; however, HEMS suffers from uncertainty in both energy generation and consumption patterns. In this paper, we propose a sequence to sequence (Seq2Seq) learning-based supply and load prediction along with reinforcement learning-based HEMS control. We investigate how the prediction method affects the HEMS operation. First, we use Seq2Seq learning to predict photovoltaic (PV) power and home devices' load. We then apply Q-learning for offline optimization of HEMS based on the prediction results. Finally, we test the online performance of the trained Q-learning scheme with actual PV and load data. The Seq2Seq learning is compared with VARMA, SVR, and LSTM in both prediction and operation levels. The simulation results show that Seq2Seq performs better with a lower prediction error and online operation performance. Mina Razghandi, Hao Zhou 0013, Melike Erol-Kantarci, Damla Turgut |
GLOBECOM | 2 |
| 2021 | Short-Term Load Forecasting for Smart Home Appliances with Sequence to Sequence LearningabstractAppliance-level load forecasting plays a critical role in residential energy management, besides having significant importance for ancillary services performed by the utilities. In this paper, we propose to use an LSTM-based sequence-to-sequence (seq2seq) learning model that can capture the load profiles of appliances. We use a real dataset collected from four residential buildings and compare our proposed scheme with three other techniques, namely VARMA, Dilated One Dimensional Convolutional Neural Network, and an LSTM model. The results show that the proposed LSTM-based seq2seq model outperforms other techniques in terms of prediction error in most cases. Mina Razghandi, Hao Zhou 0013, Melike Erol-Kantarci, Damla Turgut |
ICC | 2 |
| 2021 | RAN Resource Slicing in 5G Using Multi-Agent Correlated Q-Learningabstract5G is regarded as a revolutionary mobile network, which is expected to satisfy a vast number of novel services, ranging from remote health care to smart cities. However, heterogeneous Quality of Service (QoS) requirements of different services and limited spectrum make the radio resource allocation a challenging problem in 5G. In this paper, we propose a multi-agent reinforcement learning (MARL) method for radio resource slicing in 5G. We model each slice as an intelligent agent that competes for limited radio resources, and the correlated Q-learning is applied for inter-slice resource block (RB) allocation. The proposed correlated Q-learning based inter-slice RB allocation (COQRA) scheme is compared with Nash Q-learning (NQL), Latency-Reliability-Throughput Q-learning (LRTQ) methods, and the priority proportional fairness (PPF) algorithm. Our simulation results show that the proposed CO-QRA achieves 32.4% lower latency and 6.3% higher throughput when compared with LRTQ, and 5.8% lower latency and 5.9% higher throughput than NQL. Significantly higher throughput and lower packet drop rate (PDR) is observed in comparison to PPF. Hao Zhou 0013, Medhat H. M. Elsayed, Melike Erol-Kantarci |
PIMRC | 1 |