Xiaowen Ye

dblp:233/7377 · DBLP profile ↗
← Back
19ranked-venue papers
16as first author
17since 2021 · last 2026
0000-0002-7047-0038ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 17 · 15 first-author · 16 since 2021Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Interictal Epileptiform Discharge Detection Using Dual-Domain Features and GAN
abstract
Interictal Epileptiform Discharge is essential for identifying epilepsy. However, the unpredictable and non-stationary nature of electroencephalogram (EEG) patterns poses considerable challenges for reliable identification. Manual interpretation of EEG is subjective and time-consuming. With advancements in machine learning and deep learning, computer-aided approaches for automated IED detection have been rapidly developed. The state-of-the-art convolutional neural network (CNN)-based methods have shown promising results but struggle to capture long-term dependencies in time-series data. In contrast, Transformer excels at modeling sequential information through self-attention mechanisms, overcoming the CNN limitations. This study proposes an IED Detector (IEDD) that integrates convolutional layers and a Transformer to detect IEDs. The IEDD initially employs convolutional layers to extract local features of IEDs, followed by a Transformer to model long-term dependencies. To further extract spatial features, EEG data are represented as a three-dimensional tensor with embedded channel topology, where a CNN captures spatial features at each sampling point and a Long Short-Term Memory (LSTM) network models their temporal evolution. Additionally, due to the scarcity of IED data, a novel Transformer-based Generative Adversarial Network (GAN) is developed to augment the IED dataset. Experimental results show the proposed approach achieves an average accuracy of 96.11% on the augmented Dataset 1 and 95.25% on Dataset 2 for binary classification, with an average sensitivity of 87.26% and precision of 89.96% for multi-label classification. These findings provide valuable insights into advancing deep learning and Transformer-based approaches for automated IED detection.
Wenhao Rao, Jiayang Guo, Chunran Zhu, Meiyan Xu, Naian Xiao, Yijie Pan, Xiaowen Ye, Peipei Gu
IEEE J. Biomed. Health Informatics8
2026 Unlicensed Millimeter-Wave NR-U and WiGig Coexistence: A Hybrid Deep Reinforcement Learning Approach
abstract
This paper investigates an unlicensed millimeter-wave (mmWave) coexistence system, where the new radio-based access to unlicensed spectrum (NR-U) network and the incumbent Wireless Gigabit (WiGig) network share the same spectrum resources to transmit data packets. The total data rate of NR-U is maximized by jointly optimizing the user equipment (UE) scheduling and hybrid beamforming, subject to the constraints of the quality-of-service (QoS) requirements of all UEs and the maximum transmit power at the gNB. Besides, NR-U needs to avoid excessive interference to WiGig, without acquiring any prior information about WiGig. To circumvent this problem, we put forth a model-free joint optimization scheme, referred to as Deep Q-Policy Gradient (DQPG), based on the deep reinforcement learning (DRL) technique. Specifically, a hybrid DRL framework is first proposed to support DQPG, where the deep double Q-network (D2QN) and double-critic-based deep deterministic policy gradient (D3PG) algorithms are invoked to optimize UE scheduling in the discrete action domain and hybrid beamforming in the continuous action domain, respectively. Thereafter, a new reward function and a scaled action selection policy are judiciously designed for DQPG to satisfy various constraints. To capture the complex coupling relationship between different optimization variables, we further propose a parallel experience replay mechanism to maintain training synchronization between D2QN and D3PG. In addition, a transfer learning approach is introduce to accelerate the convergence of DQPG. Simulation results demonstrate that while satisfying the QoS requirements of all UEs, compared with other coexistence schemes, DQPG (i) attains a higher total data rate of NR-U, (ii) causes less interference to WiGig, and (iii) converges faster. Furthermore, under different numbers of UEs and QoS requirements of UEs, DQPG is more robust than benchmarks.
Xiaowen Ye, Xianxin Song, Liqun Fu 0001
IEEE Trans. Mob. Comput.1
2026 Intelligent Omni-Surface-Aided Multi-Objective ISAC: A Meta Hybrid Deep Reinforcement Learning Approach
abstract
This paper studies an intelligent omni-surface (IOS)-aided integrated sensing and communication (ISAC) system, where a base station (BS) provides simultaneous target sensing and communication services with an IOS under outdated and imperfect channel state information (CSI). Both the communication sum-rate and sensing signal-to-noise ratio are maximized through joint optimization of BS beamforming and IOS configuration. To address this problem, we propose an intelligent joint optimization scheme called meta multi-objective hybrid deep reinforcement learning (meta-MHDRL). Specifically, the meta-MHDRL framework first introduces a hybrid deep reinforcement learning (DRL) approach that integrates double-critic-based deep deterministic policy gradient with deep double Q-network algorithms, enabling parallel optimization of both continuous-domain variables (i.e., BS beamforming, IOS reflecting phase shift, and IOS reflecting/refracting amplitudes) and the discrete-domain variable (i.e., IOS refracting phase shift). Thereafter, an objective-preference weight is incorporated into the hybrid DRL framework, such that meta-MHDRL can capture the trade-off between communication and sensing performance. To address the complex coupling relationships among different optimization variables, we further put forth a synchronized experience replay mechanism for meta-MHDRL, which maintains training synchronization among different neural networks. In addition, a meta-learning approach is developed to enhance the generalization ability of meta-MHDRL across different objective-preference weights. Simulation results show that meta-MHDRL attains more Pareto-efficient solutions than other schemes under outdated and imperfect CSI while maintaining stronger robustness across various simulation setups. Besides, we demonstrate the generalization ability of meta-MHDRL for unseen tasks
Xiaowen Ye, Xianxin Song, Yi Wu 0010, Liqun Fu 0001
IEEE Trans. Mob. Comput.1
2026 Integrated Sensing and Communication for Underwater Acoustic Networks Based on Deep Reinforcement Learning
abstract
This paper investigates a new integrated sensing and communication (ISAC) scheme for underwater acoustic (UWA) networks based on deep reinforcement learning, referred to as Deep UWA-ISAC (DeepUSC). Specifically, we consider a UWA-ISAC system, where an autonomous underwater vehicle (AUV) transmits the collected environmental data to the buoy, while sensing the sea area to monitor the unauthorized mobile target. The expected communication rate over a given navigation period is maximized by jointly optimizing the AUV's beamforming and trajectory, subject to the constraints on the average signal-to-noise ratio requirement for target sensing as well as the navigation mission, collision avoidance, and maximum transmit power limit of the AUV. Three key challenges for DeepUSC are: (i) long propagation delays in the UWA-ISAC system may cause interference from the previous echo to the current ISAC signal; (ii) the mobility pattern of the target is unknown in advance; and (iii) the AUV navigation-oriented ISAC problem is a long-term optimization problem as the navigation mission typically lasts for a long period. To circumvent the above challenges, DeepUSC is developed based on a specific partially observable Markov decision process model termed episode task, where each navigation period is considered as an episode and the navigation mission corresponds to the episode task. Through judicious design of a reward function and action selection policy, DeepUSC can satisfy various preset constraints without requiring prior knowledge of the target's mobility. Besides, to enable efficient learning in episode tasks, we propose an episodic experience replay mechanism that dynamically prioritizes high-value recent experiences and utilizes all experiences generated within each episode to jointly train the neural network. Simulation results demonstrate that compared with benchmarks, DeepUSC yields a higher communication rate while satisfying all constraints, converges faster, and is more robust against different simulation setups.
Xiaowen Ye, Xianxin Song, Yi Wu 0010, Hao Xu 0003, Jun Zhang 0023
IEEE Trans. Mob. Comput.1
2026 Integrated Sensing and Communications for Low-Altitude Economy: A Deep Reinforcement Learning Approach
abstract
This paper studies an integrated sensing and communications (ISAC) system for low-altitude economy (LAE), where a ground base station (GBS) provides communication and navigation services for authorized unmanned aerial vehicles (UAVs), while sensing the low-altitude airspace to monitor the unauthorized mobile target. The expected communication sum-rate over a given flight period is maximized by jointly optimizing the beamforming at the GBS and UAVs’ trajectories, subject to the constraints on the average signal-to-noise ratio requirement for sensing, the flight mission and collision avoidance of UAVs, as well as the maximum transmit power at the GBS. Typically, this is a sequential decision-making problem with the given flight mission. Thus, we transform it to a specific Markov decision process (MDP) model called episode task. Based on this modeling, we propose a novel LAE-oriented ISAC scheme, referred to as Deep LAE-ISAC (DeepLSC), by leveraging the deep reinforcement learning (DRL) technique. In DeepLSC, a reward function and a new action selection policy termed constrained noise-exploration policy are judiciously designed to fulfill various constraints. To enable efficient learning in episode tasks, we develop a hierarchical experience replay mechanism, where the gist is to employ all experiences generated within each episode to jointly train the neural network. Besides, to enhance the convergence speed of DeepLSC, a symmetric experience augmentation mechanism, which simultaneously permutes the indexes of all variables to enrich available experience sets, is proposed. Simulation results demonstrate that compared with benchmarks, DeepLSC yields a higher sum-rate while meeting the preset constraints, achieves faster convergence, and is more robust against different settings.
Xiaowen Ye, Yuyi Mao, Xianghao Yu, Shu Sun 0001, Liqun Fu 0001, Jie Xu 0002
IEEE Trans. Wirel. Commun.1
2025 Leveraging Propagation Delays: A Delay-Aware Multiagent Reinforcement Learning MAC Protocol for Underwater Acoustic Networks
abstract
Underwater acoustic networks are typically distributed in nature and have been attracting much research interest recently. Such networks are characterized by long propagation delays, which pose challenges for the medium access control (MAC) protocol design in underwater acoustic networks. In this paper, instead of considering long propagation delay as a negative effect, we exploit it as an advantage. We propose a multi-agent reinforcement learning (MARL)-based MAC protocol without requiring acknowledgement feedback, named Delay-Aware Multi-Agent Reinforcement Learning Multiple Access (DA-MARLA), which leverages propagation delays to achieve higher throughput. Furthermore, the throughput achieved can exceed that of systems with zero propagation delay. In developing DA-MARLA, we introduce a novel MARL algorithm, termed delay-aware multi-agent proximal policy optimization (DA-MAPPO). Specifically, to leverage the long propagation delays, we propose two period-based mechanisms that coordinate nodes’ transmission schedules to reduce collisions and balance cooperation and competition among nodes. To ensure reliable operation, we incorporate a sequential policy update mechanism. This mechanism offers accurate performance evaluation for each node and establishes update sequences during centralized training. Simulation results show that our method consistently outperforms baseline methods across various network topologies while maintaining robustness, demonstrating that the propagation delays can be effectively utilized to enhance network efficiency and providing a new avenue for performance improvement in underwater acoustic networks.
Xiaowen Ye, Liqun Fu 0001
IEEE Internet Things J.2
2025 Energy-Efficient Link Adaptation for Underwater Acoustic Communications Based on Meta Deep Reinforcement Learning
abstract
Due to the harsh channel conditions and operational difficulties in battery recharging, energy-efficient transmission is critical in underwater acoustic communications (UACs). This paper investigates a new link adaptation technique for UACs that jointly optimizes transmission frequency, power, and rate to maximize energy efficiency. Conventional optimization-based approaches typically require real-time and perfect channel state information and have high computational complexity, making them difficult to implement in realistic systems. To circumvent this problem, we put forth MetaDT, a model-free link adaptation technique combining deep reinforcement learning with meta-learning. To enable powerful reasoning and fast decision-making, we further propose a dueling echo state network (ESN) with separate output architecture for incorporation into MetaDT. Besides, to enable MetaDT to quickly adapt to diverse new/unseen environments, a low-complexity meta-learning is developed to find the optimal meta-parameters of the dueling ESN architecture. Numerical results show that compared to various benchmarks, MetaDT attains significant energy efficiency gains and is more robust against different transmission distances and numbers of multi-paths. In comparison to conventional neural networks, dueling ESN shortens the run-time of MetaDT by more than 89.58% and is more efficient for temporal inference. In addition, we demonstrate the generalization capability of MetaDT with meta-learning to new/unseen environment configurations.
Xiaowen Ye, Liqun Fu 0001, Xianxin Song, Yi Wu 0010
IEEE Internet Things J.1
2025 Joint MCS Adaptation and Beamforming Design for Multiuser MISO Systems: A Constrained Hybrid Deep Reinforcement Learning Approach
abstract
This paper investigates the joint modulation-coding scheme (MCS) adaptation and beamforming design for multi-user multi-input single-output (MISO) systems, where one base station serves multiple user equipments (UEs) under imperfect and outdated channel state information (CSI). The sum-rate of the system is maximized while satisfying all UEs’ data rate requirements and the maximum transmit power constraint at the BS. Most existing beamforming designs overlooked that only a finite number of MCSs can be supported in practical communication systems. Moreover, previous works rely on perfect and real-time CSI for decision-making, neglecting processing delays and channel estimation errors. To circumvent the above issues, this paper puts forth an intelligent joint optimization scheme based on deep reinforcement learning (DRL) techniques. Specifically, a new DRL framework, termed constrained hybrid DRL (CHDRL), is first proposed, which incorporates Lagrangian primal-dual optimization theory and a constrained action selection policy into conventional DRL to tackle various constraints. By integrating deep Q-network (DQN) and deep deterministic policy gradient algorithms, CHDRL is capable of simultaneously optimizing MCS in the discrete action domain and beamforming in the continuous action domain. In addition, to handle the large discrete action space of DQN, we develop an action branch architecture for CHDRL to enable independent and concurrent MCS decisions at different UEs. Finally, a multi-parameter experience replay mechanism is designed to synchronously train Lagrangian multipliers and neural network parameters. Simulation results demonstrate that under imperfect and outdated CSI, CHDRL outperforms other benchmark schemes by (i) achieving a significantly higher sum-rate, (ii) meeting more UEs’ data rate requirements, and (iii) being more robust against different CSI delays and numbers of UEs.
Xiaowen Ye, Yuyi Mao, Xianghao Yu, Liqun Fu 0001
IEEE Internet Things J.1
2025 Digital-Twin-Enhanced Deep Reinforcement Learning for Intelligent Omni-Surface Configurations in MU-MIMO Systems
abstract
Intelligent omni-surface (IOS) is a promising technique to enhance the capacity of wireless networks, by reflecting and refracting the incident signal simultaneously. Traditional IOS configuration schemes, relying on all subchannels’ channel state information and user equipments’ mobility, are difficult to implement in complex realistic systems. Existing works attempt to address this issue employing deep reinforcement learning (DRL), but this method requires a lot of trial-and-error interactions with the external environment for efficient results and thus cannot satisfy the real-time decision making. To enable model-free and real-time IOS control, this article puts forth a new framework that integrates DRL and digital twins. As a first step, deep reinforcement learning IOS (DeepIOS), a DRL based IOS configuration scheme with the goal of maximizing the sum data rate, is developed to jointly optimize the phase-shift and amplitude of IOS in multiuser multiple-input-multiple-output (MU-MIMO) systems. Thereafter, in order to further reduce the computational complexity, DeepIOS introduces an action branch architecture, which decides two optimization variables in parallel in a separate fashion. Finally, a digital twin module is constructed through supervised learning as a preverification platform for DeepIOS, such that the decision making’s real-time can be guaranteed. The formulated framework is a closed-loop system, in which the physical space provides data to establish and calibrate the digital space, while the digital space generates a large number of experience samples for DeepIOS training and sends the trained parameters to the IOS controller for configurations. Numerical results show that compared with random and MAB schemes, the proposed framework attains a higher data rate and is more robust to different settings. Furthermore, the action branch architecture reduces DeepIOS’s computational complexity, and the digital twin module improves DeepIOS’s convergence speed and run-time.
Xiaowen Ye, Xianghao Yu, Liqun Fu 0001
IEEE Internet Things J.1
2025 Intelligent Omni-Surface-Aided Integrated Sensing and Communications Based on Deep Reinforcement Learning With Knowledge Transfer
abstract
This paper investigates an intelligent omni-surface (IOS)-assisted integrated sensing and communication (ISAC) system, where a base station provides both target sensing and communication services with an IOS. The sensing signal-to-noise ratio (SNR) is maximized while satisfying the communication requirement by optimizing IOS configurations. Conventional approaches typically need real-time and accurate channel state information (CSI) and have high computational complexity, making them difficult to implement in realistic systems. To circumvent this problem, this paper puts forth a new framework based on deep reinforcement learning (DRL) with knowledge transfer. In particular, an online learning scheme called Deep reinforcement learning IOS-ISAC (DeepOSC), is first proposed to optimize the reflecting and refracting coefficients of the IOS. Thereafter, to enable powerful reasoning and fast decision-making, we incorporate an echo state network (ESN) with separate output into DeepOSC. To further accelerate convergence, two transfer learning approaches, namely staged policy reuse (SPR) and staged policy distillation (SPD), are developed to guide the learning process of a newly deployed agent by leveraging policies of pre-trained agents. Numerical results show that compared to various benchmarks, DeepOSC attains significant sensing and communication performance gains and is more robust against outdated CSI coefficients. In addition, in comparison to conventional neural networks, ESN shortens the run-time of DeepOSC by more than ten times and is more efficient for temporal inference. Besides, we demonstrate the capabilities of SPR and SPD in accelerating the convergence of DeepOSC.
Xiaowen Ye, Yuyi Mao, Xianghao Yu, Liqun Fu 0001
IEEE Trans. Wirel. Commun.1
2024 Joint Codebook Selection and MCS Adaptation for MmWave eMBB Services Based on Deep Reinforcement Learning
abstract
This article investigates the joint codebook selection and modulation-coding-scheme (MCS) adaptation issue for the enhanced mobile broadband (eMBB) service in millimeter-wave (mmWave) cellular systems. The proposed scheme guarantees efficient mmWave eMBB service through an intelligent joint codebook selection and MCS adaptation scheme that exploits deep reinforcement learning (DRL), referred to as DeepCM. DeepCM’s objective maximizes the transmission data rate while satisfying a target block error rate (BLER) constraint. A first step formulates this joint problem into a two-time-scale system that performs MCS adaptation on a small-time scale, whereas a second step optimizes the codebook on a large-time scale. DeepCM introduces a new DRL algorithm, termed dual-deep Q-network (DQN), by incorporating the operations on two time scales into the original DQN. Dual-DQN essentially enables the operations on different time scales to benefit from each other, through closed-loop decision guidance and reward evaluation. Thereafter, to fulfill the preset BLER constraint, DeepCM uses a constrained$\epsilon $-greedy strategy for decision-making and further modifies the conventional DRL training mechanism. Basically, DeepCM continuously adjusts the agent’s feasible-action space toward the system objective. With the constrained dual-DQN, DeepCM can attain its goal even without any prior network information. Simulation results show that DeepCM, compared with Thompson Sampling-DRL, DRL-OLLA, and TS2 schemes, guarantees the target BLER requirement while yielding a much higher data rate. Various simulations demonstrate the powerful robustness of DeepCM under miscellaneous scenarios. Furthermore, DeepCM can handle well dynamic target-BLER change.
Xiaowen Ye, Liqun Fu 0001, John M. Cioffi
IEEE Internet Things J.1
2024 Joint MCS Adaptation and RB Allocation in Cellular Networks Based on Deep Reinforcement Learning With Stable Matching
abstract
Joint modulation-coding scheme (MCS) adaptation and resource block (RB) allocation is an effective approach to guarantee different quality of service (QoS) requirements of all UEs under dynamic network environments. In this article, we consider a fifth generation (5G) cellular network with time-varying wireless channels, in which the BS serves multiple user equipments (UEs) under limited available RBs. We aim to minimize the total RB consumption subject to the rigorous constraints of each UE's QoS requirement. To attain this objective, this paper puts forth an online learning technique, referred to as integrated Deep Reinforcement learning and stable Matching (DeepRM), in the sense that the MCS adaptation and RB allocation decisions are conducted without acquiring the real-time channel quality indicator (CQI) feedback. DeepRM is a closed-loop framework, where the output of deep reinforcement learning (DRL) is imported into the stable matching to guide optimal RB allocation whilst the output of stable matching is fed into the DRL framework to assist efficient MCS decision-making. Specifically, in DeepRM, we first develop a powerful DRL algorithm, termed as Action-and-Reward Branching Deep Q-network (ARBDQ), by incorporating the action branch architecture into conventional DRL and modifying the traditional deep neural network training mechanism, to perform judicious MCS decisions on different links in parallel. Then, a new many-to-one stable matching algorithm, called adaptive deferred acceptance, is exploited to dynamically adjust the RB quota of each UE in a computationally efficient fashion. Simulation results demonstrate that compared with ACO-HM, OLLA-ADA, and ARBDQ-Random algorithms, DeepRM induces much less RB consumption while guaranteeing the QoS requirements of all UEs in various network scenarios. Furthermore, under miscellaneous QoS requirement, number of UEs, and CQI reporting period setups, DeepRM is more robust than other baselines.
Xiaowen Ye, Liqun Fu 0001
IEEE Trans. Mob. Comput.1
2024 Joint Codebook Selection and UE Scheduling for Unlicensed MmWave NR-U/WiGig Coexistence Based on Deep Reinforcement Learning
abstract
Unlicensed millimeter-wave (mmWave) communication is a promising technique for the New Radio-based access to Unlicensed spectrum (NR-U) network to guarantee the ever-increasing data rate demand. A critical challenge of NR-U in unlicensed mmWave bands is to maintain equitable and harmonious coexistence with the original Wireless Gigabit (WiGig) network. In this article, we develop an intelligent joint codebook selection and user equipment (UE) scheduling scheme for mmWave NR-U and WiGig coexistence networks. Specifically, we first formulate the joint problem as a two-time scale system, wherein the codebook selection is performed on the large-time scale whilst the UE scheduling is optimized on the small-time scale. To address the multi-time scale issue, we put forth a new deep reinforcement learning (DRL) algorithm that enables operations on different time scales to benefit each other towards the target system objective, referred to as layered deep Q-network (L-DQN). Thereafter, with the judicious definitions of the state, action, and reward in L-DQN paradigms, we propose the Deep reinforcement learning based CodeBook selection and UE scheduling (DeepCBU) scheme. DeepCBU aims to attain different trade-offs between two conflicting goals, i.e., i) maximizing the total data rate of NR-U with as little interference to WiGig as possible and ii) guaranteeing the fairness among UEs, e.g., the quality of service (QoS) requirement of each UE. To fulfill this mission, we modify the conventional deep neural network architecture of DeepCBU by introducing the target branch for each objective. The gist is that different target branches evaluate the contribution of DeepCBU's strategy to different goals, and the decision of DeepCBU is determined by all target branches in a weighted fashion. Simulation results demonstrate that compared with DRL-dirLBT, TS-dirLBT, and TS-DRL schemes, DeepCBU is more Pareto efficient even without any prior network knowledge, e.g., UE mobility, random channel fading, and transmissions of WiGig, in terms of the data rate of NR-U, the data rate of WiGig, and the number of satisfied UEs. Furthermore, DeepCBU is robust to miscellaneous QoS requirement setups.
Xiaowen Ye, Liqun Fu 0001
IEEE Trans. Mob. Comput.1
2024 Deep Reinforcement Learning-Based Scheduling for NR-U/WiGig Coexistence in Unlicensed mmWave Bands
abstract
This paper investigates the coexistence of the New Radio-based access to Unlicensed spectrum (NR-U) network and the Wireless Gigabit (WiGig) network in unlicensed millimeter-wave (mmWave) bands. To enable the NR-U network to achieve equitable and harmonious spectrum sharing with WiGig systems, we develop two new classes of user equipment (UE) scheduling schemes by exploiting the deep reinforcement learning (DRL) technique. Specifically, we first propose the distributed deep reinforcement learning scheduling (DeepDS) scheme, wherein multiple deep neural networks (DNNs) are used to make decisions for different panels in an independent fashion. Thereafter, to reduce the computational cost of adopting multiple DNNs, we design the centralized deep reinforcement learning scheduling (DeepCS) scheme that introduces the shared DNN framework to perform decisions for all panels in parallel at one time. The objective of both DeepDS and DeepCS is to maximize the total data rate of the NR-U network with as little interference to WiGig systems as possible, while satisfying the quality of service (QoS) requirement for each UE. We first formulate this problem into the constrained Markov decision process framework. To address the multi-constraint issue, we put forth a new DRL algorithm that incorporates the Lagrangian primal-dual optimization into the deep Q-network framework, referred to as adaptive multi-constraint deep Q-network (AMC-DQN). With AMC-DQN, both DeepDS and DeepCS can achieve their goals even without acquiring prior operations about the WiGig network. Simulation results show that compared with the state-of-the-art omniLBT and dirLBT, both DeepDS and DeepCS yield significant performance benefits in terms of the total network data rate. We also demonstrate the ability of DeepDS and DeepCS to satisfy the QoS requirements of different UEs and their robustness against various simulation setups. Furthermore, compared with DeepDS, DeepCS can save a large amount of computational cost although at the expense of a slightly lower data rate.
Xiaowen Ye, Liqun Fu 0001
IEEE Trans. Wirel. Commun.1
2022 Deep Reinforcement Learning Based Scheduling Scheme for the NR-U/WiGig Coexistence in Unlicensed mmWave Bands
abstract
This paper considers the coexistence of the New Radio-based access to unlicensed spectrum (NR-U) network and the Wireless Gigabit (WiGig) network in unlicensed millimeter-wave (mmWave) bands. We aim to design a new scheduling scheme for the NR-U network to maximize its total data rate while satisfying the quality of service (QoS) requirement for each user equipment (UE). Specifically, we first formulate this problem into the constrained Markov decision process (CMDP) framework. Then the Lagrangian duality method is applied to relax the hard constraints in CMDP into the soft constraints. To address the multi-constraint issue, we put forth a new deep reinforcement learning (DRL) algorithm that incorporates the constraints into the DRL framework, referred to as adaptive multi-constraint deep Q-network (AMC-DQN). A prominent advantage of AMC-DQN is that it enables the NR-U network to access the shared spectrum without acquiring prior information about the WiGig network. Simulation results show that compared with the omnidirectional listen-before-talk (omniLBT) and directional LBT (dirLBT), the AMC-DQN based scheduling scheme yields the total data rate gain of the NR-U network by 158% and 38%, respectively. The results also demonstrate the ability of AMC-DQN to satisfy the QoS requirements of different UEs. Furthermore, AMC-DQN brings less interference to the WiGig network in comparison to baselines.
Xiaowen Ye, Liqun Fu 0001
ICC2
2022 Deep Reinforcement Learning Based MAC Protocol for Underwater Acoustic Networks
abstract
Long propagation delay that causes throughput degradation of underwater acoustic networks (UWANs) is a critical issue in the medium access control (MAC) protocol design in UWANs. This paper develops a deep reinforcement learning (DRL) based MAC protocol for UWANs, referred to as delayed-reward deep-reinforcement learning multiple access (DR-DLMA), to maximize the network throughput by judiciously utilizing the available time slots resulted from propagation delays or not used by other nodes. In the DR-DLMA design, we first put forth a new DRL algorithm, termed asdelayed-reward deep Q-network (DR-DQN). Then we formulate the multiple access problem in UWANs as a reinforcement learning (RL) problem by defining state, action, and reward in the parlance of RL, and thereby realizing the DR-DLMA protocol. In traditional DRL algorithms, e.g., the original DQN algorithm, the agent can get access to the “reward” from the environment immediately after taking an action. In contrast, in our design, the “reward” (i.e., the ACK packet) is only available after twice the one-way propagation delay after the agent takes an action (i.e., to transmit a data packet). The essence of DR-DQN is to incorporate the propagation delay into the DRL framework and modify the DRL algorithm accordingly. In addition, in order to reduce the cost of online training deep neural network (DNN), we provide a nimble training mechanism for DR-DQN. The optimal network throughputs in various cases are given as a benchmark. Simulation results show that our DR-DLMA protocol with nimble training mechanism can: (i) find the optimal transmission strategy when coexisting with other protocols in a heterogeneous environment; (ii) outperform state-of-the-art MAC protocols (e.g., slotted FAMA and DOTS) in a homogeneous environment; and (iii) greatly reduce energy consumption and run-time compared with DR-DLMA with traditional DNN training mechanism.
Xiaowen Ye, Yiding Yu, Liqun Fu 0001
IEEE Trans. Mob. Comput.1
2022 Multi-Channel Opportunistic Access for Heterogeneous Networks Based on Deep Reinforcement Learning
abstract
This paper investigates a new medium access control (MAC) protocol for multi-channel heterogeneous networks (HetNets) based on deep reinforcement learning (DRL), referred to as multi-channel deep-reinforcement learning multiple access (MC-DLMA). Specifically, we consider a HetNet where different radio networks adopt different MAC protocols to transmit data packets to a common access point on different wireless channels. Three key challenges for the MC-DLMA node are (i) no environmental knowledge is known in advance; (ii) the channels in HetNets are allocated to nodes using different MAC protocols; (iii) the capacities of different channels may be different. The main goal of MC-DLMA is to find an optimal access policy to transmit on those pre-allocated channels and expedite more efficient spectrum utilization. Due to the complex temporal correlation of spectrum states in HetNets, the traditional DRL technique, e.g., original deep Q-network (DQN) algorithm, is no longer applicable to our problem. In our MC-DLMA design, an advanced class of recurrent neural network, termed as Gated Recurrent Unit (GRU), is embedded into the original DQN technique to aggregate observations over time and reason the underlying temporal feature in multi-channel HetNets. Furthermore, we analytically give the optimal spectrum access patterns and derive the optimal throughputs in various HetNet scenarios. With judicious definitions of the state, action, and reward function in the parlance of the DRL framework, simulation results show that MC-DLMA can (i) find the optimal spectrum access strategies in various HetNets, (ii) outperform the random access policy, the whittle index policy, and the original DQN, (iii) perform cooperative transmission in a fully distributed manner in the presence of multiple agents, and (iv) adapt well to the environmental changes.
Xiaowen Ye, Yiding Yu, Liqun Fu 0001
IEEE Trans. Wirel. Commun.1
2020 MAC Protocol for Multi-channel Heterogeneous Networks Based on Deep Reinforcement Learning
abstract
This paper considers the problem of efficient spectrum utilization in heterogeneous wireless networks (HetNets), wherein different radio networks adopt different medium access control (MAC) protocols to transmit data packets to a common access point on different wireless channels. To allow emerging radio nodes to transmit on those pre-allocated channels and to expedite more efficient spectrum utilization, we exploit the advanced deep reinforcement learning technique to develop a new generation of MAC protocols, referred to as multi-channel deep-reinforcement learning multiple access (MC-DLMA). The emerging radio nodes that adopt MC-DLMA can make full use of the underutilized spectrum resource and maximize the sum throughput of the overall HetNet by learning the transmission patterns of the existing radio nodes. For benchmarking, we derive the optimal throughputs analytically and demonstrate that MC-DLMA can achieve the near-optimal results. Moreover, compared with other baselines (e.g., the Whittle Index policy and the random access policy), our MC-DLMA can significantly improve the sum throughput of the HetNet in various scenarios.
Xiaowen Ye, Yiding Yu, Liqun Fu 0001
GLOBECOM1
2017 Optimal design of inverter feedback device for urban rail traction power supply system
abstract
A configuration method for inverter feedback device has been proposed based on multi-train operation and traction power supply simulation. In order to ensure the safety of traction power supply, the train network voltage and rail potential are considered. And the regeneration failure ratio is sharply reduced. Based on the maximum effective power in the continuous time interval, considering the factors such as the daily energy saving and the payback period of the device, the inverter device is configured with reasonable capacity. By taking Chongqing Metro Line 9 as an example, the location and capacity of the inverter feedback device are selected by using the simulation platform.
Xiaowen Ye, Wei Liu 0132, Ying Lou
IECON1