VLDB 2026 Research / reviewers in the wild / expert
Yang Xiao 0013
dblp:181/1848-13
· DBLP profile ↗
17ranked-venue papers
6as first author
16since 2021 · last 2025
0000-0001-6897-5531ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 13 · 6 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Constraint-Aware Probabilistic Packet Forwarding Based on Deep Reinforcement Learning
Guocheng Lin, Yang Xiao 0013, Jun Liu 0014 |
Networking | 2 |
| 2025 | PVD-TD3: A Latency-Oriented Multi-Agent Reinforcement Learning Algorithm for Multipath Routing in DetNet
Yang Xiao 0013, Jun Liu 0014 |
Networking | 2 |
| 2025 | Enabling Adaptive Optimization of Energy Efficiency and Quality of Service in NR-V2X Communications via Multiagent Deep Reinforcement LearningabstractThe Third Generation Partnership Project has standardized new-radio vehicle-to-everything (NR-V2X) to facilitate advanced use cases for safety-critical message conveyance. However, there is a paucity of resource allocation research for high-performance NR-V2X communications. In this article, we investigate an adaptive optimization issue of energy efficiency (EE) and Quality of Service (QoS) in NR-V2X networks inspired by the Tchebycheff function. To address this issue, we first formulate the resource allocation task as a time-variant mixed-integer nonlinear programming (MINLP) problem. Then, we propose a fully decentralized multiagent deep reinforcement learning (MADRL)-based algorithm characterized by a multitask-actor shared-critic (MTA-SC) architecture and a localized training and distributed execution (LTDE) framework to promote efficient learning and minimize information exchange. Finally, we implement the proposed algorithm in a three-lane highway NR-V2X scenario. Numerical results demonstrate that the proposed algorithm comprehensively outperforms benchmarks in terms of convergence performance, scalability, and robustness. Yuqian Song, Yang Xiao 0013, Jun Liu 0014 |
IEEE Internet Things J. | 2 |
| 2025 | Adaptive Joint Routing and Caching in Knowledge-Defined Networking: An Actor-Critic Deep Reinforcement Learning ApproachabstractBy integrating the software-defined networking (SDN) architecture with the machine learning-based knowledge plane, knowledge-defined networking (KDN) is revolutionizing established traffic engineering (TE) methodologies. This paper investigates the challenging joint routing and caching problem in KDN-based networks, managing multiple traffic flows to improve long-term quality-of-service (QoS) performance. This challenge is formulated as a computationally expensive non-convex mixed-integer non-linear programming (MINLP) problem, which exceeds the capacity of heuristic methods to achieve near-optimal solutions. To address this issue, we present DRL-JRC, an actor-critic deep reinforcement learning (DRL) algorithm for adaptive joint routing and caching in KDN-based networks. DRL-JRC orchestrates the optimization of multiple QoS metrics, including end-to-end delay, packet loss rate, load balancing index, and hop count. During offline training, DRL-JRC employs proximal policy optimization (PPO) to smooth the policy optimization process. In addition, the learned policy can be seamlessly integrated with conventional caching solutions during online execution. Extensive experiments demonstrate the comprehensive superiority of DRL-JRC over baseline methods in various scenarios. Meanwhile, DRL-JRC consistently outperforms the heuristic baseline under partial policy deployment during execution. Compared to the average performance of the baseline methods, DRL-JRC reduces the end-to-end delay by 51.14% and the packet loss rate by 40.78%. Yang Xiao 0013, Huihan Yu, Yixing Wang, Jun Liu 0014, Nirwan Ansari |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | CGTR: Leveraging Contrastive Learning and Graph Transformer for Deep Reinforcement Learning Based Robust RoutingabstractAs a crucial role in communication networks, the routing algorithm determines how to transmit traffic from sources to destinations. In recent years, Deep Reinforcement Learning (D RL) has been introduced into routing algorithms to address dynamic traffic demands in complex networks. However, most DRL-based routing algorithms are implemented by traditional neural networks, which can only handle fixed-size matrices and operate on a fixed topology. Fortunately, Graph Neural Network (GNN) has been proposed to process graph-structured data and generalize on different graphs. Furthermore, contrastive learning has also been successfully applied to DRL for decision-making in multiple environments. In this paper, we introduce GNN and contrastive learning to DRL-based routing algorithms, and propose a Contrastive Graph Transformer Routing (CGTR) algorithm to improve the robustness of routing against link fail-ures. Aiming at probabilistic packet routing scenarios, we design Edge-enhanced Graph Transformer and Contrastive Grouping Routing in CG TR, enabling it to perceive unseen link failures without retraining. To evaluate the robustness of CG TR, we conduct extensive experiments with different patterns of traffic demands on both generated and real-world network topologies. The experimental results demonstrate that CGTR outperforms all baselines on unseen topologies with link failures, highlighting its stronger robustness. Junze Li, Yang Xiao 0013, Sixu Liu, Jun Liu 0014 |
ICC | 3 |
| 2024 | Scalable QoS-Aware Multipath Routing in Hybrid Knowledge-Defined Networking With Multiagent Deep Reinforcement LearningabstractMultipath routing remains a challenging issue in traffic engineering (TE) as existing solutions are incapable of handling the evolving network dynamics and stringent quality-of-service (QoS) requirements. To address it, multi-agent deep reinforcement learning (MADRL) is a promising technique that provides more elaborate multipath routing strategies. However, prevalent MADRL-based solutions still suffer inapplicability as they fail to ensure both scalability and QoS awareness. In this paper, we leverage the emerging hybrid knowledge-defined networking (KDN) architecture, and propose a collaborative MADRL-based multipath routing algorithm. Two novel mechanisms, i.e., parallel agent replication and periodic policy synchronization, are devised for agent design to ensure the practicality of the proposed method. In addition, an efficient communication mechanism is established to facilitate multi-agent collaboration by enabling scalable observation and reward exchange. Featuring a multi-agent twin-actor-critic (MA-TAC) learning structure and a proximal policy optimization (PPO) -based training process, the proposed algorithm consists of alternately scheduled execution and training phases for practical deployment. We compare the performance of our proposed method with those of several benchmark methods. Extensive simulation results demonstrate that the proposed method achieves significantly better scalability, QoS awareness, and stability than the benchmark methods under various environment settings. Yang Xiao 0013, Huihan Yu, Jun Liu 0014 |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | Collaborative Multi-Agent Deep Reinforcement Learning for Energy-Efficient Resource Allocation in Heterogeneous Mobile Edge Computing NetworksabstractMobile edge computing (MEC) is an enabling technology for next-generation network architectures to deliver more diversified communication services and meet more demanding quality-of-service (QoS) requirements. However, owing to the growing scale and heterogeneity of the networks, energy-efficient resource allocation in heterogeneous MEC (Het-MEC) networks faces great challenges. As an emerging research area, multi-agent deep reinforcement learning (MADRL) is expected to realize autonomous resource allocation in Het-MEC networks by learning from trial and error. Nevertheless, existing MADRL-based solutions are usually not applicable to practical scenarios by the limitations of centralized control schemes or massive signaling overhead. To address the issue, we first formulate the energy-efficient resource allocation problem in Het-MEC networks as a time-variant mixed-integer nonlinear programming (MINLP) problem. Then, we propose a fully decentralized collaborative MADRL-based algorithm featuring the multi-actor shared-critic (MASC) architecture and the regional training distributed execution (RTDE) framework, which effectively stabilizes model training and reduces information exchange, respectively. Finally, we conduct extensive simulations to evaluate the performance of the proposed algorithm. Numerical results demonstrate that the proposed algorithm comprehensively outperforms several mainstream baseline methods in terms of convergence performance, scalability, and robustness. Yang Xiao 0013, Yuqian Song, Jun Liu 0014 |
IEEE Trans. Wirel. Commun. | 1 |
| 2023 | Deep Reinforcement Learning Based Probabilistic Cognitive Routing: An Empirical Study with OMNeT++ and P4abstractThis paper presents an empirical study on deep reinforcement learning (DRL) based probabilistic cognitive routing using the OMNeT++ framework and programming protocol-independent packet processors (P4). The proposed algorithm combines the power of DRL and cognitive routing to achieve efficient and adaptive probabilistic routing in software-defined networking (SDN) environments. To facilitate the research, we develop a dedicated network simulation environment using the OMNeT++ framework and a self-developed SDN platform based on P4. The empirical study highlights the importance of a comprehensive training and validation process in both simulation and real-world SDN environments. Through closed-loop training, the cognitive routing framework provides real-time feedback from the actual network environment to the simulation environment, allowing the agent to excel in real-world network environments. Meanwhile, the results demonstrate that solely testing the algorithm in either environment is inadequate for evaluating its performance accurately. Yixing Wang, Yang Xiao 0013, Yuqian Song, Jingli Zhou, Jun Liu 0014 |
CNSM | 2 |
| 2023 | Deep Reinforcement Learning Based Dynamic Routing Optimization for Delay-Sensitive ApplicationsabstractWith the rapid development of the Internet and the approaching of the next-generation networking, the number and variety of delay-sensitive applications have increased dramatically. Nowadays, how to properly route delay-sensitive packets in complex network environment and meet the stringent quality-of-service (QoS) requirements of delay-sensitive applications remains a great challenge. Towards this end, this paper proposes a deep reinforcement learning (DRL)-based routing algorithm for delay-sensitive applications featuring the proximal policy optimization (PPO) method and the front-convergent actor-critic network (FCACN) technique. To meet the high demand of delay-sensitive applications, we consider the packet survival time (ST) to help our algorithm perform better and make up for the shortage of the time-to-live (TTL) mechanism in IP network. We conduct extensive experiments to prove the efficiency and reliability of the proposed algorithm. Experimental results show that the proposed algorithm outperforms two traditional routing protocols and two state-of-the-art DRL-based routing algorithms in terms of minimizing delay and packet loss rate. Yang Xiao 0013, Guocheng Lin, Gang He 0008, Fang Liu 0026, Jun Liu 0014 |
GLOBECOM | 2 |
| 2023 | GAPPO - A Graph Attention Reinforcement Learning based Robust Routing AlgorithmabstractRouting algorithms, which determine how to deliver traffic from the source to the destination, are essential for next-generation networks and the internet. To make optimal routing decisions in complex network environments, researchers have leveraged Deep Reinforcement Learning (DRL) to design next-generation routing mechanisms. However, most existing DRL-based routing algorithms are implemented using traditional Neural Networks (NN), which lack robustness against topology changes. In this paper, we propose a novel algorithm called GAPPO, which integrates Graph Attention Network (GAT) and Proximal Policy Optimization (PPO) to optimize routing policies with the objective of minimizing end-to-end (E2E) latency in networks affected by link failures. To evaluate the performance of the proposed algorithm, we conduct a series of experiments on dynamically changing topologies with link failures under different loads. Experimental results demonstrate that GAPPO outperforms benchmark algorithms in both simulated and real-world networks, confirming its powerful robustness against link failures. Yang Xiao 0013, Sixu Liu, Xucong Lu, Fang Liu 0026, Jun Liu 0014 |
PIMRC | 2 |
| 2023 | Multi-Agent Deep Reinforcement Learning Based Resource Allocation for Ultra-Reliable Low-Latency Internet of Controllable ThingsabstractAs a promising technology in the 5G era, the artificial intelligence (AI) enabled Internet of controllable things (IoCT) is expected to be an integral part of heterogeneous networks (HetNets) in the future. However, the realization of ultra-reliable low-latency communications (URLLC) in IoCT communications underlaid HetNet has stringent quality of service (QoS) requirements, resulting in unprecedented challenges for existing wireless resource allocation methods. In this paper, we first describe a cellular HetNet model with uplink IoCT communications, then formulate a dynamic mixed-integer nonlinear programming (MINLP) resource allocation problem for maximizing the long-term average energy efficiency under URLLC requirements including reliability, latency, and transmission rate. To solve the problem, we propose a decentralized MADRL-based resource allocation algorithm with a decentralized partially observable Markov decision process (dec-POMDP) and a mixed-centralized-decentralized (MCD) framework to address the partial observability and the scalability issues, respectively. In addition, we design a reward function featuring the objective decomposition, baseline-guided scaling, and QoS violation penalty so that the agents are coordinated. Extensive experiments demonstrate the convergence, scalability, and robustness of the proposed algorithm. Besides, the proposed algorithm substantially outperforms conventional resource allocation methods and different agent communication mechanisms in terms of maximizing energy efficiency. Yang Xiao 0013, Yuqian Song, Jun Liu 0014 |
IEEE Trans. Wirel. Commun. | 1 |
| 2022 | Towards Energy Efficient Resource Allocation: When Green Mobile Edge Computing Meets Multi-Agent Deep Reinforcement LearningabstractMobile edge computing (MEC) extends the computing power to the edge of communication networks, which has been considered as a promising technology to further improve the quality of communication services in the near future. Nevertheless, the issue of MEC-empowered energy efficient resource allocation has not been well studied. To maximize the longterm energy efficiency for green MEC-enabled heterogeneous networks (HetNets), we proposed a decentralized multi-agent deep reinforcement learning (MADRL) resource allocation algorithm. Based on the proximal policy optimization (PPO) framework, our proposed algorithm enables observation exchange to coordinate the policies of multiple agents. Simulation results show that our proposed algorithm significantly outperforms three baseline methods in terms of effectiveness, robustness, and scalability. Yang Xiao 0013, Yuqian Song, Jun Liu 0014 |
ICC | 1 |
| 2022 | Attentive Dual-Head Spatial-Temporal Generative Adversarial Networks for Crowd Flow GenerationabstractCrowd flow generation, which aims to simulate the flows of crowd in the future is of great importance to many real-life applications including epidemic spreading and traffic management. The challenges of accurate crowd flow generation come from both the complex spatial-temporal correlations of crowd flow data and the small amount of training data, which is largely not well studied and addressed in existing works. In this paper, we propose a novel attentive dual-head spatial-temporal generative adversarial network entitled ADST-GAN to simulate multi-step crowd flow. To augment the learning power of the generator, we adopt attentive temporal queue and self attention mechanism to automatically capture the complex global spatial-temporal dependencies of crowd flow data. As for the discriminator, we design a dual-head architecture with two-objective training to avoid the negative effects of quick overfitting. To evaluate the effectiveness of the proposed method, we conduct extensive experiments over the bike and taxi trip datasets in New York. The results demonstrate the proposed method outperforms seven state-of-the-art baselines significantly in terms of the quality of simulated crowd flow data. Jianxue Li, Yang Xiao 0013, Jiawei Wu 0004, Yaozhi Chen, Jun Liu 0014 |
PIMRC | 2 |
| 2022 | Deep Reinforcement Learning Enabled Energy-Efficient Resource Allocation in Energy Harvesting Aided V2X CommunicationabstractWith the commercialization of the 5th generation mobile networks, vehicle-to-everything (V2X) communication has gained tremendous attention over the last decade. However, prevailing research has not sufficiently deliberated on the energy efficiency (EE) optimization issue. This paper proposes a decentralized multi-agent deep reinforcement learning (DRL) based resource allocation algorithm. Moreover, we leverage energy harvesting (EH) to achieve long-term EE maximization. Based on the proximal policy optimization (PPO) framework, we invoke power splitting (PS) to divide the harvested energy delicately. Numerical results demonstrate that our proposed algorithm outperforms traditional and straightforward DRL-based resource allocation approaches in effectiveness and robustness. Yuqian Song, Yang Xiao 0013, Yaozhi Chen, Jun Liu 0014 |
PIMRC | 2 |
| 2022 | Deep Reinforcement Learning Based Beamforming for Throughput Maximization in Ultra-Dense NetworksabstractUltra-dense network (UDN) is a promising technology for 5G and beyond communication systems to meet the requirements of explosive data traffic. However, the dense distribution of wireless terminals potentially leads to severe interference and deteriorate network performance. To address this issue, beamforming is widely used to coordinate the interference in UDNs and improve receive gains by controlling the phase of multiple antennas. In this paper, we propose a multi-agent deep reinforcement learning (DRL) based beamforming algorithm to achieve more dynamic and fast beamforming adjustment. In the proposed algorithm, the agents inside beamforming controllers are distributively trained while exchanging partial channel state information (CSI) for better optimizing beamforming vectors to achieve maximized throughputs in UDNs. The evaluation results demonstrate that the proposed algorithm significantly improves the computation efficiency, as well as achieves the highest network throughput compared to several baselines. Huihan Yu, Yang Xiao 0013, Jiawei Wu 0004, Fang Liu 0026, Jun Liu 0014 |
WCNC | 2 |
| 2021 | Power Allocation for Device-to-Multi-Device Enabled HetNets: A Deep Reinforcement Learning ApproachabstractDevice-to-device ($D$2D) communication exploits the geographical proximity by allowing neighboring devices to di-rectly communicate with each other, which becomes one of the most promising technologies to improve the spectral and energy efficiency for 5G and beyond communication systems. To further improve the spectral efficiency and generalize ap-plication scenarios, the emerging device-to-multi-device (D2M$D$) communication enables the D2D transmitter to communicate with multiple receivers simultaneously. In this paper, we consider a heterogeneous network (HetNet) where multiple D2M$D$clusters coexist with the base station (BS) and cellular users (CUs). All D2MD clusters share the same downlink channel as the cellular network, which potentially leads to severe co-channel interfer-ence. To solve this problem, we leverage the deep reinforcement learning (DRL) and propose the deep reinforcement power allocation (DRPA) algorithm to dynamically allocate power for D2MD communication in HetNets. In addition, we apply the centralized training distributed execution (CTDE) technique to accelerate the training process and improve the robustness of DRPA. Simulation results demonstrate that the DRPA algorithm outperforms baseline methods in terms of maximizing the average sum-rate. In addition, the DRPA algorithm is robust to the changes of network environment while achieving near-optimal performance. Yang Xiao 0013, Jiawei Wu 0004, Jun Liu 0014 |
GLOBECOM | 1 |
| 2020 | Privacy preserving distributed data mining based on secure multi-party computation
Jun Liu 0014, Yang Xiao 0013, Nirwan Ansari |
Comput. Commun. | 4 |