EDBT 2026 Demo / reviewers in the wild / expert
Yang Cao 0018
dblp:25/7045-18
· DBLP profile ↗
19ranked-venue papers
11as first author
17since 2021 · last 2026
0000-0002-8061-4066ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 15 · 10 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | iRadioDiff: Physics-Informed Diffusion Model for Indoor Radio Map Construction and Localization
Xiucheng Wang, Tingwei Yuan, Yang Cao 0018, Nan Cheng 0001, Ruijin Sun, Weihua Zhuang |
ICC | 3 |
| 2026 | Foundation Model-Based Mobility Management for 6G Mobile NetworksabstractMobility management associating all user equipments (UEs) to proper base stations (BSs) (also known as handover, HO) to achieve the designated performance optimization is one of the most crucial functions in mobile networks. Although effective mobility management has received considerable research attentions, existing schemes follow event-driving operations, in which HO decisions are made based on events of performance degradation that BSs/UEs passively suffer from or proactively foresee. However, the performance and decisions of these schemes are highly subject to the identified events, to lose generalization and better performance under unidentified events. To address this issue, we propose a foundation model (FM) based mobility management for the sixth generation (6G) mobile networks inherently supporting artificial intelligence (AI) computing, in which the FM generates the HO decisions for all UEs to maximize the overall throughput under the constraints of ping-pong rate and HO failure (HOF) rate by implicitly taking moving trajectories, traffic demands, channel conditions of all UEs and available resources of BSs into account. To this end, a hierarchical model structure composed of Long Short-Term Memory (LSTM) networks with multi-head attention (MHA) is adopted, which is trained by emulated datasets with augmentation. The performance evaluation results show that the proposed scheme outperforms the state-of-the-art schemes in terms of the average throughput over all UEs while satisfying the required ping-pong rate and HOF rate, and justify the robustness of the proposed FM under different network deployment scenarios. Shao-Yu Lien, Yu-Han Huang, Chih-Cheng Tseng, Yang Cao 0018, Hui-Hsin Chin, Der-Jiunn Deng |
IEEE Internet Things J. | 4 |
| 2025 | Resilient Routing for Satellite-Terrestrial Integrated Networks via Cascaded Two-Time-Scale Deep Reinforcement Learning
Yang Cao 0018, Shao-Yu Lien, Ying-Chang Liang |
GLOBECOM | 1 |
| 2025 | Prompt-guided Semantic Communication for Image Transmission with Dynamic CompressionabstractSemantic communication (SemCom) has emerged as a key paradigm for next-generation wireless networks. However, existing SemCom systems face challenges in adapting to dynamic Semantic Entities of Interest (SEI) and suffer from significant computational overhead. This paper proposes a novel Prompt-guided SemCom (PSC) framework that leverages textual prompts to dynamically align transmitted image features with the receiver’s SEI, enabling adaptive compression and enhanced generalization. The proposed PSC system introduces two key features. Specifically, a cross-modal policy network dynamically selects image tokens based on textual prompts, thereby minimizing the transmission overhead. Moreover, a Mixture of Experts (MoE) architecture enables parallel processing with sparse expert activation, efficiently handling heterogeneous features while reducing computational complexity during inference. To train the PSC system, we introduce an iterative algorithm combining self-supervised learning and Group Relative Policy Optimization (GRPO), unifying Deep Learning (DL) and Deep Reinforcement Learning (DRL) for the first time in the context of SEI extraction and selection. Experiments demonstrate the superiority of PSC over conventional methods, achieving enhanced reconstruction performance and robustness. Furthermore, results validate the effectiveness of the policy network and MoE architecture, enabling PSC to achieve 94.80% reconstruction performance while reducing communication overhead by 44.90% and the number of activated parameters by 47.04%. Yang Cao 0018, Ying-Chang Liang |
GLOBECOM | 2 |
| 2025 | Off-grid DOA Estimation for Disturbed UAV Swarm with Nested ArrayabstractUnmanned aerial vehicles (UAVs), with the advantage in on-demand and cost-effective deployment, have emerged as an aerial communication platform for supporting constant high-data rate transmission services for terrestrial stations with different mobility patterns. To this end, developing high-accuracy localization functionality is of crucial importance in tracking the target users. However, the limited payload size of a single UAV restricts precise direction-of-arrival (DOA) estimation, and the UAV-swarm based collaborative estimation architectures have become a promising remedy. In this paper, the UAV swarm is utilized to form a nested array (NA) architecture to improve the DOA estimation accuracy. A two-stage alternating iterative approach is proposed to jointly estimate the UAV position errors and DOAs through iterative feedback. Specifically, position error estimation is performed using the gradient descent method, whereas DOA estimation is achieved through a novel NA-Bessel atomic transformation and norm minimization (NA-BATNM) framework, built upon an off-grid atomic norm minimization algorithm. Simulation results demonstrate that the proposed method achieves high-precision estimation of both DOAs and UAV position errors. Furthermore, the proposed NA-BATNM method exhibits superior performance compared to grid-based DOA estimation algorithms. Jiaxuan Gao, Yanyan Wang 0009, Yang Cao 0018, Xiaohu Tang 0004 |
VTC2025-Fall | 3 |
| 2025 | Designs and Prototypes of RICs for Policy Management in O-RANabstractToward intelligent computations in the fifth generation (5G) mobile networks, the Open Radio Access Network (O-RAN) Alliance has introduced the O-RAN architecture, which enables two unprecedented platforms: Non-Real-Time (Non-RT) RAN Intelligent Controller (RIC) and Near-Real-Time (Near-RT) RIC connected via the A1 interface. Although the standards for the functions and procedures of the A1 interface has been provided by the O-RAN Alliance, the practical implementation of A1 still suffers significant challenges due to 1) incompleteness of existing standardization procedures, 2) lack of mechanisms to ensure the integrity of procedures, and 3) lack capability to support large data storage. To address these challenges, this paper therefore provides the design and implementation of the A1 Application Protocol (A1AP) with particular focus on the A1 Policy Management Service (A1-P). To support all the functions and procedures of A1-P, this paper provides the designs and implementations of the A1 Policy Management Service Component (A1 PMS) in the Non-RT RIC, and the A1 Mediator in the Near-RT RIC, which enable not only all the A1-P procedures, but also the management of policy types, policies, policy statuses, and association of the rAPPs in the Non-RT RIC and the xAPPs in the Near-RT RIC. To ensure the integrity of A1-P, we propose additional steps in existing procedures to handle abnormal cases. To further support large data storage particularly needed for intelligent computation, all the A1-P procedures are redesigned to support InfluxDB (a database allowing data storage on hard disk drive). Through practically establishing the Non-RT RIC and Near-RT RIC platforms connected with the proposed A1-P design, and connecting the Near-RT RIC with the RAN emulator through the E2 interface, we demonstrate the practicability of our design to support the use case of energy saving (ES). Yu-Han Qiu, Shao-Yu Lien, Chih-Cheng Tseng, Yang Cao 0018 |
VTC2025-Fall | 4 |
| 2025 | Handover Mechanism based on Link Quality and Longevity for LEO-based Non-Terrestrial NetworksabstractLow Earth Orbit (LEO) satellites in the nonterrestrial networks (NTNs) are challenged by high satellite moving speeds and dynamic link conditions, which lead to frequent and suboptimal handovers. A handover mechanism based on a composite quality indicator (CQI) that balances the trade-off between the quality and longevity of a link is proposed. To capture the nonlinear behavior of the links between the LEO satellites and the user terminal (UT), two logistic functions (LFs) are employed to dynamically adjust the weight to generate the CQI and the value of the handover threshold, respectively. By using the Two-Line Element (TLE) data, a genetic algorithm (GA) is developed to find the optimal parameter values to facilitate the two LFs by searching a multidimensional solution space. Simulation results demonstrate that the proposed handover mechanism not only enhances link quality and longevity but also reduces the number of handovers compared to the elevation-based and max remaining visibility time (RVT)-based methods. Chien-Lin Yen, Chih-Cheng Tseng, Shao-Yu Lien, Fang-Chang Kuo, Yao-Jen Liang, Yang Cao 0018 |
VTC2025-Fall | 6 |
| 2025 | Integrated Distributed Semantic Communication and Over-the-Air Computation for Cooperative Spectrum SensingabstractCooperative spectrum sensing (CSS) is a promising approach to improve the detection of primary users (PUs) using multiple sensors. However, there are several challenges for existing combination methods, i.e., performance degradation and ceiling effect for hard-decision fusion (HDF), as well as significant uploading latency and non-robustness to noise in the reporting channel for soft-data fusion (SDF). To address these issues, an integrated communication and computation (ICC) framework is proposed in this paper. Specifically, distributed semantic communication (DSC) jointly optimizes multiple sensors and the fusion center to minimize the transmitted data without degrading detection performance. Moreover, over-the-air computation (AirComp) is utilized to further reduce spectrum occupation in reporting channel, taking advantage of characteristics of wireless channel to enable data aggregation. Under the ICC framework, a particular system, namely ICC-CSS, is designed and implemented, which is theoretically proved to be equivalent to the optimal estimator-correlator (E-C) detector with equal gain SDF when the PU signal samples are independent and identically distributed. Extensive simulations verify the superiority of ICC-CSS compared with various conventional CSS schemes in terms of detection performance, robustness to SNR variations in both sensing and reporting channels, as well as scalability with respect to the number of samples and sensors. Yang Cao 0018, Xin Kang 0001, Ying-Chang Liang |
IEEE Trans. Commun. | 2 |
| 2024 | Learning-Based Multitier Split Computing for Efficient Convergence of Communication and ComputationabstractWith promising benefits of splitting deep neural network (DNN) computation loads to the edge server, split computing has been a novel paradigm achieving high-quality artificial intelligence (AI) services for the energy-constrained user equipments (UEs). To satisfy the service demands of a large number of UEs, traditional edge-UE split computing evolves toward multitier split computing involving the edge and cloud servers with different capabilities, leading to a “complex” optimization involving communication and computation. To tackle this challenge, in this article, we propose a multitier deep reinforcement learning (DRL) decision-making scheme for distributed splitting point selection and computing resource allocation in the three-tier UE-edge-cloud split computing systems. With the proposed scheme, the high-dimensional optimization can be tackled by the UEs and an edge server with different control cycles through performing local decision-making tasks in a sequential manner. Based on the policies updated by the UEs and the edge server in successive stages, the overall performance of split computing can be continuously improved, which is justified through a theoretical convergence performance analysis. Comprehensive simulation studies show that the proposed multitier DRL decision-making scheme outperforms the conventional split computing schemes in terms of the overall latency, inference accuracy, and energy efficiency to practice multitier split computing. Yang Cao 0018, Shao-Yu Lien, Cheng-Hao Yeh, Der-Jiunn Deng, Ying-Chang Liang, Dusit Niyato |
IEEE Internet Things J. | 1 |
| 2024 | Collaborative Computing in Non-Terrestrial Networks: A Multi-Time-Scale Deep Reinforcement Learning ApproachabstractConstructing earth-fixed cells with low-earth orbit (LEO) satellites in non-terrestrial networks (NTNs) has been the most promising paradigm to enable global coverage. The limited computing capabilities on LEO satellites however render tackling resource optimization within a short duration a critical challenge. Although the sufficient computing capabilities of the ground infrastructures can be utilized to assist the LEO satellite, different time-scale control cycles and coupling decisions between the space- and ground-segments still obstruct the joint optimization design for computing agents at different segments. To address the above challenges, in this paper, a multi-time-scale deep reinforcement learning (DRL) scheme is developed for achieving the radio resource optimization in NTNs, in which the LEO satellite and user equipment (UE) collaborate with each other to perform individual decision-making tasks with different control cycles. Specifically, the UE updates its policy toward improving value functions of both the satellite and UE, while the LEO satellite only performs finite-step rollout for decision-makings based on the reference decision trajectory provided by the UE. Most importantly, rigorous analysis to guarantee the performance convergence of the proposed scheme is provided. Comprehensive simulations are conducted to justify the effectiveness of the proposed scheme in balancing the transmission performance and computational complexity. Yang Cao 0018, Shao-Yu Lien, Ying-Chang Liang, Dusit Niyato, Xuemin Shen |
IEEE Trans. Wirel. Commun. | 1 |
| 2024 | Deep Learning-Empowered Semantic Communication Systems With a Shared Knowledge BaseabstractDeep learning-empowered semantic communication is regarded as a promising candidate for future 6G networks. Although existing semantic communication systems have achieved superior performance compared to traditional methods, the end-to-end architecture adopted by most semantic communication systems is regarded as a black box, leading to the lack of explainability. To tackle this issue, in this paper, a novel semantic communication system with a shared knowledge base is proposed for text transmissions. Specifically, a textual knowledge base constructed by inherently readable sentences is introduced into our system. With the aid of the shared knowledge base, the proposed system integrates the message and corresponding knowledge from the shared knowledge base to obtain the residual information, which enables the system to transmit fewer symbols without semantic performance degradation. In order to make the proposed system more reliable, the semantic self-information and the source entropy are mathematically defined based on the knowledge base. Furthermore, the knowledge base construction algorithm is developed based on a similarity-comparison method, in which a pre-configured threshold can be leveraged to control the size of the knowledge base. Moreover, the simulation results have demonstrated that the proposed approach outperforms existing baseline methods in terms of transmitted data size and sentence similarity. Yang Cao 0018, Xin Kang 0001, Ying-Chang Liang |
IEEE Trans. Wirel. Commun. | 2 |
| 2023 | Efficient Communication-Computation Tradeoff for Split Computing: A Multi-Tier Deep Reinforcement Learning ApproachabstractSplitting the computation loads of a neural network (NN) training task to multiple stations, split computing has been the most promising technology to sustain high-accuracy model for resource-constrained user equipments (UEs) to empower real-time intelligent services. Nevertheless, different communication link variations and computation capabilities in different stations (including UE and servers) render the overall performance optimization in split computing a critical challenge. In this case, different stations should be able to infer the others' communication/computation capabilities to distributively decide the optimum splitting points of an NN. To this end, in this paper, we propose a multi-tier deep reinforcement learning (DRL) scheme for split computing, by which the UE and edge server can collaboratively and adaptively determine their splitting points and computation resources to optimize the long-term overall training latency through tackling different time-scale sub-optimizations in a sequential manner. With the image recognition task as experimental example, comprehensive simulations are conducted to justify the performances in terms of training latency, model accuracy and energy consumption of the proposed scheme for split computing. Yang Cao 0018, Shao-Yu Lien, Cheng-Hao Yeh, Ying-Chang Liang, Dusit Niyato |
GLOBECOM | 1 |
| 2023 | Collaborative Deep Reinforcement Learning for Resource Optimization in Non-Terrestrial NetworksabstractNon-terrestrial networks (NTNs) with low-earth orbit (LEO) satellites have been regarded as promising remedies to support global ubiquitous wireless services. Due to the rapid mobility of LEO satellite, inter-beam/satellite handovers happen frequently for a specific user equipment (UE). To tackle this issue, earth-fixed cell scenarios have been under studied, in which the LEO satellite adjusts its beam direction towards a fixed area within its dwell duration, to maintain stable transmission performance for the UE. Therefore, it is required that the LEO satellite performs real-time resource allocation, which however is unaffordable by the LEO satellite with limited computing capability. To address this issue, in this paper, we propose a two-time-scale collaborative deep reinforcement learning (DRL) scheme for beam management and resource allocation in NTNs, in which LEO satellite and UE with different control cycles update their decision-making policies through a sequential manner. Specifically, UE updates its policy subject to improving the value functions of both the agents. Furthermore, the LEO satellite only makes decisions through finite-step rollouts with a reference decision trajectory received from the UE. Simulation results show that the proposed scheme can effectively balance the throughput performance and computational complexity over traditional greedy-searching schemes. Yang Cao 0018, Shao-Yu Lien, Ying-Chang Liang, Dusit Niyato, Xuemin Shen |
PIMRC | 1 |
| 2022 | User Access Control in Open Radio Access Networks: A Federated Deep Reinforcement Learning ApproachabstractTargeting at implementing the next generation radio access networks (RANs) with virtualized network components, the open RAN (O-RAN) has been regarded as a novel paradigm towards fully open, virtualized and interoperable RANs. Through particularly introducing RAN intelligent controllers (RICs), machine learning (ML) can be unprecedentedly installed, adapting to various vertical applications and deployment environments without sophisticated planning efforts. However, the O-RAN also suffers two critical challenges of load balancing and frequent handovers in the massive base station (BS) deployment. In this paper, an intelligent user access control scheme with deep reinforcement learning (DRL) is proposed. To optimize the performance of distributed deep Q-networks (DQNs) trained by user equipments (UEs), a federated DRL-based scheme is proposed with a global model server installed in the RIC to update the DQN parameters. To further predictively train a global DQN with acceptable signaling overheads, the upper confidence bound (UCB) algorithm to select the optimal UE set and a dueling structure to decompose the DQN parameters are developed. With the proposed scheme, each UE effectively maximizes the long-term throughput and avoids frequent handovers. The simulation results well justify the outstanding performance of the proposed scheme over the-state-of-the-arts, to serve as references for the O-RAN standardization. Yang Cao 0018, Shao-Yu Lien, Ying-Chang Liang, Kwang-Cheng Chen, Xuemin Shen |
IEEE Trans. Wirel. Commun. | 1 |
| 2021 | Federated Deep Reinforcement Learning for User Access Control in Open Radio Access NetworksabstractThe Open Radio Access Network (O-RAN) introducing a particular unit known as RAN Intelligent Controllers (RICs) has been regarded as revolutionary paradigms to support multiclass wireless services required in the fifth and sixth generation (5G/6G) networks. Through unprecedentedly installing various machine learning (ML) algorithms to RICs, a RAN is able to intelligently configure resources/communications to support any vertical applications over any operating scenarios. However, to practically deploy this RAN paradigm, the O-RAN still suffers two critical issues of load balance and handover control, and therefore the very first ML algorithm for the O-RAN should effectively address these issues. In this paper, inspired by the superior performance of deep reinforcement learning (DRL) in tackling sequential decision-making tasks, we therefore develop an intelligent user access control scheme with the facilitation of deep Q-networks (DQNs). A federated DRL-based scheme is further proposed to train the parameters of multiple DQNs in the O-RAN, so as to maximize the long-term throughput and meanwhile avoid frequent user handovers with a limited amount of signaling overheads in the O-RAN. The simulation results have fully demonstrated the outstanding performance over the state-of-the-arts, to service the urgent needs in the standardization of the O-RAN. Yang Cao 0018, Shao-Yu Lien, Ying-Chang Liang, Kwang-Cheng Chen |
ICC | 1 |
| 2021 | Multi-tier Collaborative Deep Reinforcement Learning for Non-terrestrial Network Empowered Vehicular ConnectionsabstractWith the objective of supporting next generation driving services, non-terrestrial networks (NTNs) with low earth orbit (LEO) satellites have been regarded as promising paradigms to implement global ubiquitous and high-capacity vehicular connections. However, due to the high moving speed, different satellites can only service a specific set of vehicles for few minutes. In such case, due to the limited computing capability of the satellite, machine learning (ML) based and non-ML based solutions cannot be performed within such a short duration. To address these issues, in this paper, we propose a multi-tier collaborative deep reinforcement learning (DRL) scheme for resource allocation in NTN empowered vehicular networks, in which ground vehicles and LEO satellites maintain DRL-based decision model to obtain resource allocation decisions cooperatively. Specifically, ground vehicles with powerful computing capabilities can assist the satellite to tackle resource allocation optimizations, and the satellite determines final decisions and model parameters by aggregating local calculated results of vehicles. Additionally, the parameters of DRL-based decision model can be transferred from the current satellite to its successor as the starting point for future resource allocation decision-makings. Comprehensive simulations have been conducted to show the effectiveness of our proposed scheme. Yang Cao 0018, Shao-Yu Lien, Ying-Chang Liang |
ICNP | 1 |
| 2021 | Deep Reinforcement Learning For Multi-User Access Control in Non-Terrestrial NetworksabstractNon-Terrestrial Networks (NTNs) composed of space-borne (e.g., satellites) and airborne vehicles (e.g., drones and blimps) have recently been proposed by 3GPP as a new paradigm of infrastructures to enhance the capacity and coverage of existing terrestrial wireless networks. The mobility of non-terrestrial base stations (NT-BSs) however leads to a dynamic environment, which imposes unique challenges for handover and throughput optimization particularly in multi-user access control for NTNs. To achieve performance optimization, each terrestrial user equipment (UE) should autonomously estimate the dynamics of moving NT-BSs, which is different from the existing user access control schemes in terrestrial wireless networks. Consequently, new learning schemes for optimum multi-user access control are desired. In this article, we therefore propose a UE-driven deep reinforcement learning (DRL) based scheme, in which a centralized agent deployed at the backhaul side of NT-BSs is responsible for training the parameter of a deep Q-network (DQN), and each UE independently makes its own access decisions based on the parameter from the trained DQN. With the proposed scheme, each UE is able to access a proper NT-BS intelligently to enhance the long-term system throughput and avoid frequent handovers among NT-BSs. Through comprehensive simulation studies, we justify the performance of the proposed scheme, and show its effectiveness in addressing the fundamental issues in the NTNs deployment. Yang Cao 0018, Shao-Yu Lien, Ying-Chang Liang |
IEEE Trans. Commun. | 1 |
| 2019 | Deep Reinforcement Learning for Channel and Power Allocation in UAV-enabled IoT SystemsabstractUnmanned aerial vehicles (UAVs) have recently been proposed as moving base stations to collect data from ground IoT nodes in remote areas. Since IoT nodes are normally battery-limited, energy efficiency is an important metric in IoT systems. In order to improve energy efficiency in UAV-enabled IoT systems, it is necessary to allocate both channels and transmit power properly for IoT nodes. Motivated by the superior performance of deep reinforcement learning (DRL) in decision-making tasks, we propose a DRL-based channel and power allocation framework in a UAV-enabled IoT system. With the proposed framework, the UAV-BS is able to intelligently allocate both channels and transmit power for uplink transmissions of IoT nodes to maximize the minimum energy-efficiency among all the IoT nodes. Simulation results validate the effectiveness of the proposed algorithm and show its superiority over the- state-of-the-arts. Yang Cao 0018, Lin Zhang 0022, Ying-Chang Liang |
GLOBECOM | 1 |
| 2019 | Deep Reinforcement Learning for Multi-User Access Control in UAV NetworksabstractUnmanned Aerial Vehicles (UAVs) have recently been proposed as flying base stations, called UAV-BSs, to provide reliable connections and extend the coverage of the existing wireless networks. The mobility of UAV-BSs leads to a dynamic network environment, in which the global network information is hard to be obtained. Since frequent information exchanges cause huge signaling overheads, it is difficult to deploy centralized algorithms in UAV networks. Hence, we propose a distributed deep reinforcement learning (DRL) framework for multi-user access control in UAV networks. In particular, each user makes its own access decisions independently based on the local network information, and maximizes the long-term throughput while avoiding frequent handovers. Simulation results have validated the effectiveness of the proposed algorithm and shown the superiority of the proposed DRL framework over the state of arts. Yang Cao 0018, Lin Zhang 0022, Ying-Chang Liang |
ICC | 1 |