VLDB 2026 Research / reviewers in the wild / expert
Hanzhi Yu
dblp:370/2014
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2025
0009-0002-4985-400XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 5 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Contrastive Language-Image Pre-Training Model-based Semantic Communication Performance OptimizationabstractIn this paper, a novel contrastive language–image pre-training (CLIP) model based on semantic The communication framework is designed. Compared to a standard neural network (e.g., convolutional neural network) based semantic encoders and decoders that require joint training over a common dataset, Our CLIP model-based method does not require any training procedures, thus enabling a transmitter to extract data meanings of the original data without neural network model training, and the receiver to train a neural network for follow-up task implementation without the communications with the transmitter. Next, we investigate the deployment of the CLIP model-based semantic framework over a noisy wireless network. Since the semantic information generated by the CLIP model is susceptible to wireless noise and the spectrum used for semantic information transmission are limited; it is necessary to optimize CLIP jointly model architecture and spectrum resource block (RB) allocation to maximize semantic communication performance while considering wireless noise, the delay and energy used for semantic communication. To achieve this goal, we use a proximal policy optimization (PPO) based reinforcement learning (RL) algorithm to learn how wireless noise affects the semantic communication performance, thus finding optimal CLIP model and RB for each user. Simulation results show that our proposed method improves the convergence rate by up to 40%, and the accumulated reward by 4x compared to soft actor-critic. Shaoran Yang, Dongyu Wei, Hanzhi Yu, Zhaohui Yang 0001, Yuchen Liu 0001, Mingzhe Chen |
GLOBECOM | 3 |
| 2025 | Optimizing Training of Policies on Hierarchical Multi-fidelity EnvironmentsabstractIn this paper, we investigate a novel digital network twin (DNT) assisted deep learning (DL) model training framework that enables a base station (BS) to select the data collection source from both physical network and DNT for training DL models used to optimize network performance. In particular, we consider a DNT enabled cellular network that consists of a physical network where a BS uses several antennas to serve multiple mobile users, and a DNT that is a virtual representation of the physical network. The BS must adjust the antenna tilt angles to optimize the data rates of all users. Due to the user mobility, the BS may not be able to accurately track the network dynamics. Hence, a reinforcement learning (RL) approach is used to dynamically adjust the antenna tilt angles. To train the RL, we can use data collected from the physical network and the DNT. The data collected from the physical network is more accurate but incurs more communication overhead compared to the data collected from the DNT. Therefore, it is necessary to determine the ratio of data collected from the physical network and the DNT to improve the training of the RL model, so as to optimize the tilt angle of each antenna in response to user mobility. We formulate an optimization problem to jointly optimize the tilt angle adjustment policy and the data collection strategy, aiming to maximize the data rates of all users while constraining the time delay introduced by collecting data from the physical network. To solve this optimization problem, we propose a hierarchical RL framework consisting of a two-level proximal policy optimization (PPO). The first level PPO dynamically adjusts the antenna tilt angles, and the second level PPO determines the ratio of data collected from the physical network to improve the first level PPO training performance. Compared to traditional single RL algorithms such as deep Q network (DQN), the designed method optimizes the data collection ratio and the antenna tilt angles at diverse time intervals, allowing the second level PPO to adjust the data collection ratio with a large timescale using the training information provided by the first level PPO, and allowing the first level PPO focuses on adjusting the antenna tilt angles with a small timescale. Simulation results show that our proposed method reduces the physical network data collection delay by up to 71.85% compared to a hierarchical RL that uses a deep deterministic policy gradient algorithm as the second level RL. Hanzhi Yu, Hasan Farooq, Julien Forgeat, Shruti Bothe, Kristijonas Cyras, Md Moin Uddin Chowdhury, Mingzhe Chen |
GLOBECOM | 1 |
| 2025 | Joint Optimization of Communication and Device Clustering for Secure Clustered Federated LearningabstractIn this paper, a secure and communication-efficient clustered federated learning (CFL) design is investigated. In our model, several base stations (BSs) with heterogeneous task-handling capabilities and multiple users with non-independent and identically distributed (non-IID) data jointly perform CFL training using differential privacy (DP) techniques. Since each BS can process only a subset of learning tasks and has limited wireless resource blocks to allocate to users for federated learning (FL) model parameter transmission, it is necessary to jointly optimize resource block (RB) allocation and user scheduling for CFL performance optimization. Meanwhile, our considered CFL requires devices to use their limited data and FL model information to determine their task identities, which may introduce additional communication overhead. This problem is formulated as an optimization problem whose goal is to minimize the training loss of all learning tasks while considering device clustering, RB allocation, noise, and FL model transmission delay. To solve this, we propose a novel value decomposed multi-agent reinforcement learning (VD-MARL) algorithm that enables distributed BSs to independently determine their connected users, the RBs, and DP noise of the connected users but jointly minimize the training loss of all learning tasks across all BSs. Different from the existing MARL methods that assign a large penalty for invalid actions, we propose a novel penalty assignment scheme that assigns penalty depending on the number of devices that cannot meet communication constraints (e.g., delay), which can guide the MARL scheme to quickly find valid actions thus improving the convergence speed. Simulation results show that the VD-MARL can improve the convergence rate by up to 35% and the ultimate accumulated rewards by 27% compared to independent Q-learning. Dongyu Wei, Hanzhi Yu, Yuchen Liu 0001, Shiwen Mao, Mingzhe Chen |
ICC | 2 |
| 2025 | Fluid Antenna System (FAS)-Assisted 3D UAV Positioning Performance OptimizationabstractIn this paper, the framework of fluid antenna system (FAS)-assisted three dimensional (3D) passive unmanned aerial vehicle (UAV) positioning is developed. In the proposed framework, a set of controlled UAVs including an active UAV and four FAS-assisted passive UAVs, as well as a ground base station (BS) cooperatively estimate the real-time 3D position of a target UAV. Here, the active UAV transmits a measurement signal to the passive UAVs. This signal is reflected via the target UAV and received by the passive UAVs. Each passive UAV estimates the distance of the active-target-passive UAV link and selects an antenna port to share the distance information with the BS. The BS calculates the real-time position of the target UAV. As the target UAV is moving due to its task operation, the controlled UAVs must optimize their trajectories and select optimal antenna port for transmitting the positioning information, aiming to estimate the real-time position of the target UAV. We formulate an optimization problem that optimizes the trajectories of all controlled UAVs and antenna port selection of passive UAVs with the aim of minimizing the target UAV positioning error. To address this problem, an attention-based recurrent multiagent reinforcement learning (AR-MARL) scheme is proposed. In the proposed method, a recurrent neural network (RNN) acts as a local Q function of each controlled UAV to capture its historical state-action pairs, and a transformer is used to analyze the importance of these historical state-action pairs, thus improving the global$\mathbf{Q}$function approximation accuracy, thereby further improving the positioning accuracy. Simulation results show that the proposed AR-MARL scheme can reduce the average positioning error by up to 17.5 % and 58.5 % compared to the VD-MARL scheme and the proposed method without FAS. Xiaoren Xu, Hao Xu 0003, Hanzhi Yu, Yuchen Liu 0001, Mingzhe Chen |
ICC | 3 |
| 2025 | Optimizing Wireless Resource Management and Synchronization in Digital Twin NetworksabstractIn this article, we investigate an accurate synchronization between a physical network and its digital network twin (DNT), which serves as a virtual representation of the physical network. The considered network includes a set of base stations (BSs) that must allocate its limited spectrum resources to serve a set of users while also transmitting its partially observed physical network information to a cloud server to generate the DNT. Since the DNT can predict the physical network status based on its historical status, the BSs may not need to send their physical network information at each time slot, allowing them to conserve spectrum resources to serve the users. However, if the DNT does not receive the physical network information of the BSs over a large time period, the DNT’s accuracy in representing the physical network may degrade. To this end, each BS must decide when to send the physical network information to the cloud server to update the DNT, while also determining the spectrum resource allocation policy for both DNT synchronization and serving the users. We formulate this resource allocation task as an optimization problem, aiming to maximize the total data rate of all users while minimizing the asynchronization between the physical network and the DNT. The formulated problem is challenging to solve by traditional optimization methods, as each BS can only observe a partial physical network, making it difficult to find an optimal spectrum allocation strategy for the entire network. To address this problem, we propose a method based on the gated recurrent units (GRUs) and the value decomposition network (VDN). The GRU component allows the DNT to predict future status using the historical data, effectively updating itself when the BSs do not transmit the physical network information. The VDN algorithm enables each BS to learn the relationship between its local observation and the team reward of all BSs, allowing it to collaborate with others in determining whether to transmit physical network information and optimizing spectrum allocation. Simulation results show that our GRU-based and VDN-based algorithm improves the weighted sum of data rates and the similarity between the status of the DNT and the physical network by up to 28.96%, compared to a baseline method combining GRU with the independent Q learning (IQL). Hanzhi Yu, Yuchen Liu 0001, Zhaohui Yang 0001, Haijian Sun, Mingzhe Chen |
IEEE Internet Things J. | 1 |
| 2024 | Joint Communication and Synchronization Performance Optimization in Digital Twin Enabled NetworksabstractIn this paper, we investigate an accurate synchronization between a physical network and its digital network twin (DNT) that is a virtual representation of the physical network. The considered network includes a physical network where a base station (BS) serves a set of users, and a DNT that evolves with the status of both DNT and the physical network. The BS must use its limited spectrum resources to serve the users, as well as transmit the physical network information to the cloud server for DNT synchronization. Since the DNT can predict the physical network status, the BS may not need to transmit physical network information to the server at each time slot thus saving spectrum resources to serve users. However, if the BS does not transmit physical information to the DNT over a long period of time, the DNT may not be able to represent the physical network accurately. To this end, the BS must determine whether to send physical network information to the server to update DNT and the spectrum resources used for physical network information transmission and serving users. We formulate this resources allocation problem as an optimization problem aiming to maximize the sum of data rates of all users, while minimizing the gap between the states of the physical network and the DNT. The formulated problem is challenging to solve by conventional optimization methods, since the BS may not be able to know the future status of the DNT. To solve this problem, we design a gate recurrent unit (GRU) and soft action-critic (SAC) based algorithm. The GRU enables the DNT to predict its future states by using historical state data, and updating the DNT when the BS does not transmit physical network information. The SAC based algorithm enables the BS to learn the relationship between the physical network information transmission and the future status estimation accuracy of the DNT thus determining whether to transmit physical network information to the cloud server, ensuring an accuracte synchronization between the physical network and the DNT. Simulation results demonstrate that our designed algorithm can promote the weighted sum of data rates and the similarity between the status of the DNT and the physical network by up to 10.31% compared to a baseline method integrating the GRU and the deep Q network. Hanzhi Yu, Yuchen Liu 0001, Mingzhe Chen |
GLOBECOM | 1 |
| 2024 | Complex-Valued Neural-Network-Based Federated Learning for Multiuser Indoor Positioning Performance OptimizationabstractIn this article, the use of channel state information (CSI) for indoor positioning is studied. In the considered model, a server equipped with several antennas sends pilot signals to users, while each user uses the received pilot signals to estimate channel states for user positioning. To this end, we formulate the positioning problem as an optimization problem aiming to minimize the gap between the estimated positions and the ground truth positions of users. To solve this problem, we design a complex-valued neural network (CVNN) model based federated learning (FL) algorithm. Compared to standard real-valued centralized machine learning (ML) methods, our proposed algorithm has two main advantages. First, our proposed algorithm can directly process complex-valued CSI data without data transformation. Second, our proposed algorithm is a distributed ML method that does not require users to send their CSI data to the server. Since the output of our proposed algorithm is complex-valued which consists of the real and imaginary parts, we study the use of the CVNN to implement two learning tasks. First, the proposed algorithm directly outputs the estimated positions of a user. Here, the real and imaginary parts of an output neuron represent the 2D coordinates of the user. Second, the proposed method can output two CSI features (i.e., line-of-sight/non-line-of-sight transmission link classification and time of arrival (TOA) prediction) which can be used in traditional positioning algorithms. Simulation results demonstrate that our designed CVNN based FL can reduce the mean positioning error between the estimated position and the actual position by up to 36%, compared to a RVNN based FL which requires to transform CSI data into real-valued data. Hanzhi Yu, Yuchen Liu 0001, Mingzhe Chen |
IEEE Internet Things J. | 1 |
| 2023 | Complex Neural Networks for Indoor Positioning with Complex-Valued Channel State InformationabstractIn this paper, the use of channel state information (CSI) for indoor positioning is investigated. In the considered model, a base station (BS) equipped with several antennas sends pilot signals to a user that transmits the received pilot signals back to the BS. The BS will use the received CSI data to estimate the position of the user. To this end, we formulate this positioning problem as an optimization problem aiming to minimize the mean square error between the estimated position and the actual position of the user. To solve this problem, we design a complex-valued neural network (CVNN) based positioning algorithm. Compared to real-valued neural networks (RVNNs) that need to convert complex-valued CSI data into real-valued data, the proposed method uses original CSI data to train the CVNN model for user positioning. Since the output of our proposed algorithm is complex-valued and it consists of the real and imaginary parts, we can use it to implement two learning tasks. Based on this property, two use cases of the proposed algorithm are proposed: 1) the algorithm directly outputs the estimated position of the user. Here, the real and imaginary parts of an output neuron represent the 2D coordinates of the user, 2) the algorithm outputs two CSI features (i.e., line-of-sight/non-line-of-sight transmission link classification and time of arrival (TOA) prediction) which can be used in traditional positioning algorithms. Simulation results demonstrate that our designed CVNN based algorithm can reduce the mean positioning error between the estimated position and the actual position by up to 11.1%, compared to a RVNN based method which has to transform CSI data into real-valued data. Hanzhi Yu, Mingzhe Chen, Zhaohui Yang 0001, Yuchen Liu 0001 |
GLOBECOM | 1 |