VLDB 2026 Research / reviewers in the wild / expert
Tong Liu 0035
dblp:36/5558-35
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0001-7537-7070ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 4 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Event-Triggered CIO Optimization for Cellular-Connected UAVs via Deep Reinforcement LearningabstractCellular-connected unmanned aerial vehicles (UAVs) are increasingly being deployed in emerging Internet of Things (IoT) applications, where reliable handover management is critical to ensure uninterrupted communication. Unlike terrestrial users, UAVs face frequent handovers due to antenna side-lobe coverage and fragmented aerial cell overlaps. Existing optimization strategies, however, are often limited to coarse-grained “whether-to-handover” decisions or unidirectional cell individual offset (CIO) tuning, resulting in degraded handover performance. To address these problems, this paper proposes an Event-triggered Bidirectional CIO Optimization with Hybrid prioritized experience replay based Deep Reinforcement Learning (EBCO-HDRL) algorithm. Under the A3 event-triggered mechanism, EBCO-HDRL jointly configures serving-to-neighbor and neighbor-to-serving CIOs, improving handover accuracy and reducing ping-pong events. An event-triggered optimization strategy adaptively reconfigures CIOs only when needed, mitigating computational overhead and enhancing training stability. To further improve sample efficiency, a Hybrid Prioritized Experience Replay (HPER) scheme is introduced, combining temporal-difference error and high-reward sampling, supported by a heterogeneous neural network design. Simulation results show that EBCO-HDRL significantly outperforms baseline algorithms in terms of precision, stability, and robustness, offering an effective solution for UAV handover management in cellular IoT networks. Tong Liu 0035, Yimeng Shang, Nan Hu 0010, Lijun Dong, Wenying Yang, Wenzhi Li |
IEEE Internet Things J. | 1 |
| 2024 | FL2ETD: A Few-Shot Learning Framework to Electricity Theft DetectionabstractElectricity theft detection (ETD) aims to promptly identify electricity theft by vigilantly monitoring and analyzing atypical electricity consumption time series. Existing machine learning approaches to ETD demand large training sets, leading to degraded performance when limited training samples are available. In this paper, we introduce FL2ETD, a novel few-shot learning framework to ETD. The framework consists of three core components, i.e., a feature extraction module, a representation module, and a classification module. The feature extraction module processes the electricity consumption behavior of users in both the time and the frequency domains to extract distinctive features and increase the number and the diversity of features. The representation module utilizes contrast learning to pre-train unlabeled electricity consumption data for enhancing feature representation quality. The classification module integrates feature representations for making the final decision in ETD. Extensive experiments demonstrate that FL2ETD exhibits superior performance compared to baselines, and its advantage is significant when the number of available training samples is very small (with only 338 samples). Chenying Meng, Feng Lyu 0001, Jie Gao 0002, Tong Liu 0035, Xuemin Shen |
ICC | 4 |
| 2024 | Multitimescale Control and Communications With Deep Reinforcement Learning - Part II: Control-Aware Radio Resource AllocationabstractIn Part I of this two-part paper (Multitimescale Control and Communications with deep reinforcement learning (DRL)—Part I: Communication-Aware Vehicle Control), we decomposed the multitimescale control and communications (MTCCs) problem in cellular vehicle-to-everything (C-V2X) system into a communication-aware DRL-based platoon control (PC) subproblem and a control-aware DRL-based radio resource allocation (RRA) subproblem. We focused on the PC subproblem and proposed the MTCC-PC algorithm to learn an optimal PC policy given an RRA policy. In this article (Part II), we first focus on the RRA subproblem in MTCC assuming a PC policy is given, and propose the MTCC-RRA algorithm to learn the RRA policy. Specifically, we incorporate the PC advantage function in the RRA reward function, which quantifies the amount of PC performance degradation caused by observation delay. Moreover, we augment the state space of RRA with PC action history for a more well-informed RRA policy. In addition, we utilize reward shaping and reward backpropagation prioritized experience replay (RBPER) techniques to efficiently tackle the multiagent and sparse reward problems, respectively. Finally, a sample- and computational-efficient training approach is proposed to jointly learn the PC and RRA policies in an iterative process. In order to verify the effectiveness of the proposed MTCC algorithm, we performed experiments using real driving data for the leading vehicle, where the performance of MTCC is compared with those of the baseline DRL algorithms. Lei Lei 0004, Tong Liu 0035, Kan Zheng, Xuemin Shen |
IEEE Internet Things J. | 2 |
| 2024 | Multitimescale Control and Communications With Deep Reinforcement Learning - Part I: Communication-Aware Vehicle ControlabstractAn intelligent decision-making system enabled by vehicle-to-everything (V2X) communications is essential to achieve safe and efficient autonomous driving (AD), where two types of decisions have to be made at different timescales, i.e., vehicle control and radio resource allocation (RRA) decisions. The interplay between RRA and vehicle control necessitates their collaborative design. In this two-part paper (Part I and Part II), taking platoon control (PC) as an example use case, we propose a joint optimization framework of multitimescale control and communications (MTCCs) MTCCs based on deep reinforcement learning (DRL). In this article (Part I), we first decompose the problem into a communication-aware DRL-based PC subproblem and a control-aware DRL-based RRA subproblem. Then, we focus on the PC subproblem assuming an RRA policy is given, and propose the MTCC- PC algorithm to learn an efficient PC policy. To improve the PC performance under random observation delay, the PC state space is augmented with the observation delay and PC action history. Moreover, the reward function with respect to the augmented state is defined to construct an augmented state Markov decision process (MDP). It is proved that the optimal policy for the augmented state MDP is optimal for the original PC problem with observation delay. Different from most existing works on communication-aware control, the MTCC- PC algorithm is trained in a delayed environment generated by the fine-grained embedded simulation of cellular vehicle-to-everything communications rather than by a simple stochastic delay model. Finally, experiments are performed to compare the performance of MTCC- PC with those of the baseline DRL algorithms. Tong Liu 0035, Lei Lei 0004, Kan Zheng, Xuemin Shen |
IEEE Internet Things J. | 1 |
| 2023 | LEARN: Selecting Samples Without Training Verification for Communication-Efficient Vertical Federated LearningabstractIn the classical vertical federated learning (VFL) framework, feature maps and corresponding gradient information of all samples are transferred between the server and clients, which causes a significant communication burden. Therefore, to enable efficient VFL in resource-constrained wireless networks, we propose to select a part of the samples from the large training set to train models with minimal accuracy degradation. To this end, we propose LEARN, i.e., seLecting Efficient sAmples without tRaining verificatioN, to select efficient training samples for VFL. Particularly, LEARN integrates two major components named label distribution smoothing and feature center-based vertical sample filtering. The number of samples selected for each class is determined by the label distribution smoothing mechanism. Then the feature center-based vertical sample filtering component calculates the features centers and performs sample selection based on the distance between the samples and their corresponding feature center. Extensive experiments under various settings are carried out to corroborate the efficacy and robustness of LEARN. Tong Liu 0035, Feng Lyu 0001, Yongheng Deng, Qilong Tan, Yaoxue Zhang |
GLOBECOM | 1 |
| 2023 | A Prototype-Based Knowledge Distillation Framework for Heterogeneous Federated LearningabstractFederated learning (FL) is an emerging distributed machine learning paradigm, which has shown great potential in collaborative learning with privacy preservation. However, FL clients usually have disparate system resource capabilities (e.g., data, computation, and communication) for model training and aggregation, which can cause a series of system heterogeneity issues with performance degradation. To this end, we propose FedPKD, a Prototype-based Knowledge Distillation framework for FL. FedPKD integrates knowledge distillation and prototype learning with FL, which enables heterogeneous clients and the server to learn collaboratively, with different model architectures and resource capability adaptations. Specifically, FedPKD proposes to transfer dual knowledge of clients including the model output logits and prototypes to the server, and a prototype-based ensemble distillation mechanism is proposed to aggregate the logits and prototypes from clients, which can be used to train the server model with an unlabeled public dataset. The server model knowledge is then transferred back to clients to improve the performance of client models. Moreover, to improve learning performance and reduce communication overhead, we propose a prototype-based data filter mechanism to filter out the samples with low-quality knowledge. Extensive experiments under various settings demonstrate the superiority of FedPKD in learning performance and communication efficiency when compared to state-of-the-art benchmarks. Feng Lyu 0001, Yongheng Deng, Tong Liu 0035, Yongmin Zhang, Yaoxue Zhang |
ICDCS | 4 |
| 2023 | Jointly Learning V2X Communication and Platoon Control with Deep Reinforcement LearningabstractIn autonomous vehicle platooning, Vehicle-to-Everything (V2X) communications are leveraged in cooperative adaptive cruise control (CACC) to improve control performance. Since exchanging information at all times incurs significant communication overhead in vehicular networks, it is important to determine when V2X communication is necessary. To solve this problem, we propose a Deep Reinforcement Learning (DRL)-based algorithm named Attention-DDPG, which learns platoon control with Deep Deterministic Policy Gradient (DDPG), and learns when to communicate with an attention network. Specifically, each preceding vehicle is equipped with a deep neural network (DNN), which takes as input its local state and platoon control action and determines whether to transmit its acceleration or not to the following vehicle at each time step. The attention network of a preceding vehicle is trained using the feedback from the following vehicle on the value of V2X information in the form of an advantage function. In order to evaluate Attention-DDPG, simulations are performed using real driving data, and performance is compared with those of two baselines that communicate and do not communicate at all times, respectively. The results demonstrate that Attention-DDPG strikes a competitive tradeoff between control performance and communication overhead while ensuring platoon string stability. Tong Liu 0035, Lei Lei 0004, Zhiming Liu 0014, Kan Zheng |
PIMRC | 1 |
| 2023 | Autonomous Platoon Control With Integrated Deep Reinforcement Learning and Dynamic ProgrammingabstractAutonomous vehicles in a platoon determine the control inputs based on the system state information collected and shared by the Internet of Things (IoT) devices. Deep reinforcement learning (DRL) is regarded as a potential method for car-following control and has been mostly studied to support a single following vehicle. However, it is more challenging to learn an efficient car-following policy with convergence stability when there are multiple following vehicles in a platoon, especially with unpredictable leading vehicle behavior. In this context, we adopt an integrated DRL and dynamic programming (DP) approach to learn autonomous platoon control policies, which embeds the deep deterministic policy gradient (DDPG) algorithm into a finite-horizon value iteration framework. Although the DP framework can improve the stability and performance of DDPG, it has the limitations of lower sampling and training efficiency. In this article, we propose an algorithm, namely, finite-horizon-DDPG with sweeping through reduced state space using stationary approximation (FH-DDPG-SS), which uses three key ideas to overcome the above limitations, i.e., transferring network weights backward in time, stationary policy approximation for earlier time steps, and sweeping through reduced state space. In order to verify the effectiveness of FH-DDPG-SS, simulation using real driving data is performed, where the performance of FH-DDPG-SS is compared with those of the benchmark algorithms. Finally, platoon safety and string stability for FH-DDPG-SS are demonstrated. Tong Liu 0035, Lei Lei 0004, Kan Zheng, Kuan Zhang 0001 |
IEEE Internet Things J. | 1 |