VLDB 2026 Research / reviewers in the wild / expert
Yilin Xiao 0001
dblp:256/8598-1
· DBLP profile ↗
13ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-0717-7488ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 9 · 4 first-author · 7 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reinforcement Learning-Based Edge-Assisted Inference With Multimodal DataabstractDeep neural networks (DNNs) extract embeddings from multimodal data such as audio, images, and LiDAR, each with heterogeneous computational demands and representation abilities to support multimodal services such as audio-visual speech recognition. Reinforcement learning (RL)-based model selection and splitting schemes determine the model variants from DNNs and the components to offload in unimodal models to reduce inference latency, aiming for a three-fold trade-off among inference accuracy, computation cost, and communication overhead but ignore the heterogeneity of different modalities within multimodal DNNs. In this paper, we propose an RL-based edge-assisted multimodal inference scheme that optimizes model selection at modality level and edge-assisted policies, including collaborative servers and partition points for each feature extractor to perform multimodal DNNs on mobile devices. Based on the information complexity and historical influence of each modality, as well as real-time observations such as channel gain, and previous inference performance, the policy distributions are designed to maximize the utility, as a weighted sum of inference latency and energy consumption, and inference accuracy. Safe policy exploration further mitigates risks such as low inference accuracy, intolerable inference latency, and improper allocation of computational resources to specific modalities. We analyze the computational complexity affected by the number of model variants, edge servers, and partition points and derive performance bounds for inference latency, energy consumption, and utility under specific sample sizes and data rates. Experimental results show that the proposed schemes improve inference performance compared to benchmark schemes. Liang Xiao 0003, Chuxuan Wang, Zefang Lv, Yiwen Zhan 0002, Yilin Xiao 0001, Helin Yang |
IEEE Trans. Mob. Comput. | 6 |
| 2025 | Distributionally Robust Optimization for Energy Efficiency in Heterogeneous Wireless NetworksabstractEnergy efficiency (EE) is vital for 5G networks to manage the increased traffic demand while minimizing operational costs and reducing environmental impact. Optimizing energy usage also supports scalability and helps meet regulatory sustainability goals as network infrastructure expands. In a practical commercial 5G network, a widely-used method to reduce energy consumption is to dynamically shut down underutilization cells and reduce the overlapping cell coverage based on real-time traffic patterns. However, the inherent uncertainty of traffic demands, coupled with the unknown distribution function, significantly complicates the optimization of cell shutdown strategies, rendering both conventional deterministic and stochastic optimization methods ineffective. To tackle these challenges, we propose a distributionally robust optimization framework for EE optimization that does not rely on distributional knowledge. First, we construct a data-driven uncertainty set to model traffic distribution and apply Lagrangian duality to transform the infinite-dimensional optimization problem into a more tractable finite-dimensional one. Then, we employ a Bayesian optimization algorithm to efficiently solve this reformulated problem, which involves a mixed space of high-dimensional combinatorial cell shutdown and continuous Lagrangian multipliers, even with constrained black-box evaluations. Simulations on real-world field data show that our proposed solution outperforms existing benchmarks, achieving energy efficiency improvements ranging from 0.7 % to 11.05 %. Yilin Xiao 0001, Zhongji Wang, Zhizongkai Wang, Xufeng Chen, Lin Gao 0001, Fen Hou, Jianwei Huang 0001 |
ICC | 1 |
| 2025 | Deep Reinforcement Learning-Based Few-Shot Image SteganographyabstractDeep learning-based steganography techniques mainly rely on large-scale datasets with sufficient training and testing images, but the concealment performance may degrade in the few-shot-based wireless scenarios, such as vehicular networks. In this paper, we propose a few-shot image steganography framework with deep reinforcement learning (RL), which applies an adaptive generative adversarial network (GAN) to embed secret messages in cover images. This is the first work combining hierarchical deep RL with image steganography. We design an improved three-level hierarchical structure for deep RL, which determines the encoder in the GAN to improve the image quality and similarity, the number of cover images to form a few-shot training dataset, and the secret embedding depth to improve concealment performance and reduce overhead. Experiments are performed on both grayscale and RGB images from BOSSbase and Div2K datasets, demonstrating that the proposed framework surpasses SteganoGAN and HiNet benchmarks with better concealment performance and higher ability against steganalysis. Dexiang Ren, Xiaozhen Lu, Yilin Xiao 0001 |
VTC2025-Spring | 4 |
| 2025 | Learning-Based Low-Latency Collaborative Inference for Multi-Branch Models in D2D-Assisted MECabstractDeploying high-complexity deep neural network (DNN) on mobile devices presents significant challenges, stemming from the conflict between their computationally intensive demands and the constrained computational resources. Device-to-device ($D$2$D$) assisted mobile edge computing (MEC) exploits sharing resources among devices to perform DNN inference collaboratively for higher efficiency. However, most existing collaborative inference schemes ignore the structure of DNN models, thus suffering from high latency under dynamic network conditions and computational loads. In this paper, we propose a learning-based low-latency collaborative inference scheme for multi-branch models, which schedules the computations of each DNN layer to optimize the alignment between the DNN structure's characteristics and the D2D network environments. Specifically, the computational task of each branch is assigned to nearby mobile devices or the edge server for efficient collaboration. In addition, we design proximal policy optimization with task-specific architecture to choose the computational scheduling policy, which contains a two-stream NN to extract features from the D2D network and the tasks at each layer, as well as a multi-head output to obtain the collaborative devices for all tasks in the layer. Simulation results show that the proposed scheme reduces the inference latency compared with benchmarks. Yilin Xiao 0001, Xiaozhen Lu, Liang Xiao 0003 |
WCNC | 2 |
| 2025 | Spectral Co-Clustering Based Wireless Network Decomposition for Resource SchedulingabstractLarge-scale wireless networks pose significant challenges in resource scheduling, where the solution space grows exponentially with network size. While network decomposition offers a promising solution by breaking networks into manageable subnetworks, existing approaches, including spectral clustering, fail to effectively capture the complex service relationships between base stations (BSs) and users, particularly in networks with massive user populations. This paper presents BSCCD (Bidirectional Spectral Co-Clustering Based Decomposition), a new decomposition scheme that addresses these challenges through two key innovations: (i) a two-round spectral co-clustering framework that captures bidirectional BS-user relationships, and (ii) a user node merging strategy that handles massive user populations. Extensive experiments on real-world datasets from multiple Chinese cities demonstrate that BSCCD reduces computation latency by up to 61.91 % compared to global optimization, while achieving more than 10 % improvement in solution quality over traditional clustering approaches. The advantage is particularly pronounced in medium-scale networks, where BSCCD outperforms traditional methods by$\mathbf{2 8. 7 6 \%}$. Our results demonstrate BSCCD's practical viability for resource scheduling in contemporary wireless networks, especially in scenarios with complex BS-user interactions and large user populations. Yiyu Liu, Yilin Xiao 0001, Ming Tang 0006, Lin Gao 0001, Jianwei Huang 0001 |
WiOpt | 2 |
| 2025 | Reinforcement Learning-Based Accurate Worm Detection for Smart GridsabstractReinforcement learning (RL) based worm detection chooses the test threshold to evaluate the network traffic features such as the spectral flatness measure (SFM), but the detection of the evasive worm that modifies the scan rate and the worm propagation speed to manipulate the network traffic features is inaccurate due to the estimation and quantization error in the test threshold. In this paper, we propose an RL based accurate worm detection for smart grids that enables the control center to optimize the test threshold based on the number of meters, the infection time series and the number of connections to new destination IP addresses, besides the traffic log size received from each data concentrator and the number of the previously infected meters. A constraint on the maximum missed detection rate required by the smart grids is exploited in the detection policy distribution to support the reliable data transmission. A deep RL version addresses the quantization error in terms of the test threshold in the SFM evaluation and the infection time series in the state formulation, and compress the state space of the traffic log size received from a large number of data concentrators. Based on a worm detection game, the performance bound is provided under the specified detection window size and the propagation speed of evasive worm. Simulation results for 3600 meters show that the performance gain of the detection accuracy and latency against evasive worm over the benchmarks. Liang Xiao 0003, Jieling Li, Yilin Xiao 0001, Zefang Lv, Chuxuan Wang, Pengmin Li |
IEEE Internet Things J. | 3 |
| 2025 | Blockchain-Enabled Secure Offloading for VEC: A Multi-Agent Reinforcement Learning ApproachabstractVehicular edge computing (VEC) helps improve the task computational performance of vehicles on roads but has difficulty in defending against eavesdropping and selfish attacks simultaneously. In this paper, we design a reputation-based smart contract with blockchain and propose a multi-agent reinforcement learning (RL) based secure offloading scheme for VEC against both eavesdropping and selfish attacks. This scheme has a three-level hierarchical structure for each vehicle and uses the reputations obtained from the blockchain as the basis to optimize the edge node selection, offloading ratio, and power allocation, which aims to reduce the task computational latency, the vehicle energy consumption and eavesdropping rate. By using a punishment function based on the constraints, this scheme avoids exploring dangerous policies that can cause task failure or severe data leakage. A multi-agent deep RL-based secure offloading scheme is proposed for vehicles with sufficient resources, which evaluates the long-term risk rather than the punishment function to further improve the secure offloading performance. The regret bound is analyzedand the cumulative reward upper bound is provided. Simulation results verify the effectiveness of our schemes as compared with the benchmark. Xiaozhen Lu, Liang Xiao 0003, Yilin Xiao 0001, Zehui Xiong, Zhe Liu 0001, Yanyong Zhang, Weihua Zhuang |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2025 | Reinforcement Learning-Based False Data Injection Attacks in Smart GridsabstractFalse data injection (FDI) attacks construct attack vectors to inject false data into tampered meters with the goal of falsifying state estimation, but resulting in low successful attack rate with high attack costs in terms of the number of tampered meters in large-scale smart grids, because the bad data detection at the control center chooses the dynamic detection thresholds to identify the modified meter measurements. In this article, we propose a reinforcement learning-based FDI attack scheme that optimizes both the tampered meters and the false data to enhance the success attack rate and injected errors while reducing attack costs. Based on meter measurements and previous performance, the attack vector is constructed to induce more errors in state estimation and bypass bad data detection. The performance bounds regarding the successful attack rate and the injected error are derived in terms of the number of bus phase angles, the susceptance of the transmission line, and the maximum false data based on the Nash equilibrium of the FDI game. Simulations performed on both the IEEE 14-bus and IEEE 118-bus systems demonstrate the performance gain over the benchmarks. Liang Xiao 0003, Haoyu Chen 0005, Zefang Lv, Chuxuan Wang, Yilin Xiao 0001 |
IEEE Trans. Ind. Informatics | 6 |
| 2024 | Risk-Aware Federated Reinforcement Learning-Based Secure IoV CommunicationsabstractWith the rapid growth in the number of high-mobility vehicles and booming enhanced applications with restricted latency requirements, downlink communication in Internet of Vehicles (IoV) systems has become increasingly vulnerable to active eavesdropping attacks. This paper proposes a federated learning-enabled secure communication framework for IoV against active eavesdropping, in which the roadside units (RSUs) apply reinforcement learning (RL) model to optimize their downlink transmit power levels, and the server helps update the RL models of the RSUs. First, we design a multi-agent deep RL algorithm for each RSU, which designs a punishment and a blacklist mechanism to mitigate risky explorations related to severe data leakage or communication outages. Second, this framework designs a risk-aware RL for the server, which uses a two-level hierarchical structure to choose the number of participated RSUs and the corresponding local training data size for higher optimization speed. This framework considers both the reward and risk in the selection of policies to reduce the probability of exploring the risky training policies that cause defense failure of the RSUs against active eavesdropping. Third, we analyze the convergence performance, computational complexity, and reward upper bound, which reveals how the power constraint, radio bandwidth and data size affect the secure communication performance. Simulation and experimental results validate the effectiveness of our schemes, such as the reductions of the eavesdropping rate, training latency, and the loss of local models compared to the benchmarks. Xiaozhen Lu, Liang Xiao 0003, Yilin Xiao 0001, Wei Wang 0100, Nan Qi 0001, Qian Wang 0002 |
IEEE Trans. Mob. Comput. | 3 |
| 2023 | Reinforcement Learning Based Energy-Efficient Collaborative Inference for Mobile Edge ComputingabstractCollaborative inference in mobile edge computing (MEC) enables mobile devices to offload the computation tasks for the computation-intensive perception services, and the inference policy determines the inference latency and energy consumption. The optimal inference policy depends on the inference performance model of deep learning, the data generation model and the network model that are rarely known by mobile devices in time. In this paper, we propose a multi-agent reinforcement learning (RL) based energy-efficient MEC collaborative inference scheme, which enables each mobile device to choose both the partition point of deep learning and the collaborative edge of each mobile device based on the image quantity, the channel conditions and the previous inference performance. A learning experience exchange mechanism exploits the Q-values of the neighboring mobile devices to accelerate the inference policy optimization with less energy consumption. We also provide a deep multi-agent RL based inference scheme to accelerate learning for large-scale MEC networks, in which an actor network yields the collaborative inference policy probability distribution and a critic network guides the weight update of the actor network to enhance sample efficiency. We provide the inference performance bound and analyze the computational complexity. Both simulation and experimental results show that our proposed schemes reduce the inference latency and save the MEC energy consumption. Yilin Xiao 0001, Liang Xiao 0003, Kunpeng Wan, Helin Yang, Yi Zhang 0035, Yi Wu 0010, Yanyong Zhang |
IEEE Trans. Commun. | 1 |
| 2021 | Deep-Reinforcement-Learning-Based User Profile Perturbation for Privacy-Aware RecommendationabstractUser profile perturbation protects privacy in the release of user profiles to receive recommendation services, in which the privacy budget as a privacy parameter can be controlled to effect a tradeoff between the recommendation quality and privacy protection against inference attacks. In this article, we propose a deep reinforcement learning (RL)-based user profile perturbation scheme for recommendation systems. This scheme applies differential privacy to protect user privacy and uses deep RL to choose the privacy budget against inference attackers. Based on an evaluated neural network (NN) and a target NN, this scheme enables a user device to optimize the privacy budget over time based on the sensitivity level of the clicked item, the similarities among the recommended items, and the estimated privacy loss. We provide an upper bound on the privacy protection performance of this scheme in the recommendation game and evaluate its computational complexity. Simulation results for a movie recommendation system show that this scheme increases the user privacy protection level for a given recommendation quality compared with benchmark schemes. Yilin Xiao 0001, Liang Xiao 0003, Xiaozhen Lu, Hailu Zhang, Shui Yu 0001, H. Vincent Poor |
IEEE Internet Things J. | 1 |
| 2020 | Reinforcement Learning-Based Downlink Interference Control for Ultra-Dense Small CellsabstractThe dense deployment of small cells in 5G cellular networks raises the issue of controlling downlink inter-cell interference under time-varying channel states. In this paper, we propose a reinforcement learning based power control scheme to suppress downlink inter-cell interference and save energy for ultra-dense small cells. This scheme enables base stations to schedule the downlink transmit power without knowing the interference distribution and the channel states of the neighboring small cells. A deep reinforcement learning based interference control algorithm is designed to further accelerate learning for ultra-dense small cells with a large number of active users. Analytical convergence performance bounds including throughput, energy consumption, inter-cell interference, and the utility of base stations are provided and the computational complexity of our proposed scheme is discussed. Simulation results show that this scheme optimizes the downlink interference control performance after sufficient power control instances and significantly increases the network throughput with less energy consumption compared with a benchmark scheme. Liang Xiao 0003, Hailu Zhang, Yilin Xiao 0001, Xiaoyue Wan, Sicong Liu 0002, Li-Chun Wang 0001, H. Vincent Poor |
IEEE Trans. Wirel. Commun. | 3 |
| 2019 | Privacy Aware Recommendation: Reinforcement Learning Based User Profile PerturbationabstractUser profile release in recommendation systems can apply the user profile perturbation technique to protect user privacy, in which each user sends a perturbed user profile such as the a list of clicked items to receive a recommendation service from a server. The perturbation policy such as the privacy budget determines the recommendation quality and the privacy level, while its optimization usually depends on the known attack model, which is rarely known by the users. In this paper, we propose a reinforcement learning based user profile perturbation scheme that applies differential privacy to protect user privacy for recommendation systems. According to reinforcement learning, the privacy budget to perturb the released user profile depends on the features of the actual user profiles and the released user profiles, and the estimated user privacy level. This scheme enables a user to optimize his or her perturbation policy in terms of both the user privacy level and the received recommendation quality without being aware of the attack model. We evaluate the computational complexity of this scheme and analyze a case study, a privacy aware movie recommendation system. Simulation results show that this scheme improves user privacy protection for a given level of recommendation quality compared with a benchmark profile perturbation scheme. Yilin Xiao 0001, Liang Xiao 0003, Hailu Zhang, Shui Yu 0001, H. Vincent Poor |
GLOBECOM | 1 |