Liangyu Zhang

dblp:123/7110 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 ReMo-Det: A Refine-then-Modulate Framework for Efficient Dual-Modal Tiny Object Detection
Keting Jiang, Yinuo Yin, Liangyu Zhang
ICIC (16)6
2026 A Trust Assessment Method for Intelligent Connected Vehicles Based on Data Consistency Verification Matched With Adaptive Leader Election in a Zero-Trust Framework
abstract
The use of intelligent connected vehicles (ICVs) is an emerging concept in transportation systems, where the platoon leader processes various types of information from followers and roadside units to make accurate decisions. However, in complex network environments, the leader election process can become complicated and inefficient due to potential network attacks and abnormal node behaviours. To address this challenge, this paper presents a leader vehicle election mechanism based on dynamic trust assessment within a zero-trust framework. The proposed mechanism comprises two key components: dynamic trust assessment and adaptive leader election. The dynamic trust assessment method involves a trust evaluation mechanism composed of direct trust assessment and recommendation trust assessment based on the state information of vehicle nodes, where the dynamic trust assessment results are adjusted according to node behaviour information. The leader election mechanism involves periodic elections among the vehicle nodes on the basis of the results of the dynamic trust assessment, in which the vehicle node with the most votes is selected as the leader. Through comparative experiments against a full trust scenario, the proposed method demonstrates superior performance in resisting network attacks and identifying abnormal nodes, thereby significantly improving the safety, stability, and operational efficiency of the vehicle platoon.
Darong Huang 0002, Liangyu Zhang, Yuhong Na, Fawen Bu, Zhongmei Li
IEEE Trans. Dependable Secur. Comput.2
2025 A Finite Sample Analysis of Distributional TD Learning with Linear Function Approximation
abstract
In this paper, we study the finite-sample statistical rates of distributional temporal difference (TD) learning with linear function approximation. The aim of distributional TD learning is to estimate the return distribution of a discounted Markov decision process for a given policy $\pi$. Previous works on statistical analysis of distributional TD learning mainly focus on the tabular case. In contrast, we first consider the linear function approximation setting and derive sharp finite-sample rates. Our theoretical results demonstrate that the sample complexity of linear distributional TD learning matches that of classic linear TD learning. This implies that, with linear function approximation, learning the full distribution of the return from streaming data is no more difficult than learning its expectation (value function). To derive tight sample complexity bounds, we conduct a fine-grained analysis of the linear-categorical Bellman equation and employ the exponential stability arguments for products of random matrices. Our results provide new insights into the statistical efficiency of distributional reinforcement learning algorithms.
Kaicheng Jin, Liangyu Zhang
NeurIPS3
2025 Hotness Strategy based Secure Cross-user Deduplication for Cloud Storage
abstract
Random chunks generation attack is a more sophisticated form of side channel attack in cross-user chunk level deduplication, in which the target chunks are checked together with randomly generated non-duplicate ones. When the number of required chunks or linear combinations equals to that of the randomly generated chunks, this issue prevents existing schemes from achieving obfuscation in the response, thereby compromising the existence privacy of the target chunks. Moreover, this vulnerability becomes even more pronounced under the hybrid attack, in which a portion of the target chunks is replaced by public ones. To address both kinds of attacks in a lightweight way, we propose a hotness strategy based secure cross-user deduplication scheme (HSDS) for cloud storage. Specifically, the cloud service provider (CSP) records the feature hotness and chunk hotness of duplicate chunks before defining specific chunks to be non-duplicate in response generation. Both the security analysis and experimental results show that the probability to expose the existence of sensitive chunks of the proposed scheme is negligible, which is effective in resisting the random chunks generation attack and its hybrid form in a lightweight way.
Liangyu Zhang, Junyu Cheng
TrustCom1
2024 Statistical Efficiency of Distributional Temporal Difference Learning
abstract
Distributional reinforcement learning (DRL) has achieved empirical success in various domains. One of the core tasks in the field of DRL is distributional policy evaluation, which involves estimating the return distribution $\eta^\pi$ for a given policy $\pi$. The distributional temporal difference learning has been accordingly proposed, which is an extension of the temporal difference learning (TD) in the classic RL area. In the tabular case, Rowland et al. [2018] and Rowland et al. [2023] proved the asymptotic convergence of two instances of distributional TD, namely categorical temporal difference learning (CTD) and quantile temporal difference learning (QTD), respectively. In this paper, we go a step further and analyze the finite-sample performance of distributional TD. To facilitate theoretical analysis, we propose a non-parametric distributional TD learning (NTD). For a $\gamma$-discounted infinite-horizon tabular Markov decision process, we show that for NTD we need $\widetilde O\left(\frac{1}{\varepsilon^{2p}(1-\gamma)^{2p+1}}\right)$ iterations to achieve an $\varepsilon$-optimal estimator with high probability, when the estimation error is measured by the $p$-Wasserstein distance. This sample complexity bound is minimax optimal (up to logarithmic factors) in the case of the $1$-Wasserstein distance. To achieve this, we establish a novel Freedman's inequality in Hilbert spaces, which would be of independent interest. In addition, we revisit CTD, showing that the same non-asymptotic convergence bounds hold for CTD in the case of the $p$-Wasserstein distance.
Liangyu Zhang, Zhihua Zhang 0004
NeurIPS2
2024 Semi-Infinitely Constrained Markov Decision Processes and Provably Efficient Reinforcement Learning
abstract
We propose a novel generalization of constrained Markov decision processes (CMDPs) that we call the semi-infinitely constrained Markov decision process (SICMDP). Particularly, we consider a continuum of constraints instead of a finite number of constraints as in the case of ordinary CMDPs. We also devise two reinforcement learning algorithms for SICMDPs that we refer to as SI-CMBRL and SI-CPO. SI-CMBRL is a model-based reinforcement learning algorithm. Given an estimate of the transition model, we first transform the reinforcement learning problem into a linear semi-infinitely programming (LSIP) problem and then use the dual exchange method in the LSIP literature to solve it. SI-CPO is a policy optimization algorithm. Borrowing ideas from the cooperative stochastic approximation approach, we make alternative updates to the policy parameters to maximize the reward or minimize the cost. To the best of our knowledge, we are the first to apply tools from semi-infinitely programming (SIP) to solve constrained reinforcement learning problems. We present theoretical analysis for SI-CMBRL and SI-CPO, identifying their iteration complexity and sample complexity. We also conduct extensive numerical experiments to illustrate the SICMDP model and demonstrate that our proposed algorithms are able to solve complex control tasks leveraging modern deep reinforcement learning techniques.
Liangyu Zhang, Zhihua Zhang 0004
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Central Aortic Blood Pressure Waveform Estimation with a Temporal Convolutional Network
abstract
A novel temporal convolutional network (TCN) model is utilized to reconstruct the central aortic blood pressure (aBP) waveform from the radial blood pressure waveform. The method does not need manual feature extraction as traditional transfer function approaches. The data acquired by the SphygmoCor CVMS device in 1,032 participants as a measured database and a public database of 4,374 virtual healthy subjects were used to compare the accuracy and computational cost of the TCN model with the published convolutional neural network and bi-directional long short-term memory (CNN-BiLSTM) model. The TCN model was compared with CNN-BiLSTM in the root mean square error (RMSE). The TCN model generally outperformed the existing CNN-BiLSTM model in terms of accuracy and computational cost. For the measured and public databases, the RMSE of the waveform using the TCN model was 0.55 ± 0.40 mmHg and 0.84 ± 0.29 mmHg, respectively. The training time of the TCN model was 9.63 min and 25.51 min for the entire training set; the average test time was around 1.79 ms and 8.58 ms per test pulse signal from the measured and public databases, respectively. The TCN model is accurate and fast for processing long input signals, and provides a novel method for measuring the aBP waveform. This method may contribute to the early monitoring and prevention of cardiovascular disease.
Wenyan Liu 0002, Shuo Du, Na Pang, Liangyu Zhang, Guozhe Sun, Hanguang Xiao, Qi Zhao 0008, Lisheng Xu, Yu-Dong Yao, Jordi Alastruey, Alberto P. Avolio
IEEE J. Biomed. Health Informatics4
2022 Semi-infinitely Constrained Markov Decision Processes
abstract
We propose a generalization of constrained Markov decision processes (CMDPs) that we call the \emph{semi-infinitely constrained Markov decision process} (SICMDP).Particularly, in a SICMDP model, we impose a continuum of constraints instead of a finite number of constraints as in the case of ordinary CMDPs.We also devise a reinforcement learning algorithm for SICMDPs that we call SI-CRL.We first transform the reinforcement learning problem into a linear semi-infinitely programming (LSIP) problem and then use the dual exchange method in the LSIP literature to solve it.To the best of our knowledge, we are the first to apply tools from semi-infinitely programming (SIP) to solve reinforcement learning problems.We present theoretical analysis for SI-CRL, identifying its sample complexity and iteration complexity.We also conduct extensive numerical examples to illustrate the SICMDP model and validate the SI-CRL algorithm.
Liangyu Zhang, Zhihua Zhang 0004
NeurIPS1