VLDB 2026 Research / reviewers in the wild / expert
Zhicong Zheng
dblp:300/7922
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0002-7298-0381ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 4 · 1 first-author · 4 since 2021Security and privacy · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PhyFuzz: Detecting Sensor Vulnerabilities with Physical Signal Fuzzing
Zhicong Zheng, Jinghui Wu, Shilin Xiao, Yanze Ren, Chen Yan 0001, Xiaoyu Ji 0001, Wenyuan Xu 0001 |
NDSS | 1 |
| 2024 | Fast and Lightweight Voice Replay Attack Detection via Time-Frequency Spectrum DifferenceabstractDue to the open nature of voice and voice interface, an adversary can spoof voice recognition systems by replaying pre-recorded voice commands from legitimate users, known as the voice replay attack. Existing detection methods against voice replay attacks mainly rely on extra hardware to determine the sound source or require excessive computing resources to train a classifier with abundant acoustic features. In this paper, we propose Anti-Replay, a fast and lightweight detection system for voice replay attacks. To overcome the challenge of redundant classification features and complex calculation, we first investigate the time-frequency spectrum difference between the genuine human voice and the replayed audio caused by the non-linear distortion of the attacker’s microphones and speakers. Then, we design 5 types with a total of 77 features in both the time and frequency domains and propose a convolutional neural network classifier SE-ResNet50 for attack detection. Evaluations against the datasets of ASVspoof2017, ASVspoof2019, and ASVspoof2021 demonstrate that Anti-Replay can achieve an average equal error rate (EER) of 1.36% across three datasets. Meanwhile, Anti-Replay decreases the training time by 52.3% and 90.2% and decreases the model size by 83.5% and 99.9% compared with the baseline model CQCC-GMM and the state-of-the-art method Res2Net. We have also confirmed that our system is effective in detecting the adaptive replay attack. He Ruiwen, Yushi Cheng, Zhicong Zheng, Xiaoyu Ji 0001, Wenyuan Xu 0001 |
IEEE Internet Things J. | 3 |
| 2024 | Toward Pitch-Insensitive Speaker Verification via SoundfieldabstractAutomatic speaker verification systems (ASVs) verify a person’s identity by his/her voice and have been widely deployed for user authentication. However, existing ASVs are based on traditional audio spectral features and hence, perform poorly in verifying pitch-changed utterances from speakers with cold or sore throat. In this article, we propose soundfield tracker(SOFTER), a soundfield-based speaker verification system that can verify speakers regardless of the pitch changes.SOFTERis based on the observation that soundfield features reflect the speaker’s vocal tract, mouth, head, torso, etc., which are less affected by the pitch changes in speech signals.SOFTERcan be integrated into off-the-shelf smartphones without any hardware modifications. One major challenge is that the soundfield is sensitive to the distance between the speaker and the phone. To solve this problem, we propose a two-stage mechanism combining distance sensing and soundfield reconstruction, which enables to reconstruct the soundfield to a setting similar to the one in the enrollment phase, thus, the speaker can be verified from any distance to the phone. We compareSOFTERwith six state-of-the-art academic and commercial ASVs on two data sets of 134 speakers and 31000 speech samples. Results show thatSOFTERhas an equal error rate (EER) of 2.18% and 1.61% on the two data sets, respectively. Moreover,SOFTERoutperforms other ASVs by at least 24.67% on average in verifying pitch-varying or pathological speech samples, denoting an evidence ofSOFTER’s effectiveness in both normal and unhealthy user conditions. Xinfeng Li, Zhicong Zheng, Chen Yan 0001, Chaohao Li, Xiaoyu Ji 0001, Wenyuan Xu 0001 |
IEEE Internet Things J. | 2 |
| 2024 | TRCC: Transferable Congestion Control With Reinforcement LearningabstractThe breathtaking progress in machine learning has motivated studies in learning-based congestion control algorithms, which are expected to adjust congestion window choices according to the dynamic network environment. Nevertheless, existing learning-based congestion control protocols are mostly designed and trained for a specific network environment. When applied to a different network environment, the previously trained model may see a considerable degradation in performance. As the rapid development of communication technologies has given rise to the emergence of a diversity of new networks, it is desirable for learning-based congestion control models to quickly transfer to different network environments. Driven by this motivation, we propose a novel TRansferable Congestion Control (TRCC) protocol, which takes full advantage of both reinforcement learning and transfer learning to intelligently cope with network congestion in scenarios. The key idea to enable a fast transfer from the source network environment to the target network environment is to fine-tune a well-trained model in the source network to suit the target network in a time-effective way. We theoretically prove the transferability and quick convergence of our proposed transfer reinforcement learning-based congestion control algorithm by deriving the Markov transition matrix and the similarity of the reward function. Our experiments validate that TRCC can converge in a new network environment in a short time while achieving comparable performance with baseline algorithms. Zhicong Zheng, Zhenchang Xia, Yu-Cheng Chou, Yanjiao Chen |
IEEE Internet Things J. | 1 |
| 2023 | MicPro: Microphone-based Voice Privacy ProtectionabstractHundreds of hours of audios are recorded and transmitted over the Internet for voice interactions such as virtual calls or speech recognitions. As these recordings are uploaded, embedded biometric information, i.e., voiceprints, is unnecessarily exposed. This paper proposes the first privacy-enhanced microphone module (i.e., MicPro) that can produce anonymous audio recordings with biometric information suppressed while preserving speech quality for human perception or linguistic content for speech recognition. Limited by the hardware capabilities of microphone modules, previous works that modify recording at the software level are inapplicable. To achieve anonymity in this scenario, MicPro transforms formants, which are distinct for each person due to the unique physiological structure of the vocal organs, and formant transformations are done by modifying the linear spectrum frequencies (LSFs) provided by a popular codec (i.e., CELP) in low-latency communications. Shilin Xiao, Xiaoyu Ji 0001, Chen Yan 0001, Zhicong Zheng, Wenyuan Xu 0001 |
CCS | 4 |
| 2023 | The Silent Manipulator: A Practical and Inaudible Backdoor Attack against Speech Recognition SystemsabstractBackdoor Attacks have been shown to pose significant threats to automatic speech recognition systems (ASRs). Existing success largely assumes backdoor triggering in the digital domain, or the victim will not notice the presence of triggering sounds in the physical domain. However, in practical victim-present scenarios, the over-the-air distortion of the backdoor trigger and the victim awareness raised by its audibility may invalidate such attacks. In this paper, we propose SMA, an inaudible grey-box backdoor attack that can be generalized to real-world scenarios where victims are present by exploiting both the vulnerability of microphones and neural networks. Specifically, we utilize the nonlinear effects of microphones to inject an inaudible ultrasonic trigger. To accurately characterize the microphone response to the crafted ultrasound, we construct a novel nonlinear transfer function for effective optimization. We also design optimization objectives to ensure triggers' robustness in the physical world and transferability on unseen ASR models. In practice, SMA can bypass the microphone's built-in filters and human perception, activating the implanted trigger in the ASRs inaudibly, regardless of whether the user is speaking. Extensive experiments show that the attack success rate of SMA can reach nearly 100% in the digital domain and over 85% against most microphones in the physical domains by only poisoning about 0.5% of the training audio dataset. Moreover, our attack can resist typical defense countermeasures to backdoor attacks. Zhicong Zheng, Xinfeng Li, Chen Yan 0001, Xiaoyu Ji 0001, Wenyuan Xu 0001 |
ACM Multimedia | 1 |
| 2023 | MARNet: Backdoor Attacks Against Cooperative Multi-Agent Reinforcement LearningabstractRecent works have revealed that backdoor attacks against Deep Reinforcement Learning (DRL) could lead to abnormal action selections of the agent, which may result in failure or even catastrophe in crucial decision processes. However, existing attacks only consider single-agent reinforcement learning (RL) systems, in which the only agent can observe the global state and have full control of the decision process. In this article, we explore a new backdoor attack paradigm in cooperative multi-agent reinforcement learning (CMARL) scenarios, where a group of agents coordinate with each other to achieve a common goal, while each agent can only observe the local state. In the proposed MARNet attack framework, we carefully design a pipeline of trigger design, action poisoning, and reward hacking modules to accommodate the cooperative multi-agent settings. In particular, as only a subset of agents can observe the triggers in their local observations, we maneuver their actions to the worst actions suggested by an expert policy model. Since the global reward in CMARL is aggregated by individual rewards from all agents, we propose to modify the reward in a way that boosts the bad actions of poisoned agents (agents who observe the triggers) but mitigates the influence on non-poisoned agents. We conduct extensive experiments on three classical CMARL algorithms VDN, COMA, and QMIX, in two popular CMARL games Predator Prey and SMAC. The results show that the baselines extended from single-agent DRL backdoor attacks seldom work in CMARL problems while MARNet performs well by reducing the utility under attack by nearly 100%. We apply fine-tuning as a potential defense against MARNet and demonstrate that fine-tuning cannot entirely eliminate the effect of the attack. Yanjiao Chen, Zhicong Zheng, Xueluan Gong |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2021 | A Multi-objective Reinforcement Learning Perspective on Internet Congestion ControlabstractThe advent of new network architectures has resulted in the rise of network applications with different network performance requirements: live video streaming applications require low latency. In contrast, file transfer applications require high throughput. Existing congestion control protocols may fail to simultaneously meet the performance requirements of these different types of applications since their designed objective function is fixed and difficult to readjust according to the needs of the application. In this paper, we develop MOCC (Multi-Objective Congestion Control), a novel multi-objective congestion control protocol that can meet the performance requirements of different applications without the need to redesign the objective function. MOCC leverages multi-objective reinforcement learning with preferences in order to adapt to different types of applications. By addressing challenges such as slow convergence speed and the difficulty of designing the end of the episode, MOCC can quickly converge to the equilibrium point and adapt multi-objective reinforcement learning to congestion control. Through an extensive array of experiments, we discover that MOCC outperforms the most recent state-of-the-art congestion control protocols and can achieve a trade-off between throughput, latency, and packet loss, meeting the performance requirements of different types of applications by setting preferences. Zhenchang Xia, Yanjiao Chen, Yu-Cheng Chou, Zhicong Zheng, Baochun Li |
IWQoS | 5 |