Ziyi Liu 0009

dblp:09/3084-9 · also Zi-yi Liu 0009 · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0002-6796-650XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 6 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Time Shuffle: A Transferability-Booster for Multiple Audio Adversarial Tasks
abstract
Existing audio adversarial attack methods suffer from poor transferability, primarily due to insufficient exploration of model decision mechanisms and overreliance on heuristic-driven algorithm design. This paper aims to alleviate this gap. Specifically, through observations across three mainstream audio tasks (Automatic Speech Recognition, Speaker Verification, and Keyword Spotting), we reveal that these models primarily rely on local temporal features—inputs with time shuffled retain 83.7% of original accuracy. The SHAP-based visualization further validated that time shuffle leads to a significant shift in the salient regions of the model, but the samples can still be correctly identified, indicating the presence of redundant features that can affect decision-making. Inspired by these findings, we propose Time-Shuffle (TS) adversarial attack (including segments-based TS and phoneme-level-based TS-p). This method divides audio or phonemes into segments, randomly shuffles them, and computes gradients on the shuffled structure. By forcing perturbations to exploit transferable local temporal features and reduce overfitting to source-specific patterns, TS/TS-p inherently enhances transferability. As a model-agnostic framework, TS/TS-p can seamlessly integrate with existing attack methods. Comprehensive experiments demonstrate that TS-p achieved SOTA and boosts transferability by about 23%/14.7%/6.3% on ASR/ASV/KWS.
Jiacheng Deng 0001, Dengpan Ye, Zhaolin Wei, Ziyi Liu 0009
AAAI5
2026 Universal Transferable Dual Attack on Anti-Spoofing and Recognition in Facial Security Systems
Sirun Chen, Sipeng Shen, Ziyi Liu 0009, Yueyun Shang, Dengpan Ye
ICIC (16)4
2026 Perceptual V-Cloak: Generating Trainable and Unintelligible Speech Dataset
Dengpan Ye, Jiacheng Deng 0001, Zhaolin Wei, Ziyi Liu 0009
ICIC (2)5
2026 A Network Security Situation Assessment Scheme With Attack Detection Optimized by TAG-Net
abstract
The rapid development and spread of digital technology, information and communications technology, have led to an increasing number of individuals and organizations relying on the Internet for their daily work and life. However, a variety of emerging threats pose significant obstacles to traditional defense strategies. Network Security Situation Assessment (NSSA) is an effective measure to protect network systems from malicious attacks and provides an effective solution to protect network security. However, the existing NSSA schemes suffer from low accuracy and poor efficiency when dealing with network traffic data, which is characterized by large scale, nonlinearity, irregularity, high dimensionality, and temporal correlation. To solve these problems, we propose a novel integration of Time-Attention and residual connections, and design a NSSA scheme for network in this paper. Specifically, we first design a Time-Attention mechanism, which can achieve sufficient feature extraction with linear complexity. Subsequently, we integrate the Time-Attention and residual connection to improve the Gated Recurrent Unit (GRU) and design a novel integration of Time-Attention and residual connections neural network model called Time Attention Gated Network (TAG-Net). TAG-Net uses Time-Attention and residual connection as a reset gate and an update gate to reduce conflicts between reset and update gates. Meanwhile, we propose a TAG-Net-based NSSA scheme for network, which can improve the assessment accuracy and efficiency. Finally, we implement our proposed scheme and provide a performance evaluation. The experimental results show an accuracy of 81.87% for the NSL-KDD dataset, 98.28% for the UNSW-NB15 dataset, and 99.99% for the Bot-IoT dataset compared to the state-of-the-art models.
Yingcong Lan, Yong Ding 0005, Yueling Liu, Ziyi Liu 0009
IEEE Internet Things J.5
2026 MSFT-Net: Mixture Semantic-Agnostic Manipulation Trace Enhanced Architecture for Robust Image Manipulation Localization
abstract
Since the proliferation of image manipulation methods, effective image manipulation localization (IML) in scenarios with post-processing operations gradually becomes a core challenge. For a long time, IML either relies on strongly semantically related features, resulting in semantic relevance bias in the localization results, or only uses a single semantic-agnostic space feature, which is unable to maintain effective localization capabilities after image post-processing operations. Inspired by this, we propose a novel mixture semantic-agnostic manipulation trace robust localization network (MSFT-Net), which specifically utilizes mixture semantic-agnostic information to achieve effective and robust IML. The MSFT-Net introduces two new modules, the mixture shared manipulation trace enhancement module (MISE) and the Multiscale Feature Association Module (FAM). MISE dynamically links multiple semantic-agnostic feature extractors using a sparsity-enhanced mixture of shared experts, enabling the extraction of diverse manipulation features for accurate localization. Furthermore, keeping the high resolution of the localization features is very important in the mask prediction stage. Therefore, FAM outputs high-resolution fused manipulation features by using the correlation of features at the same level and the spatial context information from different levels. This further improves the effectiveness of IML in post-processing scenarios. Comprehensive experiments on five datasets demonstrate that our model significantly improves both in localization accuracy (average F1 score and IoU increasing by over 9.9% and 4.0%) and robustness. The codes will be made available.
Dengpan Ye, Yunming Zhang, Jiacheng Deng 0001, Ziyi Liu 0009, Yueyun Shang, Zhihong Tian 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 DIP-Watermark: A Double Identity Protection Method Based on Robust Adversarial Watermark
abstract
The wide deployment of Face Recognition (FR) systems poses privacy risks. One countermeasure is adversarial attack, deceiving unauthorized malicious FR, but it also disrupts regular identity verification of trusted authorizers, exacerbating the potential threat of identity impersonation. To address this, we propose the first double identity protection scheme based on traceable adversarial watermarking, termed DIP-Watermark. DIP-Watermark employs a one-time watermark embedding to deceive unauthorized FR models and allows authorizers to perform identity verification by extracting the watermark. Specifically, we propose an information-guided adversarial attack against FR models. The encoder embeds an identity-specific watermark into the deep feature space of the carrier, guiding recognizable features of the image to deviate from the source identity. We further adopt a collaborative meta-optimization strategy compatible with sub-tasks, which regularizes the joint optimization direction of the encoder and decoder. This strategy enhances the representation of universal carrier features, mitigating multi-objective optimization conflicts in watermarking. Extensive experiments on two large-scale facial datasets demonstrate that DIP-Watermark achieves significant attack success rates and traceability accuracy on state-of-the-art FR models and commercial APIs. It also exhibits superior robustness against a wide range of real-world simulated distortions, outperforming existing privacy protection methods based on adversarial attacks, deep watermarking, or their simple combination. Our work potentially opens up new insights into proactive protection for FR privacy.
Yunming Zhang, Dengpan Ye, Caiyun Xie, Sipeng Shen, Ziyi Liu 0009, Jiacheng Deng 0001, Yueyun Shang, Zhihong Tian 0001
IEEE Trans. Dependable Secur. Comput.5
2026 Take Attention as Gate: An Associative Recurrent Network-Based Intrusion Detection Method for Industrial Control Network
abstract
The Industrial Control Network (ICN), which is characterized by real-time responsiveness and reliability, plays a key role in increasing production speed, ensuring efficient processing, and managing industrial processes. Despite tremendous advantages, ICN inevitably struggles with some challenges, such as malicious user intrusion and hacker attacks. To detect malicious intrusions in ICN, Intrusion Detection Systems (IDS) have been deployed. However, network traffic in ICN often exhibits significant temporal periodicity, and computational resources are limited on edge nodes and infrastructure gateway devices. These characteristics pose significant challenges to the design and performance of IDS. To properly solve these problems, we design a new intrusion detection method for ICN. Specifically, we first design a novel neural network model called Associative Recurrent Network (ARN), which can properly handle the relationship between previous hidden state and current input. Then, we construct a novel intrusion detection method based on the ARN, which avoids gating conflicts in traditional Recurrent Neural Network (RNN), effectively captures the temporal characteristics of ICN traffic, and maintains slightly higher computational overhead than GRU, thus demonstrating good adaptability to industrial control networks. Subsequently, through theoretical analysis of computational complexity, we demonstrate that the proposed method achieves high computational efficiency, comparable to mainstream RNN methods and superior to Transformer methods. Finally, we implement a prototype system to evaluate detection accuracy. Experimental results show that our method achieves state-of-the-art performance on the industrial control systems datasets (ICS-ADD and SWaT) and the conventional network dataset (UNSW-NB15), with average accuracies of 98.93%, 95.57%, and 98.27%, respectively.
Ziyi Liu 0009, Dengpan Ye, Yong Ding 0005, Yueling Liu, Chuanxi Chen
IEEE Trans. Netw. Serv. Manag.1
2025 PhonoFence: A Cross-Task Defense Framework for DeepFake via Phoneme-Level Adversarial Perturbations
Zhaolin Wei, Xiuwen Shi, Dengpan Ye, Jiacheng Deng 0001, Ziyi Liu 0009
ACM Multimedia7
2025 AdvLUT: Cloaking Geographic Location With Semantic-Based Adversarial 3-D Lookup Tables
abstract
The proliferation of Internet of Things (IoT) devices equipped with cameras, such as those in electric vehicles, has increased the collection of personal image data. However, the potential misuse of cross-view geo-localization (CVGL) models, which can infer precise locations from ground view images, has been overlooked and seriously threatens individual location privacy. In this article, we introduce AdvLUT, a novel semantic-based adversarial 3-D lookup tables (3DLUTs) privacy protection framework designed to safeguard geographic location privacy against CVGL models. The AdvLUT employs a geographic feature encoder to extract semantic features rich in geographic information from the ground view input. These features then guide a specialized adversarial 3DLUT generator in producing a 3DLUT that alters the color properties of the input image, thereby obstructing accurate location inference. Furthermore, AdvLUT is designed with a generative architecture that enables rapid image processing within milliseconds, eliminating the need for the corresponding satellite image or CVGL model. Experimental results on multiple benchmark datasets and CVGL models demonstrate that our method achieves up to a 65.48% reduction in R@1 localization accuracy, with performance further improving to 69.25% after JPEG compression.
Yiheng He, Dengpan Ye, Ziyi Liu 0009, Chuanxi Chen
IEEE Internet Things J.4
2025 Toward a Universal, Transferable, and Robust Adversarial Perturbation Framework Against Deep Hashing-Based Facial Image Retrieval
abstract
Deep Hashing (DH) based image retrieval is commonly used in facial recognition systems for its precision and effectiveness. However, this convenience is accompanied by a mounting threat to privacy. The DH model possesses vulnerability to adversarial attacks, which can be leveraged to prevent the retrieval of private images. Current adversarial attacks on DH models commonly focus on individual images or specific categories, lacking universal perturbations for the entire hashing dataset. This paper introduces the UTAP series, the first universal, transferable, and robust adversarial perturbation against DH facial image retrieval, safeguarding all images with a single perturbation. We explore the relationships between clusters learned by different DH models and define the optimization goal for optimizing UTAP series as moving away from the voted overall hashcenter. To alleviate the challenges of single-objective optimization, we randomly vote for sub-cluster centers and propose sub-task-based meta-learning to aid global optimization. Furthermore, we dissect the functional roles of key components in DH models and introduce UTAP++, a feature-hashing two-stage attack that is readily adaptable to cross-model and cross-scheme ensemble adversarial attacks. Extensive experiments conducted on renowned face datasets and DH models under varied complex scenarios, encompassing cross-image, cross-model, cross-bit, cross-algorithm, model ensemble, algorithm ensemble, and image compression, reveal that the UTAP series demonstrate remarkable universality, transferability, and robustness in preventing facial image retrieval. Compared to existing state-of-the-art methods, the UTAP series excel in white-box settings and exhibits significant transferability improvements of$10\%-70\%$in all black-box settings, with 20% and 55% average robustness improvements in white-box and black-box settings, respectively. These findings underscore the practical value of the UTAP series in real-world, presenting novel effective defense strategies against unauthorized facial image retrieval.
Yunna Lv, Dengpan Ye, Yiheng He, Ziyi Liu 0009, Caiyun Xie
IEEE Trans. Circuits Syst. Video Technol.5
2025 Towards Invisible Decision-Based Adversarial Attacks Against Visual Object Tracking
abstract
Adversarial attacks have become a critical focus in visual object tracking (VOT) research. Small, carefully crafted adversarial perturbations to video frames can easily disrupt the visual object tracker, leading to tracking failure. Therefore, studying adversarial attacks contributes to the development of more robust and reliable trackers. Considering that trackers are agnostic in real-world scenarios, research on decision-based black-box attacks is straightforward and practical. However, existing decision-based black-box attacks neither comprehensively analyze the unique characteristics of object tracking nor sufficiently consider the imperceptibility of adversarial perturbations. In this paper, we propose invisible local attack (ILA), a novel decision-based adversarial attack specifically for VOT with imperceptible perturbations. We assume that a significant number of pixels in a frame, irrelevant to the tracked object, do not substantially contribute to the functioning mechanism of a deep tracker. Based on this consideration, we propose a search algorithm to identify the pixel set focused on by the tracker during object tracking. The adversarial noise is then confined to these pixels and iteratively optimized through a heuristic algorithm of ILA. By perturbing only the key pixels, ILA significantly enhances both the attack performance and imperceptibility when it is applied to visual object trackers. Extensive experiments demonstrate that our ILA method achieves a 121% increase in the robustness metric and a 137% improvement in the structural similarity index measure (SSIM) across multiple datasets for various trackers compared with the state-of-the-art (SOTA) method.
Ziyi Liu 0009, Caiyun Xie, Wenbing Ding, Dengpan Ye, Qian Wang 0002
IEEE Trans. Multim.1
2025 The Interpretable and Transferable Adversarial Attack against Synthetic Speech Detectors
abstract
Existing work finds it challenging for adversarial examples to transfer among different synthetic speech detectors because of cross-feature and cross-model. To enhance the transferability of adversarial examples, we propose a spectral saliency analysis method and gain insight into the underlying detection mechanisms of existing detectors for the first time. These insights offer an interpretable basis for why adversarial examples are challenging to transfer between synthetic speech detection models. Then we further propose a two-stage adversarial attack framework. Specifically, the first stage leverages insights into the model detection mechanism to design a random time-frequency masking module, the random offset module, and 1D convolution to generate transferable and robust adversarial examples. In the second stage, to mitigate the problem of obvious noise in the low-energy frames of the carrier in existing adversarial attacks, we perform secondary optimization on frames below the Signal-Noise-Rate threshold to enhance its auditory quality. Extensive experimental results demonstrate that the proposed method significantly enhances the transferability and robustness of adversarial examples, while simultaneously preserving the acoustic quality compared to typical approaches.
Jiacheng Deng 0001, Dengpan Ye, Jizhi Li, Ziyi Liu 0009, Yunming Zhang
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Dual Defense: Adversarial, Traceable, and Invisible Robust Watermarking Against Face Swapping
abstract
Malicious applications of deep face swapping technology pose security threats such as misinformation dissemination and identity fraud. Some research propose the utilization of robust watermarking methods to track the copyright of facial images, facilitating post-forgery identity attribution. However, these methods cannot fundamentally prevent or eliminate the adverse impacts of face swapping. To address this issue, we present Dual Defense, an innovative framework based on robust adversarial watermarking. It simultaneously tracks image copyrights and disrupts the face swapping model by one-time embedding the robust adversarial watermark. Specifically, we propose an Original-domain Feature Emulation Attack (OFEA) method, which makes the traceable watermark adversarial through specially designed original domain adversarial loss. Additionally, we conduct a wavelet domain image structural information compensation loss, combined with a channel attention mechanism, to jointly balance watermark invisibility, adversariality, and traceability. Furthermore, we design a more comprehensive and rational evaluation method to thoroughly assess the effectiveness of adversarial attacks against face swapping models. Extensive experiments demonstrate that Dual Defense exhibits exceptional cross-task generality and dataset generalization. It maintains impressive adversariality and traceability in both original and robust settings, surpassing current forgery defense methods that possess only one of these capabilities.
Yunming Zhang, Dengpan Ye, Caiyun Xie, Xin Liao 0001, Ziyi Liu 0009, Chuanxi Chen, Jiacheng Deng 0001
IEEE Trans. Inf. Forensics Secur.6
2021 An Efficient Network Security Situation Assessment Method Based on AE and PMU
abstract
Network security situation assessment (NSSA) is an important and effective active defense technology in the field of network security situation awareness. By analyzing the historical network security situation awareness data, NSSA can evaluate the network security threat and analyze the network attack stage, thus fully grasping the overall network security situation. With the rapid development of 5G, cloud computing, and Internet of things, the network environment is increasingly complex, resulting in diversity and randomness of network threats, which directly determine the accuracy and the universality of NSSA methods. Meanwhile, the indicator data is characterized by large scale and heterogeneity, which seriously affect the efficiency of the NSSA methods. In this paper, we design a new NSSA method based on the autoencoder (AE) and parsimonious memory unit (PMU). In our novel method, we first utilize an AE‐based data dimensionality reduction method to process the original indicator data, thus effectively removing the redundant part of the indicator data. Subsequently, we adopt a PMU deep neural network to achieve accurate and efficient NSSA. The experimental results demonstrate that the accuracy and efficiency of our novel method are both greatly improved.
Xiaoling Tao, Ziyi Liu 0009
Wirel. Commun. Mob. Comput.2
2020 An Improved Parallel Network Traffic Anomaly Detection Method Based on Bagging and GRU
Xiaoling Tao, Feng Zhao 0002, Sufang Wang, Ziyi Liu 0009
WASA (1)5