Zipeng Ye

dblp:190/1787 · DBLP profile ↗
← Back
20ranked-venue papers
11as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 5 since 2021Security and privacy · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Reconstruction Attack-Resistant Inference Paradigm for LLM Cloud Services
abstract
Large language models (LLMs) have seen remarkable growth in recent years. To leverage convenient LLM cloud services, users are inevitably to upload their prompts. Additionally, for tasks such as translation, reading comprehension, and summarization, associated files or context are inherently needed, whether or not they contain user privacy information. Despite the rapid progress in LLM capabilities, research on preserving user privacy during inference has been relatively scarce. To this end, this paper conducts some exploratory research in this domain. Firstly, we show that (1) the embedding space of tokens is highly sparse, and (2) LLMs primarily function in the orthogonal subspace of embedding space, these two factors making privacy extremely vulnerable. Then, we analyze the structural characteristics of LLMs and design a distributed privacy-preserving inference paradigm which can effectively resist privacy attacks. Finally, we perform a thorough evaluation of the defended models on mainstream tasks and find that low-bit quantization techniques can be effectively combined with our inference paradigm, achieving a balance between privacy, utility, and runtime memory efficiency.
Zipeng Ye, Wenjian Luo, Yubo Tang
AAAI1
2025 MIR: Efficient Exploration in Episodic Multi-agent Reinforcement Learning via Mutual Intrinsic Reward
Kesheng Chen, Wenjian Luo, Bang Zhang, Zeping Yin, Zipeng Ye
ICIC (14)5
2025 NID: A privacy-preserving operator based on Neural Information Diffusion
Muhammad Luqman Naseem, Zipeng Ye, Wenjian Luo
J. Inf. Secur. Appl.2
2025 Gradient Inversion of Text-Modal Data in Distributed Learning
abstract
Gradient inversion attacks (GIAs) pose significant challenges to the privacy-preserving paradigm of distributed learning. These attacks employ carefully designed strategies to reconstruct victim’s private training data from their shared gradients. However, existing work mainly focuses on attacks and defenses for image-modal data, while the study for text-modal data remains scarce. Furthermore, the performance of the limited attack researches on text-modal data is also unsatisfactory, which can be partially attributed to the finer granularity of text data compared to image. To bridge the existing research gap, we propose a high-fidelity attack method tailored for Transformer-based language models (LMs). In our method, we initially reconstruct the label space of the victim’s training data by leveraging the characteristics of the Transformer architecture. After that, we propose a shallow-to-deep paradigm to facilitate gradient matching, which can significantly improve the attack performance. Furthermore, we develop a weighted surrogate loss that resolves the consistent deviation issue present in current attack researches. A substantial number of experiments on Transformer-based LMs (e.g., Bert and GPT) demonstrate that our attack is competitive and significantly outperforms existing methods. In the final part of this paper, we investigate the influence of the inherent position embedding module within the Transformer architecture on attack performance, and based on the analysis results, we propose a countermeasure to alleviate part of the privacy leakage issue in distributed learning.
Zipeng Ye, Wenjian Luo, Yubo Tang, Zhenqian Zhu, Yuhui Shi 0001, Yan Jia 0001
IEEE Trans. Inf. Forensics Secur.1
2024 High-Fidelity Gradient Inversion in Distributed Learning
abstract
Distributed learning frameworks aim to train global models by sharing gradients among clients while preserving the data privacy of each individual client. However, extensive research has demonstrated that these learning frameworks do not absolutely ensure the privacy, as training data can be reconstructed from shared gradients. Nevertheless, the existing privacy-breaking attack methods have certain limitations. Some are applicable only to small models, while others can only recover images in small batch size and low resolutions, or with low fidelity. Furthermore, when there are some data with the same label in a training batch, existing attack methods usually perform poorly. In this work, we successfully address the limitations of existing attacks by two steps. Firstly, we model the coefficient of variation (CV) of features and design an evolutionary algorithm based on the minimum CV to accurately reconstruct the labels of all training data. After that, we propose a stepwise gradient inversion attack, which dynamically adapts the objective function, thereby effectively and rationally promoting the convergence of attack results towards an optimal solution. With these two steps, our method is able to recover high resolution images (224*224 pixel, from ImageNet and Web) with high fidelity in distributed learning scenarios involving complex models and larger batch size. Experiment results demonstrate the superiority of our approach, reveal the potential vulnerabilities of the distributed learning paradigm, and emphasize the necessity of developing more secure mechanisms. Source code is available at https://github.com/MiLab-HITSZ/2023YeHFGradInv.
Zipeng Ye, Wenjian Luo, Yubo Tang
AAAI1
2024 Data-Free Backdoor Model Inspection: Masking and Reverse Engineering Loops for Feature Counting
abstract
Deep Neural Networks (DNNs) are widely used for the outstanding performance in many fields. However, the training of DNN models has high requirements for the users’ data and computation resources, so many users with limited resources tend to download pre-trained models from some platforms and then finetune the pre-trained models to match their own tasks. However, the pre-trained models are under the threat of the backdoor attack. The backdoor attackers inject backdoors in the models, leading the backdoor models to predict target predictions designed by the attackers in advance. However, most existing backdoor model inspection methods rely on the clean data samples from the dataset of the model, which are difficult to get for users who just download the pre-trained models from the platforms. There are also a few defense methods not dependent on the data, but they also have their limits in practice. We propose Data-Free Masking and Reverse Engineering Loops (DF-MREL), a simple yet efficient data-free method for backdoor model inspection, which is widely applicable when resources are limited. Our experiments show its excellent performance in detecting backdoor models. Source code will be published after accepted.
Wenjian Luo, Zipeng Ye, Yubo Tang
IJCNN3
2024 Gradient Inversion Attacks: Impact Factors Analyses and Privacy Enhancement
abstract
Gradient inversion attacks (GIAs) have posed significant challenges to the emerging paradigm of distributed learning, which aims to reconstruct the private training data of clients (participating parties in distributed training) through the shared parameters. For counteracting GIAs, a large number of privacy-preserving methods for distributed learning scenario have emerged. However, these methods have significant limitations, either compromising the usability of global model or consuming substantial additional computational resources. Furthermore, despite the extensive efforts dedicated to defense methods, the underlying causes of data leakage in distributed learning still have not been thoroughly investigated. Therefore, this paper tries to reveal the potential reasons behind the successful implementation of existing GIAs, explore variations in the robustness of models against GIAs during the training process, and investigate the impact of different model structures on attack performance. After these explorations and analyses, this paper propose a plug-and-play GIAs defense method, which augments the training data by a designed vicinal distribution. Sufficient empirical experiments demonstrate that this easy-to-implement method can ensure the basic level of privacy without compromising the usability of global model.
Zipeng Ye, Wenjian Luo, Zhenqian Zhu, Yuhui Shi 0001, Yan Jia 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 C2FMI: Corse-to-Fine Black-Box Model Inversion Attack
abstract
Privacy-preserving machine learning requires that models do not reveal any private information about their training data. However, model inversion attacks (MIAs), which aim to recover the features of training data, pose a huge threat to the security of AI models. Most existing MIAs assume that the target model is white-box, but most models deployed in reality are black-box, and these models can only be accessed like an oracle. There are a few studies for black-box scenarios, but their performance is limited. In this paper, we firstly formulate the MIA problem completely in Bayesian perspective. Second, we propose a novel two-stage MIA approach, the Coarse-to-Fine Model Inversion Attack (C2FMI), which efficiently addresses the MIA problem in the black-box scenario. In stage I of C2FMI, we design a reverse network that constrains the recovered images (also named attacked images) to fall near the manifold space of the training data. In stage II, we design a black-box oriented strategy which further facilitates the attacked images to approach the training data. Empirically, C2FMI achieves a performance that even surpasses existing white-box attack methods. Furthermore, we design the stability analysis method for analyzing the stability of C2FMI along with existing MIAs. Finally, we explore the potential countermeasures which could defend against our attacks.
Zipeng Ye, Wenjian Luo, Muhammad Luqman Naseem, Xiangkai Yang, Yuhui Shi 0001, Yan Jia 0001
IEEE Trans. Dependable Secur. Comput.1
2024 An Evolutionary Attack for Revealing Training Data of DNNs With Higher Feature Fidelity
abstract
Model inversion attacks aim to reveal information about sensitive training data of AI models, which may lead to serious privacy leakage. However, existing attack methods have limitations in reconstructing training data with higher feature fidelity. In this paper, we propose an evolutionary model inversion attack approach (EvoMI) and empirically demonstrate that combined with the systematic search in the multi-degree-of-freedom latent space of the generative model, the simple use of an evolutionary algorithm can effectively improve the attack performance. Concretely, at first, we search for latent vectors which can generate images close to the attack target in the latent space with low-degree of freedom. Generally, the low-freedom constraint will reduce the probability of getting a local optima compared to existing methods that directly search for latent vectors in the high-freedom space. Consequently, we introduce a mutation operation to expand the search domain, thus further reduce the possibility of obtaining a local optima. Finally, we treat the searched latent vectors as the initial values of the post-processing and relax the constraint to further optimize the latent vectors in a higher-freedom space. Our proposed method is conceptually simple and easy to implement, yet it achieves substantial improvements and outperforms the state-of-the-art methods significantly.
Zipeng Ye, Wenjian Luo, Ruizhuo Zhang, Yuhui Shi 0001, Yan Jia 0001
IEEE Trans. Dependable Secur. Comput.1
2023 Audio-Driven Talking Face Video Generation With Dynamic Convolution Kernels
abstract
In this paper, we present a dynamic convolution kernel (DCK) strategy for convolutional neural networks. Using a fully convolutional network with the proposed DCKs, high-quality talking-face video can be generated from multi-modal sources (i.e., unmatched audio and video) in real time, and our trained model is robust to different identities, head postures, and input audios. Our proposed DCKs are specially designed for audio-driven talking face video generation, leading to a simple yet effective end-to-end system. We also provide a theoretical analysis to interpret why DCKs work. Experimental results show that our method can generate high-quality talking-face video with background at 60 fps. Comparison and evaluation between our method and the state-of-the-art methods demonstrate the superiority of our method.
Zipeng Ye, Mengfei Xia, Ran Yi 0002, Juyong Zhang, Yukun Lai, Xuwei Huang, Guo-Xin Zhang, Yong-Jin Liu 0001
IEEE Trans. Multim.1
2023 Predicting Personalized Head Movement From Short Video and Speech Signal
abstract
Audio-driven talking face video generation has attracted much attention recently. However, few existing works pay attention to machine learning of talking head movement, especially based on the phonetic study. Observing that real-world talking faces often accompany natural head movement, in this paper, we model the relation between speech signal and talking head movement, which is a typical one-to-many mapping problem. To solve this problem, we propose a novel two-step mapping strategy: (1) in the first step, we train an encoder that predicts a head motion behavior pattern (modeled as a feature vector) from the head motion sequence of a short video of 10–15 seconds, and (2) in the second step, we train a decoder that predict a unique head motion sequence from both the motion behavior pattern and the auditory features of an arbitrary speech signal. Based on the proposed mapping strategy, we build a deep neural network model that takes a speech signal of a source person and a short video of a target person as input, and outputs a synthesized high-fidelity talking face video with personalized head pose. Extensive experiments and a user study show that our method can generate high-quality personalized head movement in synthesized talking face videos, and meanwhile, has comparable facial animation quality (e.g., lip synchronization and expression) with the state-of-the-art methods.
Ran Yi 0002, Zipeng Ye, Zhiyao Sun, Juyong Zhang, Guo-Xin Zhang, Pengfei Wan 0001, Hujun Bao, Yong-Jin Liu 0001
IEEE Trans. Multim.2
2023 3D-CariGAN: An End-to-End Solution to 3D Caricature Generation From Normal Face Photos
abstract
Caricature is a type of artistic style of human faces that attracts considerable attention in the entertainment industry. So far a few 3D caricature generation methods exist and all of them require some caricature information (e.g., a caricature sketch or 2D caricature) as input. This kind of input, however, is difficult to provide by non-professional users. In this paper, we propose an end-to-end deep neural network model that generates high-quality 3D caricatures directly from a normal 2D face photo. The most challenging issue for our system is that the source domain of face photos (characterized by normal 2D faces) is significantly different from the target domain of 3D caricatures (characterized by 3D exaggerated face shapes and textures). To address this challenge, we: (1) build a large dataset of 5,343 3D caricature meshes and use it to establish a PCA model in the 3D caricature shape space; (2) reconstruct a normal full 3D head from the input face photo and use its PCA representation in the 3D caricature shape space to establish correspondences between the input photo and 3D caricature shape; and (3) propose a novel character loss and a novel caricature loss based on previous psychological studies on caricatures. Experiments including a novel two-level user study show that our system can generate high-quality 3D caricatures directly from normal face photos.
Zipeng Ye, Mengfei Xia, Yanan Sun 0006, Ran Yi 0002, Minjing Yu, Juyong Zhang, Yukun Lai, Yong-Jin Liu 0001
IEEE Trans. Vis. Comput. Graph.1
2022 Generating Smooth and Facial-Details-Enhanced Talking Head Video: A Perspective of Pre and Post Processes
abstract
Talking head video generation has received increasing attention recently. So far the quality (especially the facial details) of the videos output from state-of-the-art deep learning methods is limited by either the quality of training data or the performance of generators, and needs to be further improved. In this paper, we propose a data pre- and post- processing strategy based on a key observation: generating talking head video from multi-modal input is a challenging problem and generating smooth video with fine facial details makes the problem even harder. Then we propose to decompose the problem solution into a main deep model, a pre- and a post- processing. The main deep model generates a reasonably good talking face video, with the aid of a pre-process, which also contributes to a post-process for restoring smooth and fine facial details in the final video. In particular, our main deep model reconstructs a 3D face from an input reference frame, and then uses an AudioNet to generate a sequence of facial expression coefficients with an input audio clip. To ensure final facial details in the generated video, we sample the original texture from the reference frame in the pre-process with the aid of reconstructed 3D face and a predefined UV map. Accordingly, in the post-process, we smooth the expression coefficients of adjacent frames to alleviate jitters and apply a pretrained face restoration module to recover the fine facial details. Experimental results and ablation study show the advantage of our proposed method.
Tian Lv, Yu-Hui Wen, Zhiyao Sun, Zipeng Ye, Yong-Jin Liu 0001
ACM Multimedia4
2022 GAN-Based Multi-Style Photo Cartoonization
abstract
Cartoon is a common form of art in our daily life and automatic generation of cartoon images from photos is highly desirable. However, state-of-the-art single-style methods can only generate one style of cartoon images from photos and existing multi-style image style transfer methods still struggle to produce high-quality cartoon images due to their highly simplified and abstract nature. In this article, we propose a novel multi-style generative adversarial network (GAN) architecture, called MS-CartoonGAN, which can transform photos into multiple cartoon styles. MS-CartoonGAN uses only unpaired photos and cartoon images of multiple styles for training. To achieve this, we propose to use (1) a hierarchical semantic loss with sparse regularization to retain semantic content and recover flat shading in different abstract levels, (2) a new edge-promoting adversarial loss for producing fine edges, and (3) a style loss to enhance the difference between output cartoon styles and make training process more stable. We also develop a multi-domain architecture, where the generator consists of a shared encoder and multiple decoders for different cartoon styles, along with multiple discriminators for individual styles. By observing that cartoon images drawn by different artists have their unique styles while sharing some common characteristics, our shared network architecture exploits the common characteristics of cartoon styles, achieving better cartoonization and being more efficient than single-style cartoonization. We show that our multi-domain architecture can theoretically guarantee to output desired multiple cartoon styles. Through extensive experiments including a user study, we demonstrate the superiority of the proposed method, outperforming state-of-the-art single-style and multi-style image style transfer methods.
Yezhi Shu, Ran Yi 0002, Mengfei Xia, Zipeng Ye, Wang Zhao 0001, Yukun Lai, Yong-Jin Liu 0001
IEEE Trans. Vis. Comput. Graph.4
2021 Feature-Aware Uniform Tessellations on Video Manifold for Content-Sensitive Supervoxels
abstract
Over-segmenting a video into supervoxels has strong potential to reduce the complexity of downstream computer vision applications. Content-sensitive supervoxels (CSSs) are typically smaller in content-dense regions (i.e., with high variation of appearance and/or motion) and larger in content-sparse regions. In this paper, we propose to compute feature-aware CSSs (FCSSs) that are regularly shaped 3D primitive volumes well aligned with local object/region/motion boundaries in video. To compute FCSSs, we map a video to a 3D manifold embedded in a combined color and spatiotemporal space, in which the volume elements of video manifold give a good measure of the video content density. Then any uniform tessellation on video manifold can induce CSS in the video. Our idea is that among all possible uniform tessellations on the video manifold, FCSS finds one whose cell boundaries well align with local video boundaries. To achieve this goal, we propose a novel restricted centroidal Voronoi tessellation method that simultaneously minimizes the tessellation energy (leading to uniform cells in the tessellation) and maximizes the average boundary distance (leading to good local feature alignment). Theoretically our method has an optimal competitive ratio O(1), and its time and space complexities are O(NK) and O(N+K) for computing K supervoxels in an N-voxel video. We also present a simple extension of FCSS to streaming FCSS for processing long videos that cannot be loaded into main memory at once. We evaluate FCSS, streaming FCSS and ten representative supervoxel methods on four video datasets and two novel video applications. The results show that our method simultaneously achieves state-of-the-art performance with respect to various evaluation criteria.
Ran Yi 0002, Zipeng Ye, Wang Zhao 0001, Minjing Yu, Yukun Lai, Yong-Jin Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 Dirichlet energy of Delaunay meshes and intrinsic Delaunay triangulations
Zipeng Ye, Ran Yi 0002, Wen-Yong Gong, Ying He 0001, Yong-Jin Liu 0001
Comput. Aided Des.1
2020 Ranking-Preserving Cross-Source Learning for Image Retargeting Quality Assessment
abstract
Image retargeting techniques adjust images into different sizes and have attracted much attention recently. Objective quality assessment (OQA) of image retargeting results is often desired to automatically select the best results. Existing OQA methods train a model using some benchmarks (e.g., RetargetMe), in which subjective scores evaluated by users are provided. Observing that it is challenging even for human subjects to give consistent scores for retargeting results of different source images (diff-source-results), in this paper we propose a learning-based OQA method that trains a General Regression Neural Network (GRNN) model based on relative scores-which preserve the ranking-of retargeting results of the same source image (same-source-results). In particular, we develop a novel training scheme with provable convergence that learns a common base scalar for same-source-results. With this source specific offset, our computed scores not only preserve the ranking of subjective scores for same-source-results, but also provide a reference to compare the diff-source-results. We train and evaluate our GRNN model using human preference data collected in RetargetMe. We further introduce a subjective benchmark to evaluate the generalizability of different OQA methods. Experimental results demonstrate that our method outperforms ten representative OQA methods in ranking prediction and has better generalizability to different datasets.
Yong-Jin Liu 0001, Yiheng Han, Zipeng Ye, Yukun Lai
IEEE Trans. Pattern Anal. Mach. Intell.3
2019 Fast Computation of Content-Sensitive Superpixels and Supervoxels Using Q-Distances
abstract
State-of-the-art researches model the data of images and videos as low-dimensional manifolds and generate superpixels/supervoxels in a content-sensitive way, which is achieved by computing geodesic centroidal Voronoi tessellation (GCVT) on manifolds. However, computing exact GCVTs is slow due to computationally expensive geodesic distances. In this paper, we propose a much faster queue-based graph distance (called q-distance). Our key idea is that for manifold regions in which q-distances are different from geodesic distances, GCVT is prone to placing more generators in them, and therefore after few iterations, the q-distance-induced tessellation is an exact GCVT. This idea works well in practice and we also prove it theoretically under moderate assumption. Our method is simple and easy to implement. It runs 6-8 times faster than state-of-the-art GCVT computation, and has an optimal approximation ratio O(1) and a linear time complexity O(N) for N-pixel images or N-voxel videos. A thorough evaluation of 31 superpixel methods on five image datasets and 8 supervoxel methods on four video datasets shows that our method consistently achieves the best over-segmentation accuracy. We also demonstrate the advantage of our method on one image and two video applications.
Zipeng Ye, Ran Yi 0002, Minjing Yu, Yong-Jin Liu 0001, Ying He 0001
ICCV1
2019 DE-Path: A Differential-Evolution-Based Method for Computing Energy-Minimizing Paths on Surfaces
Zipeng Ye, Yong-Jin Liu 0001, Jianmin Zheng, Kai Hormann, Ying He 0001
Comput. Aided Des.1
2019 LineUp: Computing Chain-Based Physical Transformation
abstract
In this article, we introduce a novel method that can generate a sequence of physical transformations between 3D models with different shape and topology. Feasible transformations are realized on a chain structure with connected components that are 3D printed. Collision-free motions are computed to transform between different configurations of the 3D printed chain structure. To realize the transformation between different 3D models, we first voxelize these input models into a similar number of voxels. The challenging part of our approach is to generate a simple path—as a chain configuration to connect most voxels. A layer-based algorithm is developed with theoretical guarantee of the existence and the path length. We find that collision-free motion sequence can always be generated when using a straight line as the intermediate configuration of transformation. The effectiveness of our method is demonstrated by both the simulation and the experimental tests taken on 3D printed chains.
Minjing Yu, Zipeng Ye, Yong-Jin Liu 0001, Ying He 0001, Charlie C. L. Wang
ACM Trans. Graph.2