VLDB 2026 Research / reviewers in the wild / expert
Yi-Xing Peng
dblp:318/9387
· DBLP profile ↗
18ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0003-4248-9820ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Negative Semantic Guided Identity Boundary Construction for Open-World Person Re-IdentificationabstractPerson re-identification (ReID) aims to match person images of the same identity under different camera views. Conventional ReID models mainly consider a closed-world setting where person identities in query and gallery are exactly the same. However, in real-world applications, query identities and gallery identities usually do not exactly contain the same persons. Therefore, open-world ReID has been proposed to match the images of gallery identities (targets) with a large number of non-gallery identities ( non-targets). Since some non-targets are quite similar to the targets, the ReID model may make incorrect judgments when verifying these non-targets. To solve this problem, we leverage the impressive cross-modal matching capabilities of the large vision-language model (VLM) to constructNegativeSemantic guided identity boundaries for each person to develop the open-worldReIDmodel (NS-ReID). To construct the identity boundary, we propose Virtual Non-target Repulsion that utilizes negative semantics to prompt the ReID model to push virtual non-targets away from the targets. The prompts expressing negative semantics offer a different perspective to guide the training process to avoid contradictory optimization. Moreover, we propose the Dual-Boost Refinement Learning strategy to train learnable identity prompts to capture detailed identity information, which is essential for constructing the identity boundary since the variations among identities are comparatively small. These facilitate the model in constructing the wide identity boundary of each person. Extensive experiments on two benchmark ReID datasets demonstrate that our proposed NS-ReID achieves state-of-the-art performance compared with existing methods. Xiao-Wen Zhang, Delong Zhang, Yi-Xing Peng, Jingke Meng, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Person De-reidentification: A Variation-guided Identity Shift ModelingabstractPerson re-identification (ReID) is to associate images of individuals from different camera views against cross-view variations. Like other surveillance technologies, Re-ID faces serious privacy challenges, particularly the potential for unauthorized tracking. Although various tasks (e.g., face recognition) have developed machine unlearning techniques to address privacy concerns, such methods have not yet been explored within the Re-ID field. In this work, we pioneer the exploration of the person de-reidentification (De-ReID) problem and present its inherent challenges. In the context of ReID, De-ReID is to unlearn the knowledge about accurately matching specific persons so that these "unlearned persons" cannot be re-identified across cameras for privacy guarantee. The primary challenge is to achieve the unlearning without degrading the identity-discriminative feature embeddings to ensure the model’s utility. To address this, we formulate a De-ReID framework that utilizes a labeled dataset of un-learned persons for unlearning and an unlabeled dataset of accessible persons for knowledge preservation. Instead of unlearning based on (pseudo) identity labels, we introduce a variation-guided identity shift mechanism that unlearns the specific persons by fitting the variations in their images while preserving ReID ability on other persons by overcoming the variations in images of accessible persons. As a result, the model shifts the unlearned persons to a feature space that is vulnerable to cross-view variations. Extensive experiments on benchmarks demonstrate the superiority of our method. Yi-Xing Peng, Yu-Ming Tang, Kun-Yu Lin, Qize Yang, Jingke Meng, Xihan Wei, Wei-Shi Zheng 0001 |
CVPR | 1 |
| 2025 | ViSpeak: Visual Instruction Feedback in Streaming Videos
Shenghao Fu, Qize Yang, Yuan-Ming Li, Yi-Xing Peng, Kun-Yu Lin, Xihan Wei, Jianfang Hu, Xiaohua Xie, Wei-Shi Zheng 0001 |
ICCV | 4 |
| 2025 | Viperson: Flexibly Generating Virtual Identity for Person Re-Identification
Xiao-Wen Zhang, Delong Zhang, Yi-Xing Peng, Zhi Ouyang, Jingke Meng, Wei-Shi Zheng 0001 |
ICCV | 3 |
| 2025 | Supplementary Material for "NoiseActor: A Noise-Action Collaborative Framework for Privacy-Preserving Action Recognition without Privacy Labels"abstractOur supplementary material is organized into the following sections: • Section II provides the details of the LOCATION module. • Section III provides the evaluation protocol for the SBU dataset. • Section IV provides the evaluation protocol for cross-dataset experiment on the UCF101 dataset and the VISPR dataset. • Section V provides implementation details for each datasets. • Section VI provides more anonymized frames on the SBU dataset. • Section VII provides more anonymized frames on the UCF101 dataset. • Section VIII provides details about anonymized video file. • Section IX provides ablation studies of the adapter. Xiao Li 0074, Xiao-Ming Wu 0002, Delong Zhang, Kun-Yu Lin, Yi-Xing Peng, Ling-An Zeng, Wei-Shi Zheng 0001 |
ICME | 5 |
| 2025 | Protecting Feature Privacy in Person Re-IdentificationabstractPerson re-identification (ReID) is to identify the same person across non-overlapping camera views. After a decade of development, the methods based on deep networks have achieved high performance on benchmarks and become mainstream. In applications, the features of gallery images extracted by deep learning-based methods are stored to speed up the query process and protect the sensitive information contained in the images. Unfortunately, it is demonstrated that turning the images into features cannot properly protect privacy, as these features could be reversed to the corresponding images, revealing the sensitive information they contain. Therefore, for preventing privacy leakage, recent methods learn their features against some feature reversal methods, and most conventional reversal methods focus on minimizing the difference between a reconstruction and its original image. However, there could be many reasonable reconstruction results from a single feature, and the conventional reversal methods will inevitably generate reconstruction results that lie in a different distribution from one of the original images, which cannot properly assess the private information for learning to protect and thus hamper the privacy-protected feature learning. To mitigate this problem, we enforce the reconstructions to follow the same distribution as the original images by the generative adversarial network (GAN). We operate this GAN-based feature reversal module accompanied by the conventional ReID feature extraction module and form a novel GAN-based feature privacy-protected person ReID model, which is expected to protect feature privacy so as against reversal attack and maintain ReID utility. We demonstrate that optimizing ReID model to accommodate privacy protection faces a double adversarial objective and is thus challenging. As a remedy, we design a novel two-step training and lazy update strategy that alternatively optimizes the feature extraction module and stabilizes the update process of the GAN-based feature reversal module. To evaluate the efficiency of the model in balancing its ReID utility and feature privacy protection, we introduce a novel metric called utility-reversibility ratio (URR). Compared with existing privacy-protected feature extraction models, the proposed method achieves a better balance between privacy protection and person ReID performance. Extensive experiments validate that our model can effectively protect feature privacy at a tiny accuracy cost, and validate the effectiveness of our model with the emerging diffusion model. Xiao Li 0074, Yi-Xing Peng, Wei-Shi Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | A Versatile Framework for Multi-Scene Person Re-IdentificationabstractPerson Re-identification (ReID) has been extensively developed for a decade in order to learn the association of images of the same person across non-overlapping camera views. To overcome significant variations between images across camera views, mountains of variants of ReID models were developed for solving a number of challenges, such as resolution change, clothing change, occlusion, modality change, and so on. Despite the impressive performance of many ReID variants, these variants typically function distinctly and cannot be applied to other challenges. To our best knowledge, there is no versatile ReID model that can handle various ReID challenges at the same time. This work contributes to the first attempt at learning a versatile ReID model to solve such a problem. Our main idea is to form a two-stage prompt-based twin modeling framework called VersReID. Our VersReID firstly leverages the scene label to train a ReID Bank that contains abundant knowledge for handling various scenes, where several groups of scene-specific prompts are used to encode different scene-specific knowledge. In the second stage, we distill a V-Branch model with versatile prompts from the ReID Bank for adaptively solving the ReID of different scenes, eliminating the demand for scene labels during the inference stage. To facilitate training VersReID, we further introduce the multi-scene properties into self-supervised learning of ReID via a multi-scene prioris data augmentation (MPDA) strategy. Through extensive experiments, we demonstrate the success of learning an effective and versatile ReID model for handling ReID tasks under multi-scene conditions without manual assignment of scene labels in the inference stage, including general, low-resolution, clothing change, occlusion, and cross-modality scenes. Wei-Shi Zheng 0001, Junkai Yan, Yi-Xing Peng |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Distilling consistent relations for multi-source domain adaptive person re-identification
Yuqiao Xian, Yi-Xing Peng, Xing Sun 0001, Wei-Shi Zheng 0001 |
Pattern Recognit. | 2 |
| 2024 | Rethinking Few-Shot Class-Incremental Learning: Learning from Yourself
Yu-Ming Tang, Yi-Xing Peng, Jingke Meng, Wei-Shi Zheng 0001 |
ECCV (61) | 2 |
| 2024 | Patch-Based Privacy Attention for Weakly-Supervised Privacy-Preserving Action RecognitionabstractPrivacy-preserving action recognition aims to prevent privacy leakage by learning to anonymize video frames for action recognition. However, apart from video-level action labels, supervised methods require costly frame-level privacy labels. To achieve privacy-preserving action recognition in the absence of privacy labels, weakly-supervised privacy-preserving action recognition proposes to merely utilize action labels for learning without using any privacy labels during training, and therefore the challenge lies in removing the privacy information without privacy annotations. Inspired by the fact that private information such as the identity of the participant is not significant to action recognition, our main idea is to utilize the attention mechanism to automatically discover the action-sensitive information in frames and remove the other information to prevent privacy leakage without using the privacy labels. Our method contains a novel patch-based privacy attention module and an action recognition module. The patch-based privacy attention module splits raw frames into patches and exploits self-attention among patches to adaptively discover action-sensitive but privacy-less information. Our patch-based privacy attention module mines the action-sensitive information from both individual frames and adjacent frames to generate anonymized frames. A distance correlation loss is introduced to enforce that the generated anonymized frames are distinguished from the original frames and contain less private information. In addition, the action recognition module learns to recognize actions based on anonymized frames. Extensive experiments demonstrate that our model can effectively alleviate privacy leakage and maintain the performance of action recognition without using privacy labels. Xiao Li 0074, Yukun Qiu, Yi-Xing Peng, Wei-Shi Zheng 0001 |
FG | 3 |
| 2024 | PixelFade: Privacy-preserving Person Re-identification with Noise-guided Progressive Replacement
Delong Zhang, Yi-Xing Peng, Xiao-Ming Wu 0002, Ancong Wu, Wei-Shi Zheng 0001 |
ACM Multimedia | 2 |
| 2024 | Privacy-Preserving Action Recognition: A Survey
Xiao Li 0074, Yukun Qiu, Yi-Xing Peng, Ling-An Zeng, Wei-Shi Zheng 0001 |
PRCV (7) | 3 |
| 2024 | Revisiting Person Re-Identification by Camera SelectionabstractPerson re-identification (Re-ID) is a fundamental task in visual surveillance. Given a query image of the target person, conventional Re-ID focuses on the pairwise similarities between the candidate images and the query. However, conventional Re-ID does not evaluate the consistency of the retrieval results of whether the most similar images ranked in each place contain the same person, which is risky in some applications such as missing out a place where the patient passed will hinder the epidemiological investigation. In this work, we investigate a more challenging task: consistently and successfully retrieving the target person in all camera views. We define the task as continuous person Re-ID and propose a corresponding evaluation metric termed overall Rank-K accuracy. Different from the conventional Re-ID, any incorrect retrieval under an individual camera view that raises an inconsistency will fail the continuous Re-ID. Consequently, the defective cameras, in which the images are hard to be automatically associated with the images from other views, strongly degrade the performance of continuous person Re-ID. Since the camera deployment is crucial for continuous tracking across camera views, we rethink person Re-ID from the perspective of camera deployment and assess the quality of a camera network by performing continuous Re-ID. Moreover, we propose to automatically detect the defective cameras that greatly hamper the continuous Re-ID. Because brute-force search is costly when the camera network becomes complicated, we explicitly model the visual relations as well as the spatial relations among cameras and develop a relational deep Q-network to select the properly deployed cameras and the un-selected cameras are regarded as the defective cameras. Since most existing datasets do not provide topology information about the camera network, they are unsuitable for investigating the importance of spatial relations on camera selection. Thus, we collect a new dataset including 20 cameras with topology information. Compared with randomly removing cameras, the experimental results show that our method can effectively detect the defective cameras so that people could take further operations on these cameras in practice (https://www.isee-ai.cn/∼yixing/MCCPD.html). Yi-Xing Peng, Yuanxun Li, Wei-Shi Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Generalized Intra-Camera Supervised Person Re-IdentificationabstractPerson re-identification (Re-ID) is to match the images of the same person from different camera views, which demands a view-invariant feature embedding. Recently, intra-camera supervised (ICS) Re-ID develops the Re-ID models without cross-view annotated data. Existing ICS methods are developed based on the assumptions, such as assuming each person in the training set appears under multiple cameras. However, there is no guarantee that the assumptions are true without cross-view annotations, and their performance degrades when the assumptions are violated. In this work, we generalize the ICS Re-ID and develop an ICS Re-ID model without the assumptions. The absence of prior assumptions and cross-view annotations poses a challenge in exploiting the discriminative information among cross-view images. To this end, we propose to mine the view-invariant relations between cross-view images for Re-ID model to exploit discriminative information and overcome the cross-view variations. Specifically, we learn composited view-aware features by compositing the identity information with different camera view information in the feature composition module. Then, we exploit the composited features to model various view-aware relations between pairwise images. By mining the common patterns among the view-aware relations, we obtain the view-invariant pairwise relation for learning. Besides, leveraging the composited view-aware features, we develop a view-aware marginal constraint for robust cross-view learning. To facilitate learning the feature composition module, we augment an auxiliary network to exploit the camera view information at the feature level. Extensive experimental results show the effectiveness of our method under different scenarios. Yi-Xing Peng, Yu-Ming Tang, Kun-Yu Lin, Wei-Shi Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Cross-Modal Adaptive Dual Association for Text-to-Image Person RetrievalabstractText-to-image person re-identification (ReID) aims to retrieve images of a person based on a given textual description. The key challenge is to learn the relations between detailed information from visual and textual modalities. Existing work focuses on learning a latent space to narrow the modality gap and further build local correspondences between two modalities. However, these methods assume that image-to-text and text-to-image associations are modality-agnostic, resulting in suboptimal associations. In this work, we demonstrate the discrepancy between image-to-text association and text-to-image association and proposecross-modal adaptive dual association (CADA) to build fine bidirectional image-text detailed associations. Our approach features a decoder-based adaptive dual association module that enables full interaction between visual and textual modalities, enabling bidirectional and adaptive cross-modal correspondence associations. Specifically, this paper proposes a bidirectional association mechanism: Association of text Tokens to image Patches (ATP) and Association of image Regions to text Attributes (ARA). We adaptively model the ATP based on the fact that aggregating cross-modal features based on mistaken associations will lead to feature distortion. For modeling the ARA, since attributes are typically the first distinguishing cues of a person, we explore attribute-level associations by predicting the masked text phrase using the related image region. Finally, we learn the dual associations between texts and images, and the experimental results demonstrate the superiority of our dual formulation. The code used in this article will be made publicly available athttps://github.com/LinDixuan/CADA. Dixuan Lin, Yi-Xing Peng, Jingke Meng, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | When Prompt-based Incremental Learning Does Not Meet Strong PretrainingabstractIncremental learning aims to overcome catastrophic forgetting when learning deep networks from sequential tasks. With impressive learning efficiency and performance, prompt-based methods adopt a fixed backbone to sequential tasks by learning task-specific prompts. However, existing prompt-based methods heavily rely on strong pretraining (typically trained on ImageNet-21k), and we find that their models could be trapped if the potential gap between the pretraining task and unknown future tasks is large. In this work, we develop a learnable Adaptive Prompt Generator (APG). The key is to unify the prompt retrieval and prompt learning processes into a learnable prompt generator. Hence, the whole prompting process can be optimized to reduce the negative effects of the gap between tasks effectively. To make our APG avoid learning ineffective knowledge, we maintain a knowledge pool to regularize APG with the feature distribution of each class. Extensive experiments show that our method significantly outperforms advanced methods in exemplar-free incremental learning without (strong) pretraining. Besides, under strong pretraining, our method also has comparable performance to existing prompt-based models, showing that our method can still benefit from pretraining. Codes can be found at https://github.com/TOM-tym/APG Yu-Ming Tang, Yi-Xing Peng, Wei-Shi Zheng 0001 |
ICCV | 2 |
| 2023 | Consistent Discrepancy Learning for Intra-Camera Supervised Person Re-IdentificationabstractSince annotating pedestrians across different views is extremely costly, intra-camera supervised person re-identification (ReID) aims to learn a ReID model from the intra-view labeled data. Under this setting, the most challenge lies in learning a view-invariant feature embedding in the absence of the cross-view annotations. Previous works focus on assigning a pseudo identity label for each image based on the feature similarity and learn view-invariant features by classification loss. However, because of the cross-view variations in lighting, background, etc., the pseudo labels are often noisy, and therefore not reliable for classification. In this paper, we explore learning a consistent discrepancy for pairwise images. Our main idea is that the discrepancy between pedestrian images should be consistent across different views regardless of view change so that it mainly depicts the identity difference. Due to the lack of cross-view annotations, we project images into different views and obtain likelihood prototypes for cross-view learning. These likelihood prototypes are used to measure the discrepancies between pairwise images under different views. And then, we propose an intra-view discrepancy preservation module to enforce the discrepancy to be view-consistent so as to encourage the model to distinguish the images based on the identities regardless of view change. Extensive experiments on multiple datasets show that our method outperforms existing related methods by clear margins and our method is comparable to supervised counterparts. Code will be made publicly available. Yi-Xing Peng, Jile Jiao, Xuetao Feng, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Learning to Imagine: Diversify Memory for Incremental Learning using Unlabeled DataabstractDeep neural network (DNN) suffers from catastrophic forgetting when learning incrementally, which greatly limits its applications. Although maintaining a handful of samples (called “exemplars”) of each task could alleviate forgetting to some extent, existing methods are still limited by the small number of exemplars since these exemplars are too few to carry enough task-specific knowledge, and therefore the forgetting remains. To overcome this problem, we propose to “imagine” diverse counterparts of given exemplars referring to the abundant semantic-irrelevant information from unlabeled data. Specifically, we develop a learnable feature generator to diversify exemplars by adaptively generating diverse counterparts of exemplars based on semantic information from exemplars and semantically-irrelevant information from unlabeled data. We introduce semantic contrastive learning to enforce the generated samples to be semantic consistent with exemplars and perform semantic-decoupling contrastive learning to encourage diversity of generated samples. The diverse generated samples could effectively prevent DNN from forgetting when learning new tasks. Our method does not bring any extra inference cost and outperforms state-of-the-art methods on two benchmarks CIFAR-100 and ImageNet-Subset by a clear margin. Yu-Ming Tang, Yi-Xing Peng, Wei-Shi Zheng 0001 |
CVPR | 2 |