VLDB 2026 Research / reviewers in the wild / expert
Pengcheng Zhang 0003
dblp:84/6775-3
· DBLP profile ↗
13ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0001-8585-7545ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual-label noise filtering for weakly supervised person search
Huadong Lin, Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001, Chen Wang 0026 |
Neurocomputing | 2 |
| 2025 | Visual Perturbation for Text-Based Person SearchabstractText-based person search aims at locating a person described by natural language in uncropped scene images. Recent works for TBPS mainly focus on aligning multi-granularity vision and language representations, neglecting a key discrepancy between training and inference where the former learns to unify vision and language features where the visual side covers all clues described by language, yet the latter matches image-text pairs where the images may capture only part of the described clues due to perturbations such as occlusions, background clutters and misaligned boundaries. To alleviate this issue, we present ViPer: a Visual Perturbation network that learns to match language descriptions with perturbed visual clues. On top of a CLIP-driven baseline, we design three visual perturbation modules: (1) Spatial ViPer that varies person proposals and produces visual features with misaligned boundaries, (2) Attentive ViPer that estimates visual attention on the fly and manipulates attentive visual tokens within a proposal to produce global features under visual perturbations, and (3) Fine-grained ViPer that learns to recover masked visual clues from detailed language descriptions to encourage matching language features with perturbed visual features at the fine granularity. This overall framework thus simulates real-world scenarios at the training stage to minimize the discrepancy and improve the generalization ability of the model. Experimental results demonstrate that the proposed method clearly surpasses previous TBPS methods on the PRW-TBPS and CUHK-SYSU-TBPS datasets. Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001 |
AAAI | 1 |
| 2025 | Revisiting Continual Ultra-fine-grained Visual Recognition with Pre-trained ModelsabstractContinual ultra-fine-grained visual recognition (C-UFG) aims to continuously learn to categorize the increasing number of cultivates (VC-UFG) and consistently recognize crops across reproductive stages (HC-UFG), which is a fundamental goal of intelligent agriculture. Despite the progress made in general continual learning, C-UFG remains an underexplored issue. This work establishes the first comprehensive C-UFG benchmark using massive soy leaf data. By analyzing recent pre-trained model (PTM) based continual learning methods on the proposed benchmark, we propose two simple yet effective PTM-based methods to boost the performance of VC-UFG and HC-UFG, respectively. On top of those, we integrate the two methods into one unified framework and propose the first unified model, Unic, that is capable of tackling the C-UFG problem where VC-UFG and HC-UFG co-exist in a single continual learning sequence. To understand the effectiveness of the proposed methods, we first evaluate the models on VC-UFG and HC-UFG challenges and then test the proposed Unic on a unified C-UFG challenge. Experimental results demonstrate the proposed methods achieve superior performance for C-UFG. The code is available at https://github.com/PatrickZad/unicufg. Pengcheng Zhang 0003, Xiaohan Yu 0001, Meiying Gu, Yongsheng Gao 0001, Xiao Bai 0001 |
IJCAI | 1 |
| 2025 | View-aware Decomposition and Unification for Fast Ground-to-Aerial Person SearchabstractGround-to-aerial person search leverages cooperative efforts between unmanned aerial vehicles (UAV) and ground surveillance cameras to locate person individuals. Despite the progress made by recent works, the impact of the discrepancy between the two views is underestimated. This limits the overall person search performance when training the model in a view-agnostic way. To address this, we propose a view-aware decomposition and unification (VADU) framework for ground-to-aerial person search. Specifically, we decompose the person search model to learn view-oriented modules for image feature encoding and person proposal generation. The data sampling and retrieval feature learning are also composed to cope with the decomposed model. This decomposition improves both person detection and discriminative feature learning within each view. On top of the decomposition, we propose view-aware unification to produce unified cross-view person features. Cross-view prototypical contrastive learning is introduced to enhance the unification between different views, enhancing model robustness to retrieve a target person in cameras of a different view. As the decomposed parts of the model are deployed on different devices for inference, this overall framework adds no extra computation cost in real-world applications. Extensive experiments demonstrate that the proposed method achieves superior person search performance and guarantees the efficiency of inference. The source code is available at https://github.com/QFWang-11/vadu. Qifei Wang, Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001, Yongsheng Gao 0001 |
IROS | 2 |
| 2025 | Fully Decoupled End-to-End Person Search: An Approach without Conflicting Objectives
Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001, Xin Ning 0001, Edwin R. Hancock |
Int. J. Comput. Vis. | 1 |
| 2024 | Occluded Person Retrieval with Hierarchical Feature OptimizationabstractOccluded person retrieval aims to match images from occluded pedestrians. It pushes forward progress of person retrieval towards applications in real-world scenarios, thus attracting increasing attention in recent years. A key challenge is to learn discriminative representation within limited informative regions due to obstacle or pedestrian occlusion. To that end, we propose a hierarchical feature optimization model (HFO) that jointly optimizes image-level, object-level and part-level features for improved occluded person retrieval. A hierarchical discriminative feature grouping (HDFG) module is developed to generate hierarchical object/part masks for comprehensive feature extraction. Via learning a set of part prototypes, HDFG localizes hierarchical informative object/parts by grouping intermediate feature vectors based on their similarity to these prototypes. The proposed HFO is trained in an end-to-end manner using only identity labels, making it a practical solution for occluded person retrieval. We verify the effectiveness of the proposed method on three challenging occluded datasets and two holistic datasets, i.e., Occluded-DukeMTMC, Occluded-REID, P-DukeMTMC-reID, Market1501, and DukeMTMC-reID. Extensive experiments and ablation studies demonstrate superior or comparable performance of the proposed method over the state-of-the-art methods. The code is available at https://github.com/Patrickzad/HFO. Yang Zhao 0019, Pengcheng Zhang 0003, Xiaohan Yu 0001, Zhibin Liao, Johan Verjans, Xiao Bai 0001 |
FG | 2 |
| 2024 | Prompting Continual Person SearchabstractThe development of person search techniques has been greatly promoted in recent years for its superior practicality and challenging goals. Despite their significant progress, existing person search models still lack the ability to continually learn from increasing real-world data and adaptively process input from different domains. To this end, this work introduces the continual person search task that sequentially learns on multiple domains and then performs person search on all seen domains. This requires balancing the stability and plasticity of the model to continually learn new knowledge without catastrophic forgetting. For this, we propose a Prompt-based Continual Person Search (PoPS) model in this paper. First, we design a compositional person search transformer to construct an effective pre-trained transformer without exhaustive pre-training from scratch on large-scale person search data. This serves as the fundamental for prompt-based continual learning. On top of that, we design a domain incremental prompt pool with a diverse attribute matching module. For each domain, we independently learn a set of prompts to encode the domain-oriented knowledge. Meanwhile, we jointly learn a group of diverse attribute projections and prototype embeddings to capture discriminative domain attributes. By matching an input image with the learned attributes across domains, the learned prompts can be properly selected for model inference. Extensive experiments are conducted to validate the proposed method for continual person search. The source code is available at https://github.com/PatrickZad/PoPS. Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001, Xin Ning 0001 |
ACM Multimedia | 1 |
| 2024 | Consistent prototype contrastive learning for weakly supervised person search
Huadong Lin, Xiaohan Yu 0001, Pengcheng Zhang 0003, Xiao Bai 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2024 | Learning From Human Attention for Attribute-Assisted Visual RecognitionabstractWith prior knowledge of seen objects, humans have a remarkable ability to recognize novel objects using shared and distinct local attributes. This is significant for the challenging tasks of zero-shot learning (ZSL) and fine-grained visual classification (FGVC), where the discriminative attributes of objects have played an important role. Inspired by human visual attention, neural networks have widely exploited the attention mechanism to learn the locally discriminative attributes for challenging tasks. Though greatly promoted the development of these fields, existing works mainly focus on learning the region embeddings of different attribute features and neglect the importance of discriminative attribute localization. It is also unclear whether the learned attention truly matches the real human attention. To tackle this problem, this paper proposes to employ real human gaze data for visual recognition networks to learn from human attention. Specifically, we design a unified Attribute Attention Network (A$^{2}$Net) that learns from human attention for both ZSL and FGVC tasks. The overall model consists of an attribute attention branch and a baseline classification network. On top of the image feature maps provided by the baseline classification network, the attribute attention branch employs attribute prototypes to produce attribute attention maps and attribute features. The attribute attention maps are converted to gaze-like attentions to be aligned with real human gaze attention. To guarantee the effectiveness of attribute feature learning, we further align the extracted attribute features with attribute-defined class embeddings. To facilitate learning from human gaze attention for the visual recognition problems, we design a bird classification game to collect real human gaze data using the CUB dataset via an eye-tracker device. Experiments on ZSL and FGVC tasks without/with real human gaze data validate the benefits and accuracy of our proposed model. This work supports the promising benefits of collecting human gaze datasets and automatic gaze estimation algorithms learning from human attention for high-level computer vision tasks. Xiao Bai 0001, Pengcheng Zhang 0003, Xiaohan Yu 0001, Edwin R. Hancock, Jun Zhou 0001, Lin Gu 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Joint discriminative representation learning for end-to-end person search
Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001, Chen Wang 0026, Xin Ning 0001 |
Pattern Recognit. | 1 |
| 2024 | Towards effective person search with deep learning: A survey from systematic perspective
Pengcheng Zhang 0003, Xiaohan Yu 0001, Chen Wang 0026, Xin Ning 0001, Xiao Bai 0001 |
Pattern Recognit. | 1 |
| 2023 | Information bottleneck and selective noise supervision for zero-shot learning
Lei Zhou 0008, Yang Liu 0357, Pengcheng Zhang 0003, Xiao Bai 0001, Lin Gu 0003, Jun Zhou 0001, Yazhou Yao, Tatsuya Harada, Edwin R. Hancock |
Mach. Learn. | 3 |
| 2022 | Where to Focus: Investigating Hierarchical Attention Relationship for Fine-Grained Visual Classification
Yang Liu 0357, Lei Zhou 0008, Pengcheng Zhang 0003, Xiao Bai 0001, Lin Gu 0003, Xiaohan Yu 0001, Jun Zhou 0001, Edwin R. Hancock |
ECCV (24) | 3 |