VLDB 2026 Research / reviewers in the wild / expert
Liying Gao
dblp:300/9120
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0003-4204-2092ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Post-enhancing viewpoint robustness for pre-training generalizable re-identification
Yunlong Wang 0008, Bingliang Jiao, Liying Gao, Wenxuan Wang 0003, Peng Wang 0015 |
Knowl. Based Syst. | 3 |
| 2026 | Words Strip Wardrobe: Learning Invariant Textual Prompts for Cloth-Changing Person Re-IdentificationabstractCloth-changing person re-identification (CC-ReID) aims to match individuals wearing varying clothes across camera views. Existing CC-ReID methods typically focus on extracting clothing-invariant features such as body shape, pose, gait,etc. However, these features are diverse and often entangled with clothing-related visual clues, posing significant challenges for comprehensively and effectively separating and leveraging them for re-identification. To address these challenges, we propose a text-guided clothing generalizable (Tex-CG) model, which employs multi-modal large language models (MLLMs) to comprehensively and explicitly decouple clothing-invariant features from pedestrian images in the textual domain. By shifting feature disentanglement to the textual domain, the interference caused by visual entanglement between clothing and clothing-invariant clues can be significantly reduced. Additionally, to ensure compatibility between offline-decoupled features from MLLMs and our online-trained Tex-CG model, we utilize CLIP-based image-text matching to train implicit clothing-invariant prompts that embed discriminative pedestrian information. A dynamic fusion module is subsequently introduced to leverage these implicit prompts for selectively integrating valuable and compatible components from the MLLM’s explicitly decoupled features, constructing robust guidance to direct our model to effectively capture clothing-invariant visual clues for re-identification. Extensive experiments demonstrate the effectiveness of our method, and the Tex-CG model achieves state-of-the-art performance on five mainstream CC-ReID benchmarks. Our code is available at https://github.com/JiaoBL1234/Tex-CG. Bingliang Jiao, Liying Gao, Feiyue Zhao, Dapeng Oliver Wu, Wenxuan Wang 0003, Peng Wang 0015 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Generalizable Person Re-Identification From a 3D Perspective: Addressing Unpredictable Viewpoint ChangesabstractMost existing Domain Generalizable Person Re-identification (DG-ReID) methods focus on addressing style disparities between domains but often overlook the impact of unpredictable camera view changes, which we have identified as a significant factor responsible for poor generalization performance. To address this issue, we propose a novel approach from a 3D perspective, utilizing a customized 2D-to-3D reconstruction model to convert images captured from arbitrary camera views into canonical view images. However, merely applying a 3D reconstruction model in isolation may not result in improved DG-ReID performance, as reconstruction quality can be influenced by multiple factors, such as insufficient image resolution, extreme viewpoint, and environmental variations. These factors may lead to error accumulation and the loss of critical discriminative clues in the reconstructed results. To address this difficulty, we propose fusing the canonical view image with the original image using a transformer-based module. The transformer’s cross-attention mechanism is ideal for aligning and fusing the key semantic clues of the original image with the canonical view image, compensating for reconstruction errors. We demonstrate the effectiveness of our method through extensive experiments in various evaluation settings, achieving superior DG-ReID performance compared to existing approaches. Our approach addresses the impact of unpredictable camera view changes and provides a new perspective for designing DG-ReID methods. Bingliang Jiao, Lingqiao Liu, Liying Gao, Dapeng Oliver Wu, Guosheng Lin, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Contrastive Pedestrian Attentive and Correlation Learning Network for Occluded Person Re-IdentificationabstractOccluded person Re-identification (ReID) aims to match occluded and holistic pedestrian images across different camera views. This task presents two primary challenges. First, it is crucial to accurately capture pedestrian foregrounds from seriously occluded person images. Second, a noticeable information asymmetry exists between the partial body in occluded images and the complete body in corresponding holistic images, which could cause the ReID model to underestimate their similarities. To address these challenges, we introduce a contrastive pedestrian attentive and correlation learning (CpaCol) model. Within CpaCol, we first design a Contrastive Pedestrian Attention (ContrastAttn) module to capture pedestrian foregrounds from occluded images. In this process, we notice that most existing attention-based methods only supervise the final predictions with identity loss yet neglect its causality with the generated attention maps, which could mislead the model to capture some salient yet pedestrian-irrelevant noises as discriminative clues. To rectify this, we integrate contrastive learning into our ContrastAttn module to guide it to learn the semantic divergence between pedestrian foregrounds and noises, thereby capturing pedestrian foregrounds more accurately. Besides, we propose a correlation learning module, where we tailor an effective dense feature correlation learning tool, 4D convolution, to enable it to adapt to pedestrian images and capture corresponding clues between comparing images. By focusing more on corresponding clues, our model could avoid overemphasizing the inherent information asymmetry between occluded and holistic images, thereby improving re-identification. Empowered by these modules, our CpaCol achieves state-of-the-art performance on three relevant ReID settings,i.e., occluded, partial, and holistic ReID. Our code is available in https://github.com/nwpugaoliying/CpaCol. Liying Gao, Bingliang Jiao, Yuzhou Long, Kai Niu 0002, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Vehicle Re-Identification in Aerial Images and Videos: Dataset and ApproachabstractIn this work, we propose a large-scale dataset, VRAI, and an effective Orientation Adaptive and Salience Attentive (OASA) Network for vehicle re-identification (ReID) in aerial imagery. The VRAI dataset includes two subsets: VRAI-Image, which contains over 137,000 images of 13,000 vehicle instances, and VRAI-Video, which comprises more than 14,000 video trajectories of 7,000 identities. To our best knowledge, this is the largest dataset for UAV-based vehicle ReID, and the first dataset proposed for video-based ReID under UAV views. Based on the VRAI dataset, we design an OASA network to address two crucial challenges of vehicle ReID in aerial imagery. Firstly, the significant vehicle orientation variations in aerial images could cause great vehicle pattern deformations, making it difficult to identify vehicles across UAV views. To overcome this challenge, in our OASA, an orientation adaptive dynamic convolution module is designed, which constructs customized kernels for each vehicle instance to extract orientation-invariant features. Besides, the unique vertical view and long focal length of the UAV platform often render many salient vehicle attributes, such as logos and license plates, invisible, which brings a great challenge to ReID models to extract distinguishable vehicle features. To address this issue, in the OASA, we design a transformer-based salience attentive module (Trans-Attn) that guides the model to focus on subtle yet discriminative clues of vehicle instances in aerial imagery. Through extensive experiments, both of our designed modules are verified effective. Besides, our OASA model outperforms state-of-the-art algorithms both on our VRAI dataset and other surveillance-based datasets. Our VRAI dataset is available in https://github.com/JiaoBL1234/VRAI-Dataset. Bingliang Jiao, Lu Yang 0016, Liying Gao, Peng Wang 0015, Shizhou Zhang, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Scale Adaptive Network for Partial Person Re-identification: Counteracting Scale Variance
Bingliang Jiao, Liying Gao, Peng Wang 0015 |
BMVC | 3 |
| 2023 | Toward Re-Identifying Any AnimalabstractThe current state of re-identification (ReID) models poses limitations to their applicability in the open world, as they are primarily designed and trained for specific categories like person or vehicle. In light of the importance of ReID technology for tracking wildlife populations and migration patterns, we propose a new task called ``Re-identify Any Animal in the Wild'' (ReID-AW). This task aims to develop a ReID model capable of handling any unseen wildlife category it encounters. To address this challenge, we have created a comprehensive dataset called Wildlife-71, which includes ReID data from 71 different wildlife categories. This dataset is the first of its kind to encompass multiple object categories in the realm of ReID. Furthermore, we have developed a universal re-identification model named UniReID specifically for the ReID-AW task. To enhance the model's adaptability to the target category, we employ a dynamic prompting mechanism using category-specific visual prompts. These prompts are generated based on knowledge gained from a set of pre-selected images within the target category. Additionally, we leverage explicit semantic knowledge derived from the large-scale pre-trained language model, GPT-4. This allows UniReID to focus on regions that are particularly useful for distinguishing individuals within the target category. Extensive experiments have demonstrated the remarkable generalization capability of our UniReID model. It showcases promising performance in handling arbitrary wildlife categories, offering significant advancements in the field of ReID for wildlife conservation and research purposes. Bingliang Jiao, Lingqiao Liu, Liying Gao, Ruiqi Wu 0001, Guosheng Lin, Peng Wang 0015, Yanning Zhang 0001 |
NeurIPS | 3 |
| 2023 | Addressing Information Inequality for Text-Based Person Search via Pedestrian-Centric Visual Denoising and Bias-Aware AlignmentsabstractText-based person search is an important task in video surveillance, which aims to retrieve the corresponding pedestrian images with a given description. In this fine-grained retrieval task, accurate cross-modal information matching is an essential yet challenging problem. However, existing methods usually ignore the information inequality between modalities, which could introduce great difficulties to cross-modal matching. Specifically, in this task, the images inevitably contain some pedestrian-irrelevant noise like background and occlusion, and the descriptions could be biased to partial pedestrian content in images. With that in mind, in this paper, we propose a Text-Guided Denoising and Alignment (TGDA) model to alleviate the information inequality and realize effective cross-modal matching. In TGDA, we first design a prototype-based denoising module, which integrates pedestrian knowledge from textual features into a prototype vector and uses it as guidance to filter out pedestrian-irrelevant noise from visual features. Thereafter, a bias-aware alignment module is introduced, which guides our model to focus on the description-biased pedestrian content in cross-modal features consistently. Through extensive experiments, the effectiveness of both modules has been validated. Besides, our TGDA achieves state-of-the-art performance on various related benchmarks. Liying Gao, Kai Niu 0002, Bingliang Jiao, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Dynamically Transformed Instance Normalization Network for Generalizable Person Re-Identification
Bingliang Jiao, Lingqiao Liu, Liying Gao, Guosheng Lin, Lu Yang 0016, Shizhou Zhang, Peng Wang 0015, Yanning Zhang 0001 |
ECCV (14) | 3 |
| 2022 | Temporal-Consistent Visual Clue Attentive Network for Video-Based Person Re-IdentificationabstractVideo-based person re-identification (ReID) aims to match video trajectories of pedestrians across multi-view cameras and has important applications in criminal investigation and intelligent surveillance. Compared with single image re-identification, the abundant temporal information contained in video sequences makes it describe pedestrian instances more precisely and effectively. Recently, most existing video-based person ReID algorithms have made use of temporal information by fusing diverse visual contents captured in independent frames. However, these algorithms only measure the salience of visual clues in each single frame, inevitably introducing momentary interference caused by factors like occlusion. Therefore, in this work, we introduce a Temporal-consistent Visual Clue Attentive Network (TVCAN), which is designed to capture temporal-consistently salient pedestrian contents among frames. Our TVCAN consists of two major modules, the TCSA module, and the TCCA module, which are responsible for capturing and emphasizing consistently salient visual contents from the spatial dimension and channel dimension, respectively. Through extensive experiments, the effectiveness of our designed modules has been verified. Additionally, our TVCAN outperforms all compared state-of-the-art methods on three mainstream benchmarks. Bingliang Jiao, Liying Gao, Peng Wang 0015 |
ICMR | 2 |
| 2021 | Text-Guided Visual Feature Refinement for Text-Based Person SearchabstractText-based person search is a task to retrieve the corresponding person in a large-scale image database given a textual description, which has important value in various fields like video surveillance. In the inferring phase, language descriptions, serving as queries, guide to search the corresponding person images. Most existing methods apply cross-modal signals to guide feature refinement. However, they employ visual features from the gallery to refine textual features, which may cause high similarity between unmatched pairs. Besides, the similarity-based cross-modal attention could disturb the choice of interested areas for descriptions. In this paper, we analyze the deficiency of previous methods and carefully design a Text-guided Visual Feature Refinement network (TVFR), which utilizes text as reference to refine visual representations. Firstly, we divide each visual feature into several horizontal stripes for fine-grained refinement. After that, we employ a text-based filter generation module to generate description-customized filters, which are used to indicate the corresponding stripes mentioned in the textual input. Thereafter, we employ a text-guided visual feature refinement module to fuse part-level visual features adaptively for each description. In experiments, we validate our TVFR through extensive experiments on CUHK-PEDES, which is the only available dataset for text-based person search. To the best of our knowledge, the TVFR outperforms other state-of-the-art methods. Liying Gao, Kai Niu 0002, Zehong Ma, Bingliang Jiao, Tonghao Tan, Peng Wang 0015 |
ICMR | 1 |