EDBT 2026 Demo / reviewers in the wild / expert
Min Cao 0005
dblp:92/5343-5
· DBLP profile ↗
32ranked-venue papers
11as first author
28since 2021 · last 2026
0000-0002-6628-0370ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 8 first-author · 18 since 2021Artificial intelligence and machine learning · 20 · 6 first-author · 16 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Facial-R1: Aligning Reasoning and Recognition for Facial Emotion AnalysisabstractFacial Emotion Analysis (FEA) extends traditional facial emotion recognition by incorporating explainable, fine-grained reasoning. The task integrates three subtasks—emotion recognition, facial Action Unit (AU) recognition, and AU-based emotion reasoning—to jointly model affective states. While recent approaches leverage Vision-Language Models (VLMs) and achieve promising results, they face two critical limitations: (1) hallucinated reasoning, where VLMs generate plausible but inaccurate explanations due to insufficient emotion-specific knowledge; and (2) misalignment between emotion reasoning and recognition, caused by fragmented connections between observed facial features and final labels. We propose Facial-R1, a three-stage alignment framework that effectively addresses both challenges with minimal supervision. First, we employ instruction fine-tuning to establish basic emotional reasoning capability for reducing hallucinations. Second, we introduce reinforcement training guided by emotion and AU labels as reward signals, which explicitly aligns the generated reasoning process with the predicted emotion. Third, we design a data synthesis pipeline that iteratively leverages the prior stages to expand the training dataset, enabling scalable self-improvement of the model. Built upon this framework, we introduce FEA-20K, a benchmark dataset comprising 17,737 training and 1,688 test samples with fine-grained emotion analysis annotations. Extensive experiments across eight standard benchmarks demonstrate that Facial-R1 achieves state-of-the-art performance in FEA, with strong generalization and robust interpretability. Jiulong Wu, Yucheng Shen, Lingyong Yan, Haixin Sun 0004, Deguo Xia, Jizhou Huang, Min Cao 0005 |
AAAI | 7 |
| 2026 | Text-based Aerial-Ground Person RetrievalabstractThis work introduces Text-based Aerial-Ground Person Retrieval (TAG-PR), which aims to retrieve person images from heterogeneous aerial and ground views with textual descriptions. Unlike traditional Text-based Person Retrieval (T-PR), which focuses solely on ground-view images, TAG-PR introduces greater practical significance and presents unique challenges due to the large viewpoint discrepancy across images. To support this task, we contribute: (1) TAG-PEDES dataset, constructed from public benchmarks with automatically generated textual descriptions, enhanced by a diversified text generation paradigm to ensure robustness under view heterogeneity; and (2) TAG-CLIP, a novel retrieval framework that addresses view heterogeneity through a hierarchically-routed mixture of experts module to learn view-specific and view-agnostic features and a viewpoint decoupling strategy to decouple view-specific features for better cross-modal alignment. We evaluate the effectiveness of TAG-CLIP on both the proposed TAG-PEDES and existing T-PR benchmarks. Yu Wu 0023, Min Cao 0005, Mang Ye |
AAAI | 5 |
| 2026 | Multilingual Text-to-Image Person Retrieval via Bidirectional Relation Reasoning and AligningabstractText-to-image person retrieval (TIPR) aims to identify the target person using textual descriptions, facing challenge in modality heterogeneity. Prior works have attempted to address it by developing cross-modal global or local alignment strategies. However, global methods typically overlook fine-grained cross-modal differences, whereas local methods require prior information to explore explicit part alignments. Additionally, current methods are English-centric, restricting their application in multilingual contexts. To alleviate these issues, we pioneer a multilingual TIPR task by developing a multilingual TIPR benchmark, for which we leverage large language models for initial translations and refine them by integrating domain-specific knowledge. Correspondingly, we propose Bi-IRRA: a Bidirectional Implicit Relation Reasoning and Aligning framework to learn alignment across languages and modalities. Within Bi-IRRA, a bidirectional implicit relation reasoning module enables bidirectional prediction of masked image and text, implicitly enhancing the modeling of local relations across languages and modalities, a multi-dimensional global alignment module is integrated to bridge the modality heterogeneity. The proposed method achieves new state-of-the-art results on all multilingual TIPR datasets. Min Cao 0005, Ding Jiang, Bo Du 0001, Mang Ye, Min Zhang 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Toward Real-World Holistic Privacy-Preserving Person Re-IdentificationabstractReal-world person re-identification (Re-ID) systems are susceptible to malicious attacks, leading to the leakage of pedestrian images and the Re-ID model, posing severe threats to the privacy of both system owners and pedestrians. Existing privacy-preserving person re-identification (PPPR) methods fail to simultaneously resist data leakage, model leakage, and data & model leakage while compromising the normal functionality of Re-ID systems. In this paper, we begin with an in-depth analysis of prior methodologies and identify the gap between existing works and the ideal PPPR paradigm. Inspired by the concept of "Let the invisible perturbation become the system trigger", we propose SHIELD, a pioneering and comprehensive two-stage privacy-preserving framework. To resist data leakage, we propose a self-supervised method for Protected Dataset Generation in the first stage, which obviates the dependence on identity labels and ensures image quality. To resist model leakage without compromising the normal retrieval accuracy, we propose Original Feature Deconstruction and Protected Feature Alignment to train the system model with paired protected and original images. Extensive experiments substantiate that SHIELD significantly outperforms existing PPPR methods, offering robust and holistic protection for Re-ID systems while maintaining decent retrieval accuracy for authorized users. The code will be released soon. Qianxiang Meng, He Li 0054, Min Cao 0005, Mang Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Beyond action units: Towards multi-cue facial emotion analysis
Yucheng Shen, Jiulong Wu, Lingyong Yan, Dawei Yin 0001, Min Cao 0005, Mang Ye |
Pattern Recognit. | 6 |
| 2026 | An Empirical Study of Validating Synthetic Data for Text-Based Person Retrieval
Min Cao 0005, Ziyin Zeng, Dong Yi, Jinqiao Wang, Mang Ye |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2026 | SA-Person: Text-Based Person Retrieval With Scene-Aware Re-Ranking
Yingjia Xu, Jinlin Wu, Daming Gao, Zhen Chen 0018, Yang Yang 0062, Min Cao 0005, Mang Ye, Zhen Lei 0001 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Chat-based Person Retrieval via Dialogue-Refined Cross-Modal AlignmentabstractTraditional text-based person retrieval (TPR) relies on a single-shot text as query to retrieve the target person, assuming that the query completely captures the user’s search intent. However, in real-world scenarios, it can be challenging to ensure the information completeness of such single-shot text. To address this limitation, we propose chat-based person retrieval (ChatPR), a new paradigm that takes an interactive dialogue as query to perform the person retrieval, engaging the user in conversational context to progressively refine the query for accurate person retrieval. The primary challenge in ChatPR is the lack of available dialogue-image paired data. To overcome this challenge, we establish ChatPedes, the first dataset designed for ChatPR, which is constructed by leveraging large language models to automate the question generation and simulate user responses. Additionally, to bridge the modality gap between dialogues and images, we propose a dialogue-refined cross-modal alignment (DiaNA) framework, which leverages two adaptive attribute refiner modules to bottleneck the conversational and visual information for fine-grained cross-modal alignment. Moreover, we propose a dialogue-specific data augmentation strategy, random round retaining, to further enhance the model’s generalization ability across varying dialogue lengths. Extensive experiments demonstrate that DiaNA significantly outperforms existing TPR approaches, highlighting the effectiveness of conversational interactions for person retrieval. Yucheng Ji, Min Cao 0005, Jinqiao Wang, Mang Ye |
CVPR | 3 |
| 2025 | Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference OptimizationabstractLarge Visual Language Models (LVLMs) have demonstrated impressive capabilities across multiple tasks.However, their trustworthiness is often challenged by hallucinations, which can be attributed to the modality misalignment and the inherent hallucinations of their underlying Large Language Models (LLMs) backbone.Existing preference alignment methods focus on aligning model responses with human preferences while neglecting image-text modality alignment, resulting in over-reliance on LLMs and hallucinations.In this paper, we propose Entity-centric Multimodal Preference Optimization (EMPO), which achieves enhanced modality alignment than existing human preference alignment methods.Besides, to overcome the scarcity of high-quality multimodal preference data, we utilize open-source instruction datasets to automatically construct highquality preference data across three aspects: image, instruction, and response.Experiments on two human preference datasets and five multimodal hallucination benchmarks demonstrate the effectiveness of EMPO, e.g., reducing hallucination rates by 85. Jiulong Wu, Zhengliang Shi, Shuaiqiang Wang, Jizhou Huang, Dawei Yin 0001, Lingyong Yan, Min Cao 0005, Min Zhang 0006 |
EMNLP | 7 |
| 2025 | RareCLIP: Rarity-Aware Online Zero-Shot Industrial Anomaly Detection
Jianfang He, Min Cao 0005, Silong Peng, Qiong Xie |
ICCV | 2 |
| 2025 | Mitigating Hallucinations in Large Vision-Language Models via Dual Contrastive DecodingabstractLarge Vision-Language Models (LVLMs) have shown impressive performance across diverse multimodal tasks. However, their trustworthiness is challenged by two types of hallucinations: over-reliance on the language priors of the underlying Large Language Model (LLM) and misalignment between visual and textual modalities. Most existing hallucination mitigation methods often require high training costs, limiting their effectiveness and robustness. To address this, we propose Dual Contrastive Decoding (DCD), which mitigates hallucinations stemming from language priors by subtracting image-irrelevant distributions and further aligns image-text modalities by enabling LVLMs to focus on correct image entities. This strategy requires no additional training and can be applied during the decoding phase of any LVLM, effectively improving the inference capability of LVLMs at minimal cost. Extensive experiments on discriminative (POPE, MME) benchmarks demonstrate that DCD significantly mitigates hallucinations at both the object and attribute levels, while also enhancing overall model performance on general vision-language perception and recognition tasks. The code and dataset will be made available. Jiulong Wu, Yucheng Shen, Haixin Sun 0004, Min Cao 0005 |
MMAsia | 4 |
| 2025 | Rethinking visual prompt learning as masked visual token modeling
Ning Liao, Bowen Shi 0003, Xiaopeng Zhang 0008, Min Cao 0005, Junchi Yan, Qi Tian 0001 |
Artif. Intell. | 4 |
| 2025 | Prompt-based Weakly-supervised Vision-language Pre-trainingabstractWeakly-supervised Vision-Language Pre-training (W-VLP) explores methods leveraging weak cross-modal supervision, typically relying on object tags generated by a pre-trained object detector (OD) from images. However, training such an OD necessitates dense cross-modal information, including images paired with numerous object-level annotations. To alleviate that requirement, this paper addresses W-VLP in two stages: (1) creating data with weaker cross-modal supervision and (2) pre-training a vision-language (VL) model with the created data. The data creation process involves collecting knowledge from large language models (LLMs) to describe images. Given a category label of an image, its descriptions generated by an LLM are used as the language counterpart. This knowledge supplements what can be obtained using an OD, such as spatial relationships among objects most likely appearing in a scene. To mitigate the noise in the LLM-generated descriptions that destabilizes the training process and may lead to overfitting, we incorporate knowledge distillation and external retrieval-augmented knowledge during pre-training. Furthermore, we present an effective VL model pre-trained with the created data. Empirically, despite its weaker cross-modal supervision, our pre-trained VL model notably outperforms other W-VLP works in image and text retrieval tasks, e.g., VLMixer by 17.7% on MSCOCO and RELIT by 11.25% on Flickr30K relatively in Recall@1 in text-to-image retrieval task. It also shows superior performance on other VL downstream tasks, making a big stride towards matching the performances of strongly supervised VLP models. The results reveal the effectiveness of the proposed W-VLP methodology. Zixin Guo, Julius Wang, Selen Pehlivan, Abduljalil Radman, Min Cao 0005, Jorma Laaksonen |
Pattern Recognit. Lett. | 5 |
| 2025 | Contextual Graph Reconstruction and Emotional Variation Learning for Conversational Emotion RecognitionabstractConversational Emotion Recognition (CER) significantly benefits from the integration of multiple modalities. However, real-world scenarios are often plagued by hardware malfunctions and network failures that lead to missing modalities and incomplete emotional representations. Existing methods primarily focus on modeling inter-modal relationships within isolated utterances to generate missing data, so they do not adequately capture conversational context and dynamic emotional evolution. In conversations, emotional expressions are inherently context-dependent and dynamic. Insufficient modeling of these properties can substantially degrade performance. To address these challenges, we propose Contextual Graph Reconstruction and Emotional Variation Learning (CGR-EVL). Our approach flexibly constructs an utterance graph with diverse coverage based on speaker activity and modality availability, thereby capturing more contextual information. To enhance the handling of missing modalities, we further integrate temporal relations through a relational graph convolutional network to reconstruct missing features and introduce a speaker emotion-aware constraint to ensure emotional coherence. Additionally, we propose the concept of emotional entropy to quantify variation patterns and develop a novel loss function that aligns predicted and actual variations. Experiments on the IEMOCAP and MELD datasets show that CGR-EVL outperforms state-of-the-art methods, particularly under conditions of incomplete modalities. Yujing Rao, Min Cao 0005, Mang Ye |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | M-Tuning: Prompt Tuning With Mitigated Label Bias in Open-Set ScenariosabstractIn realistic open-set scenarios where labels of a part of testing data are totally unknown, when vision-language (VL) prompt learning methods encounter inputs related to unknown classes (i.e., not seen during training), they always predict them as one of the training classes. The exhibited label bias causes difficulty in open set recognition (OSR), in which an image should be correctly predicted as one of the known classes or the unknown one. To achieve this goal, we propose a vision-language prompt tuning method with mitigated label bias (M-Tuning). It introduces open words from the WordNet to extend the range of words forming the prompt texts from only closed-set label words to more, and thus prompts are tuned in a simulated open-set scenario. Besides, inspired by the observation that classifying directly on large datasets causes a much higher false positive rate than on small datasets, we propose a Combinatorial Tuning and Testing (CTT) strategy for improving performance. CTT decomposes M-Tuning on large datasets as multiple independent group-wise tuning on fewer classes, then makes accurate and comprehensive predictions by selecting the optimal sub-prompt. Finally, given the lack of VL-based OSR baselines in the literature, especially for prompt methods, we contribute new baselines for fair comparisons. Our method achieves the best performance on datasets with various scales, and extensive ablation studies also validate its effectiveness. Ning Liao, Xiaopeng Zhang 0008, Min Cao 0005, Junchi Yan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Semi-Supervised Text-Based Person SearchabstractText-based person search (TBPS) aims to retrieve images of a specific person from a large image gallery based on a natural language description. Existing methods rely on massive annotated image-text data to achieve satisfactory performance in fully-supervised learning. This presents a substantial practical challenge, given the difficulty in obtaining annotated texts for person images. This work undertakes a pioneering initiative to explore TBPS under the semi-supervised setting, where only a limited number of person images are annotated with textual descriptions while the majority of images lack annotations. We present a two-stage basic solution based on generation-then-retrieval for semi-supervised TBPS. The generation stage enriches annotated data by applying an image captioning model to generate pseudo-texts for unannotated images. Later, the retrieval stage performs fully-supervised retrieval learning using the augmented data. Crucially, considering the noise interference of the pseudo-texts on retrieval learning, we propose a noise-robust retrieval framework that enhances the ability of the retrieval model to handle noisy data. The framework integrates two key strategies: Hybrid Patch-Channel Masking (PC-Mask) to refine the model architecture, and Noise-Guided Progressive Training (NP-Train) to enhance the training process. PC-Mask performs masking on the input data at both the patch-level and the channel-level to prevent overfitting noisy supervision. NP-Train introduces a progressive training schedule based on the noise level of pseudo-texts to facilitate noise-robust learning. Extensive experiments on multiple TBPS benchmarks show that the proposed framework achieves promising performance under the semi-supervised setting. Daming Gao, Min Cao 0005, Hao Dou, Mang Ye, Min Zhang 0005 |
IEEE Trans. Image Process. | 3 |
| 2024 | An Empirical Study of CLIP for Text-Based Person SearchabstractText-based Person Search (TBPS) aims to retrieve the person images using natural language descriptions. Recently, Contrastive Language Image Pretraining (CLIP), a universal large cross-modal vision-language pre-training model, has remarkably performed over various cross-modal downstream tasks due to its powerful cross-modal semantic learning capacity. TPBS, as a fine-grained cross-modal retrieval task, is also facing the rise of research on the CLIP-based TBPS. In order to explore the potential of the visual-language pre-training model for downstream TBPS tasks, this paper makes the first attempt to conduct a comprehensive empirical study of CLIP for TBPS and thus contribute a straightforward, incremental, yet strong TBPS-CLIP baseline to the TBPS community. We revisit critical design considerations under CLIP, including data augmentation and loss function. The model, with the aforementioned designs and practical training tricks, can attain satisfactory performance without any sophisticated modules. Also, we conduct the probing experiments of TBPS-CLIP in model generalization and model compression, demonstrating the effectiveness of TBPS-CLIP from various aspects. This work is expected to provide empirical insights and highlight future CLIP-based TBPS research. Min Cao 0005, Ziyin Zeng, Mang Ye, Min Zhang 0005 |
AAAI | 1 |
| 2024 | MAIR: A Massive Benchmark for Evaluating Instructed RetrievalabstractWeiwei Sun, Zhengliang Shi, Wu Jiu Long, Lingyong Yan, Xinyu Ma, Yiding Liu, Min Cao, Dawei Yin, Zhaochun Ren. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Weiwei Sun 0001, Zhengliang Shi, Wu Long, Lingyong Yan, Xinyu Ma 0001, Min Cao 0005, Dawei Yin 0001, Zhaochun Ren |
EMNLP | 7 |
| 2024 | LAIP: Learning Local Alignment from Image-Phrase Modeling for Text-based Person SearchabstractText-based person search aims at retrieving images of a particular person based on a given textual description. A common solution for this task is to directly match the entire images and texts, i.e., global alignment, which fails to deal with discerning specific details that discriminate against appearance-similar people. As a result, some works shift their attention towards local alignment. One group matches fine-grained parts using forward attention weights of the transformer yet underutilizes information. Another implicitly conducts local alignment by reconstructing masked parts based on unmasked context yet with a biased masking strategy. All limit performance improvement. This paper proposes the Local Alignment from Image-Phrase modeling (LAIP) framework, with Bidirectional Attention-weighted local alignment (BidirAtt) and Mask Phrase Modeling (MPM) module. BidirAtt goes beyond the typical forward attention by considering the gradient of the transformer as backward attention, utilizing two-sided information for local alignment. MPM focuses on mask reconstruction within the noun phrase rather than the entire text, ensuring an unbiased masking strategy. Extensive experiments conducted on the CUHK-PEDES, ICFG-PEDES, and RSTPReid datasets demonstrate the superiority of the LAIP framework over existing methods. Yu Wu 0023, Mengxia Wu, Min Cao 0005, Min Zhang 0005 |
ICME | 4 |
| 2024 | Context-aided unicity matching for person re-identification
Min Cao 0005, Cong Ding 0013, Chen Chen 0036, Silong Peng |
J. Vis. Commun. Image Represent. | 1 |
| 2024 | Efficient Image-Text Retrieval via Keyword-Guided Pre-ScreeningabstractImage-text retrieval is a fundamental task to model a connection between images and natural language. Under its flourishing development in performance, most current methods suffer fromN-related time complexity, which hinders their application in practice to a certain extent. Targeting efficiency improvement, we propose a simple and effective keyword-guided pre-screening framework for image-text retrieval. Specifically, we convert the image and text data into keywords and perform keyword matching across the modalities to exclude a large number of irrelevant gallery samples prior to the retrieval network. For the keyword prediction, we transfer it into a multi-label classification problem and propose a multi-task learning scheme by appending the multi-label classifiers to the image-text retrieval network to achieve a lightweight and high-performance keyword prediction. For keyword matching, we introduce the inverted index from the search engine and thus create a win-win situation on both time and space complexities for the pre-screening. Extensive experiments on the two widely-used datasets,i.e., Flickr30K and MS-COCO, verify the effectiveness of the proposed framework. The proposed framework equipped with only two embedding layers achievesO(1) querying time complexity, while improving the retrieval efficiency and maintaining performance, when applied prior to the common image-text retrieval methods. Min Cao 0005, Ziqiang Cao, Liqiang Nie, Min Zhang 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Unpaired Multi-domain Attribute Translation of 3D Facial Shapes with a Square and Symmetric Geometric MapabstractWhile impressive progress has recently been made in image-oriented facial attribute translation, shape-oriented 3D facial attribute translation remains an unsolved issue. This is primarily limited by the lack of 3D generative models and ineffective usage of 3D facial data. We propose a learning framework for 3D facial attribute translation to relieve these limitations. Firstly, we customize a novel geometric map for 3D shape representation and embed it in an end-to-end generative adversarial network. The geometric map represents 3D shapes symmetrically on a square image grid, while preserving the neighboring relationship of 3D vertices in a local least-square sense. This enables effective learning for the latent representation of data with different attributes. Secondly, we employ a unified and unpaired learning framework for multi-domain attribute translation. It not only makes effective usage of data correlation from multiple domains, but also mitigates the constraint for hardly accessible paired data. Finally, we propose a hierarchical architecture for the discriminator to guarantee robust results against both global and local artifacts. We conduct extensive experiments to demonstrate the advantage of the proposed framework over the state-of-the-art in generating high-fidelity facial shapes. Given an input 3D facial shape, the proposed framework is able to synthesize novel shapes of different attributes, which covers some downstream applications, such as expression transfer, gender translation, and aging. Code at https://github.com/NaughtyZZ/3D_facial_shape_attribute_translation_ssgmap. Zhenfeng Fan, Chongyang Zhong, Min Cao 0005, Shihong Xia |
ICCV | 5 |
| 2023 | RaSa: Relation and Sensitivity Aware Representation Learning for Text-based Person SearchabstractText-based person search aims to retrieve the specified person images given a textual description. The key to tackling such a challenging task is to learn powerful multi-modal representations. Towards this, we propose a Relation and Sensitivity aware representation learning method (RaSa), including two novel tasks: Relation-Aware learning (RA) and Sensitivity-Aware learning (SA). For one thing, existing methods cluster representations of all positive pairs without distinction and overlook the noise problem caused by the weak positive pairs where the text and the paired image have noise correspondences, thus leading to overfitting learning. RA offsets the overfitting risk by introducing a novel positive relation detection task (i.e., learning to distinguish strong and weak positive pairs). For another thing, learning invariant representation under data augmentation (i.e., being insensitive to some transformations) is a general practice for improving representation's robustness in existing methods. Beyond that, we encourage the representation to perceive the sensitive transformation by SA (i.e., learning to detect the replaced words), thus promoting the representation's robustness. Experiments demonstrate that RaSa outperforms existing state-of-the-art methods by 6.94%, 4.45% and 15.35% in terms of Rank@1 on CUHK-PEDES, ICFG-PEDES and RSTPReid datasets, respectively. Code is available at: https://github.com/Flame-Chasers/RaSa. Min Cao 0005, Daming Gao, Ziqiang Cao, Chen Chen 0036, Zhenfeng Fan, Liqiang Nie, Min Zhang 0005 |
IJCAI | 2 |
| 2023 | Text-based Person Search without Parallel Image-Text DataabstractText-based person search (TBPS) aims to retrieve the images of the target person from a large image gallery based on a given natural language description. Existing methods are dominated by training models with parallel image-text pairs, which are very costly to collect. In this paper, we make the first attempt to explore TBPS without parallel image-text data (μ-TBPS), in which only non-parallel images and texts, or even image-only data, can be adopted. Towards this end, we propose a two-stage framework, generation-then-retrieval (GTR), to first generate the corresponding pseudo text for each image and then perform the retrieval in a supervised manner. In the generation stage, we propose a fine-grained image captioning strategy to obtain an enriched description of the person image, which firstly utilizes a set of instruction prompts to activate the off-the-shelf pretrained vision-language model to capture and generate fine-grained person attributes, and then converts the extracted attributes into a textual description via the finetuned large language model or the hand-crafted template. In the retrieval stage, considering the noise interference of the generated texts for training model, we develop a confidence score-based training scheme by enabling more reliable texts to contribute more during the training. Experimental results on multiple TBPS benchmarks (i.e., CUHK-PEDES, ICFG-PEDES and RSTPReid) show that the proposed GTR can achieve a promising performance without relying on parallel image-text data. Min Cao 0005, Chen Chen 0036, Ziqiang Cao, Liqiang Nie, Min Zhang 0005 |
ACM Multimedia | 3 |
| 2023 | Progressive Context-Aware Graph Feature Learning for Target Re-IdentificationabstractThis paper aims at robust and discriminative feature learning for target re-identification (Re-ID). In addition to paying attention to the individual appearance information as in most Re-ID methods, we further utilize the abundant contextual information as additional clues to guide the feature learning. Graph as a format of structured data is used to represent the target sample with its context. It describes the first-order appearance information of the samples and the second-order topological relationship information among samples, based on which we compute the feature representation by learning a graph feature embedding. We provide a detailed analysis of graph convolutional network mechanism applied in target Re-ID and propose a novel progressive context-aware graph feature learning method, in which the message passing is dominated by a pre-defined adjacency relationship followed by a learned relationship in a self-adaptive way. The proposed method fully exploits and utilizes contextual information at a low cost for Re-ID. Extensive experiments on five Re-ID benchmarks demonstrate the state-of-the-art performance of the proposed method. Min Cao 0005, Cong Ding 0013, Chen Chen 0036, Hao Dou, Xiyuan Hu, Junchi Yan |
IEEE Trans. Multim. | 1 |
| 2022 | Learning Semantic-Aligned Feature Representation for Text-Based Person SearchabstractText-based person search aims to retrieve images of a certain pedestrian by a textual description. The key challenge of this task is to eliminate the inter-modality gap and achieve the feature alignment across modalities. In this paper, we propose a semantic-aligned embedding method for text-based person search, in which the feature alignment across modalities is achieved by automatically learning the semantic-aligned visual features and textual features. First, we introduce two Transformer-based backbones to encode robust feature representations of the images and texts. Second, we design a semantic-aligned feature aggregation network to adaptively select and aggregate features with the same semantics into part-aware features, which is achieved by a multi-head attention module constrained by a cross-modality part alignment loss and a diversity loss. Experimental results on the CUHK-PEDES and Flickr30K datasets show that our method achieves state-of-the-art performances. Shiping Li, Min Cao 0005, Min Zhang 0005 |
ICASSP | 2 |
| 2022 | Image-text Retrieval: A Survey on Recent Research and DevelopmentabstractIn the past few years, cross-modal image-text retrieval (ITR) has experienced increased interest in the research community due to its excellent research value and broad real-world application. It is designed for the scenarios where the queries are from one modality and the retrieval galleries from another modality. This paper presents a comprehensive and up-to-date survey on the ITR approaches from four perspectives. By dissecting an ITR system into two processes: feature extraction and feature alignment, we summarize the recent advance of the ITR approaches from these two perspectives. On top of this, the efficiency-focused study on the ITR system is introduced as the third perspective. To keep pace with the times, we also provide a pioneering overview of the cross-modal pre-training ITR approaches as the fourth perspective. Finally, we outline the common benchmark datasets and evaluation metric for ITR, and conduct the accuracy comparison among the representative ITR approaches. Some critical yet less studied issues are discussed at the end of the paper. Min Cao 0005, Shiping Li, Juntao Li 0005, Liqiang Nie, Min Zhang 0005 |
IJCAI | 1 |
| 2021 | Progressive Bilateral-Context Driven Model for Post-Processing Person Re-IdentificationabstractMost existing person re-identification methods compute pairwise similarity by extracting robust visual features and learning the discriminative metric. Owing to visual ambiguities, these content-based methods that determine the pairwise relationship only based on the similarity between them, inevitably produce a suboptimal ranking list. Instead, the pairwise similarity can be estimated more accurately along the geodesic path of the underlying data manifold by exploring the rich contextual information of the sample. In this paper, we propose a lightweight post-processing person re-identification method in which the pairwise measure is determined by the relationship between the sample and the counterpart's context in an unsupervised way. We translate the point-to-point comparison into the bilateral point-to-set comparison. The sample's context is composed of its neighbor samples with two different definition ways: the first order context and the second order context, which are used to compute the pairwise similarity in sequence, resulting in a progressive post-processing model. The experiments on four large-scale person re-identification benchmark datasets indicate that (1) the proposed method can consistently achieve higher accuracies by serving as a post-processing procedure after the content-based person re-identification methods, showing its state-of-the-art results, (2) the proposed lightweight method only needs about 6 milliseconds for optimizing the ranking results of one sample, showing its high-efficiency. Code is available at: https://github.com/123ci/PBCmodel. Min Cao 0005, Chen Chen 0036, Hao Dou, Xiyuan Hu, Silong Peng, Arjan Kuijper |
IEEE Trans. Multim. | 1 |
| 2019 | Towards fast and kernelized orthogonal discriminant analysis on person re-identification
Min Cao 0005, Chen Chen 0036, Xiyuan Hu, Silong Peng |
Pattern Recognit. | 1 |
| 2018 | Ranking Loss: A Novel Metric Learning Method for Person Re-identification
Min Cao 0005, Chen Chen 0036, Xiyuan Hu, Silong Peng |
ACCV (2) | 1 |
| 2018 | Region-specific Metric Learning for Person Re-identificationabstractPerson re-identification addresses the problem of matching individual images of the same person captured by different non-overlapping camera views. Distance metric learning plays an effective role in addressing the problem. With the features extracted on several regions of person image, most of distance metric learning methods have been developed in which the learnt cross-view transformations are region-generic, i.e all region-features share a homogeneous transformation. The spatial structure of person image is ignored and the distribution difference among different region-features is neglected. Therefore in this paper, we propose a novel region-specific metric learning method in which a series of region-specific sub-models are optimized for learning cross-view region-specific transformations. Additionally, we also present a novel feature pre-processing scheme that is designed to improve the features' discriminative power by removing weakly discriminative features. Experimental results on the publicly available VIPeR, PRID450S and QMUL GRID datasets demonstrate that the proposed method performs favorably against the state-of-the-art methods. Min Cao 0005, Chen Chen 0036, Xiyuan Hu, Silong Peng |
ICPR | 1 |
| 2017 | Key Person Aided Re-identification in Partially Ordered Pedestrian Set
Chen Chen 0036, Min Cao 0005, Silong Peng |
BMVC | 2 |