EDBT 2026 Demo / reviewers in the wild / expert
Chengji Wang
dblp:222/6009
· DBLP profile ↗
15ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0002-3770-8892ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 5 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Consensus Labeling: Prompt-Guided Clustering Refinement for Weakly Supervised Text-Based Person Re-IdentificationabstractWeakly supervised text-based person re-identification aims to retrieve specific pedestrians based on textual descriptions without identity labels available during training. This task remains challenging due to the inherent cross-modal heterogeneity and lack of identity annotations. There is a common issue of modality gap in vision language models, which in turn affects the performance of downstream tasks such as cross-modal retrieval and multimodal clustering. Specifically, in our research and experiments, we found that there is a problem of inter-modal misalignment between image and text modalities. However, existing methods rely on mutual enhancement strategies between image and text clustering, leading to the accumulation of clustering noise and affecting the final retrieval performance. To address this issue, we propose a Consensus Labelling: Prompt-guided Clustering refinement (CLPC) framework for weakly supervised text-based person re-identification. Specifically, we introduce a textual inversion network to learn a pseudo token that captures visual context, which is then integrated into natural language sentences as personalized textual prompt. To further improve clustering quality, we introduce a Nearest Neighbor-Guided Pseudo Label Mining (NGPM) method, which uses the clusters derived from personalized textual prompts to refine the clustering of image features. Additionally, we design a Dynamic Margin Triplet (DMT) loss, where the margin is adaptively adjusted using a sigmoid-based function to enhance the model’s ability to distinguish hard negative samples. We have also introduce a Normalized Distribution Matching (NDM) loss to minimize the KL divergence between the image-text matching scores and the normalized soft matching scores. The extensive experimental results on three public datasets have demonstrated the superiority of our method. Our code is available at https://github.com/LeviWeiZhi/CLPC. Chengji Wang, Weizhi Nie, Hongbo Zhang 0002, Hao Sun 0014, Mang Ye |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2026 | P-CLIP: Progressive Discrepancy Learning for One-Shot Text-to-Image Person Re-IdentificationabstractOne-shot Text-to-Image Person Re-Identification (One-shot TIReID) aims to construct a TIReID model using only a single labeled image-text pair per identity, along with a large pool of unlabeled person images. While supervised learning in text-to-image person re-identification has demonstrated high effectiveness, the requirement for extensive annotated data, both in terms of identities and corresponding textual descriptions, makes it impractical for large-scale camera networks. One-shot TIReID presents a promising approach to reduce the annotation burden. The primary challenge in one-shot TIReID lies in establishing consistent visual-textual correspondences across diverse viewing conditions, particularly in the absence of cross-view paired data. To address this challenge, we propose a novel progressive discrepancy learning framework, termed P-CLIP, which aims to establish a shared embedding space that is robust to view-specific biases. To achieve this goal, we dynamically construct multi-view image-text pairs based on a single labeled pair and simultaneously project the multi-view data into a unified embedding space. Specifically, we propose a Progressive Multi-View Generation method (MVG) to generate multiple noisy views from a single labeled instance for training. To mitigate cross-view ambiguities, we introduce a Cross-View Discrepancy Learning module (CDL) that leverages the discrepancies among different views to guide the learning of cross-view visual-textual correspondences. This approach effectively integrates multimodal error correction into the person re-identification domain. Furthermore, to enhance the effectiveness of visual-textual correspondence learning, we propose a Compact Cross-Modal Matching Loss (CCM), which suppresses unmatched pairs while emphasizing matched ones. Extensive experiments were conducted on three benchmark datasets, and the experimental results demonstrate the effectiveness of our proposed method. The data and codes are available at https://github.com/Itachjw/P-CLIP/tree/main. Chengji Wang, Ming Dong 0004, Mang Ye, Hao Sun 0014, Xingpeng Jiang |
IEEE Trans. Image Process. | 1 |
| 2025 | PC-Net: Weakly Supervised Compositional Moment Retrieval via Proposal-Centric NetworkabstractWith the exponential growth of video content, aiming at localizing relevant video moments based on natural language queries, video moment retrieval (VMR) has gained significant attention. Existing weakly supervised VMR methods focus on designing various feature modeling and modal interaction modules to alleviate the reliance on precise temporal annotations. However, these methods have poor generalization capabilities on compositional queries with novel syntactic structures or vocabulary in real-world scenarios. To this end, we propose a new task: weakly supervised compositional moment retrieval (WSCMR). This task trains models using only video-query pairs without precise temporal annotations, while enabling generalization to complex compositional queries. Furthermore, a proposal-centric network (PC-Net) is proposed to tackle this challenging task. First, video and query features are extracted through frozen feature extractors, followed by modality interaction to obtain multimodal features. Second, to handle compositional queries with explicit temporal associations, a dual-granularity proposal generator decodes multimodal global and frame-level features to obtain query-relevant proposal boundaries with fine-grained temporal perception. Third, to improve the discrimination of proposal features, a proposal feature aggregator is constructed to conduct semantic alignment of frames and queries, and employ a learnable peak-aware Gaussian distributor to fit the frame weights within the proposals to derive proposal features from the video frame features. Finally, the proposal quality is assessed based on the results of reconstructing the masked query using the obtained proposal features. To further enhance the model's ability to capture semantic associations between proposals and queries, a quality margin regularizer is constructed to dynamically stratify proposals into high and low query-relevance subsets and enhance the association between queries and common elements within proposals, and suppress spurious correlations via inter-subset contrastive learning. Notably, PC-Net achieves superior performance with 54\% fewer parameters than prior works by parameter-efficient design. Experiments on Charades-CG and ActivityNet-CG demonstrate PC-Net’s ability to generalize across diverse compositional queries. Code is available at https://github.com/mingyao1120/PC-Net. Mingyao Zhou, Hao Sun 0014, Wei Xie 0008, Ming Dong 0004, Chengji Wang, Mang Ye |
NeurIPS | 5 |
| 2025 | Learning Positive-Negative Prompts for Open-Set Remote Sensing Scene Classification
Hao Sun 0014, Hanlizi Chen, Wenjing Chen 0003, Chengji Wang, Wei Xie 0008, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Modality-Dependent Sentiments Exploring for Multi-Modal Sentiment ClassificationabstractRecognizing human feelings from image and text is a core challenge of multi-modal data analysis, often applied in personalized advertising. Previous works aim at exploring the shared features, which are the matched contents between images and texts. However, the modality-dependent sentiment information (private features) in each modality is usually ignored by cross-modal interactions, the real sentiment is often reflected in one modality. In this paper, we propose a Modality-Dependent Sentiment Exploring framework (MDSE). First, to exploit the private features, we compare shared features with original image or text features, identifying previously overlooked unimodal features. Fusing the private and shared features can make the model more robust. Second, in order to obtain unified sentiment representations, we treat unimodal features and multi-modal fused features equally. We introduce a Modality-Agnostic Contrastive Loss (MACL) that performs contrastive learning between unimodal features and multi-modal fused features. The MACL can fully exploit sentiment information from multi-modal data and reduce the modality gap. Experiments on four public datasets demonstrate the effectiveness of our MDSE compared with existing methods. The full codes are available at https://github.com/royal-dargon/MDSE. Jingzhe Li, Chengji Wang, Zhiming Luo, Yuxian Wu, Xingpeng Jiang |
ICASSP | 2 |
| 2024 | Omni-Granularity Embedding Network for Text-to-Image Person RetrievalabstractText-to-image person retrieval aims to identify the desired individual based on a textual description. As an instance-level retrieval problem, it has a large intra-class variance and a small inter-class variance. Although significant progress has been made, the omni-granularity matching issue remains unaddressed. Omni-granularity matching involves aligning words with multi-granularity image regions, challenging models to learn in an omni-granularity embedding space. In this paper, we introduce a novel Omni-Granularity Embedding Network (OGEN) for person representation learning. It addresses the omni-granularity matching issue by developing a Cross-Granularity Aggregation Module (CGAM). This module dynamically consolidates diverse granularity features for learning granularity-dependent and omni-granularity person representations. Additionally, a teacher-student knowledge transfer framework is introduced to minimize the inter-modality discrepancy, allowing CGAM to focus on modality-shared semantics. Due to the effectiveness of CGAM and the knowledge transfer framework, our OGEN enhances the Rank-1 accuracy of the Baseline by 8.54%, 9.89%, and 11.09% on three public datasets, respectively. Chengji Wang, Zhiming Luo, Shaozi Li |
ICME | 1 |
| 2024 | Image-Centered Pseudo Label Generation for Weakly Supervised Text-Based Person Re-Identification
Weizhi Nie, Chengji Wang, Hao Sun 0014, Wei Xie 0008 |
PRCV (12) | 2 |
| 2024 | Uncertainty-Aware Gradient Modulation and Feature Masking for Multimodal Sentiment Analysis
Yuxian Wu, Chengji Wang, Jingzhe Li, Xingpeng Jiang |
PRCV (11) | 2 |
| 2024 | Reconstruct incomplete relation for incomplete modality brain tumor segmentation
Jiawei Su, Zhiming Luo, Chengji Wang, Sheng Lian, Xuejuan Lin, Shaozi Li |
Neural Networks | 3 |
| 2024 | Cross-Modal Feature Fusion-Based Knowledge Transfer for Text-Based Person SearchabstractText-based person search aims to retrieve corresponding images of person from a large gallery based on text descriptions. Existing methods strive to bridge the modality gap between images and texts and have made promising progress. However, these approaches disregard the knowledge imbalance between images and texts caused by the reporting bias. To resolve this issue, we present a cross-modal feature fusion-based knowledge transfer network to balance identity information between images and texts. First, we design an identity information emphasis module to enhance person-relevant information and suppress person-irrelevant information. Second, we design an intermediate modal-guided knowledge transfer module to balance the knowledge between images and texts. Experimental results on CUHK-PEDES, ICFG-PEDE, and RSTPReid datasets demonstrate that our method achieves state-of-the-art performance. Kaiyang You, Wenjing Chen 0003, Chengji Wang, Hao Sun 0014, Wei Xie 0008 |
IEEE Signal Process. Lett. | 3 |
| 2022 | Improving embedding learning by virtual attribute decoupling for text-based person search
Chengji Wang, Zhiming Luo, Yaojin Lin, Shaozi Li |
Neural Comput. Appl. | 1 |
| 2021 | Text-based Person Search via Multi-Granularity Embedding LearningabstractMost existing text-based person search methods highly depend on exploring the corresponding relations between the regions of the image and the words in the sentence. However, these methods correlated image regions and words in the same semantic granularity. It 1) results in irrelevant corresponding relations between image and text, 2) causes an ambiguity embedding problem. In this study, we propose a novel multi-granularity embedding learning model for text-based person search. It generates multi-granularity embeddings of partial person bodies in a coarse-to-fine manner by revisiting the person image at different spatial scales. Specifically, we distill the partial knowledge from image scrips to guide the model to select the semantically relevant words from the text description. It can learn discriminative and modality-invariant visual-textual embeddings. In addition, we integrate the partial embeddings at each granularity and perform multi-granularity image-text matching. Extensive experiments validate the effectiveness of our method, which can achieve new state-of-the-art performance by the learned discriminative partial embeddings. Chengji Wang, Zhiming Luo, Yaojin Lin, Shaozi Li |
IJCAI | 1 |
| 2021 | Divide-and-Merge the embedding space for cross-modality person search
Chengji Wang, Zhiming Luo, Zhun Zhong, Shaozi Li |
Neurocomputing | 1 |
| 2021 | SAFD: single shot anchor free face detector
Chengji Wang, Zhiming Luo, Zhun Zhong, Shaozi Li |
Multim. Tools Appl. | 1 |
| 2018 | Anchor Free Network for Multi-Scale Face DetectionabstractAnchor-based deep methods are the most widely used methods for face detection and have reached the state-of-the-art result. Compared with anchor-based methods that estimates the bounding-box rely on some pre-defined anchor boxes, anchor-free methods perform the localization by predicting the offsets of a pixel inside a face to its outside boundaries whose accuracies are much more precise. However, anchor-free methods suffer the drawback of low recall-rate mainly because 1) only using single scale features lead to miss detection of small faces, 2) the highly intra-class imbalance problem among different size faces. In this paper, to address these problems, we propose a unified anchor-free network for detecting multi-scale faces by leveraging the local and global contextual information of multi-layer features. We also utilize a scale aware sampling strategy to mitigate the intra-class imbalance issue which can adaptivity select the positive samples. Furthermore, a revised focal loss function is adopted to deal with the foreground/background imbalance issue. Experimental results on two benchmark datasets demonstrate the effective of our proposed method. Chengji Wang, Zhiming Luo, Sheng Lian, Shaozi Li |
ICPR | 1 |