EDBT 2026 Demo / reviewers in the wild / expert
Neng Dong
dblp:319/8601
· DBLP profile ↗
14ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0001-5523-1082ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical Prompt Learning for Image- and Text-Based Person Re-IdentificationabstractPerson re-identification (ReID) aims to retrieve target pedestrian images given either visual queries (image-to-image, I2I) or textual descriptions (text-to-image, T2I). Although both tasks share a common retrieval objective, they pose distinct challenges: I2I emphasizes discriminative identity learning, while T2I requires accurate cross-modal semantic alignment. Existing methods often treat these tasks separately, which may lead to representation entanglement and suboptimal performance. To address this, we propose a unified framework named Hierarchical Prompt Learning (HPL), which leverages task-aware prompt modeling to jointly optimize both tasks. Specifically, we first introduce a Task-Routed Transformer, which incorporates dual classification tokens into a shared visual encoder to route features for I2I and T2I branches respectively. On top of this, we develop a hierarchical prompt generation scheme that integrates identity-level learnable tokens with instance-level pseudo-text tokens. These pseudo-tokens are derived from image or text features via modality-specific inversion networks, injecting fine-grained, instance-specific semantics into the prompts. Furthermore, we propose a Cross-Modal Prompt Regularization strategy to enforce semantic alignment in the prompt token space, ensuring that pseudo-prompts preserve source-modality characteristics while enhancing cross-modal transferability. Extensive experiments on multiple ReID benchmarks validate the effectiveness of our method, achieving state-of-the-art performance on both I2I and T2I tasks. Linhan Zhou, Neng Dong, Yonghang Tai, Huafeng Li 0001 |
AAAI | 3 |
| 2026 | Modalities collaboration and granularities interaction for fine-grained sketch-based image retrieval
Junchao Ge, Jiaman Ding, Neng Dong, Shaojie Qiao, Zhengtao Yu 0001, Huafeng Li 0001 |
Pattern Recognit. | 4 |
| 2026 | Spatial-Temporal High-Frequency Learning for Video-Based Visible-Infrared Person Re-IdentificationabstractVideo-based Visible-Infrared Person Re-Identification (VVI-ReID) aims to learn consistent person feature representations across video sequences in different modalities. Existing methods that use an intermediate modality to bridge the gap between visible (RGB) and infrared (IR) sequences tend to be limited by high construction costs, loss of high-frequency details, and lack of temporal cues. Moreover, they typically focus on refining global representation using high-level features, neglecting the enhancement of local details through low-level features. To address these challenges, we propose the novel Spatial-Temporal High-Frequency Learning (STHF) framework, which constructs an appropriate intermediate modality for the VVI-ReID task and alleviates the modality gap via hierarchical feature enhancement. Specifically, we introduce the Spatial-Temporal High-Pass Filter (ST-HPF), which filters out spatial-temporal Low-Frequency Components (LFC), preserving high-frequency details to construct an intermediate modality at the sequence level. We then enhance the local details with low-level features through the Shallow Detail Compensation (SDC) module, which reduces local noise interference. Finally, the Deep Semantic Refinement (DSR) module refines the global representation by modeling spatial-temporal high-frequency semantic associations using high-level features. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art approaches on the publicly available HITSZ-VCM and BUPTCampus datasets. The code is available at https://github.com/TSC95720/STHF. Sichen Tao, Neng Dong, Fan Li 0006, Huafeng Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Richer Semantics, Better Alignment: Aligning Visual Features with Explicit and Enriched Semantics for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VIReID) retrieves pedestrian images with the same identity across different modalities. Existing methods learn visual features solely from images, failing to align them into the modality-invariant semantic space. In this paper, we propose a novel framework, termed Richer Semantics, Better Alignment (RSBA), to align visual features with explicit and enriched semantics. Specifically, we first develop an Explicit Semantics-Guided Feature Alignment (ESFA) module, which supplements textual descriptions for cross-modality images and aligns image-text pairs within each modality, alleviating the distribution discrepancy of visual features. We then devise a Consistent Similarity-Guided Indirect Alignment (CSIA) module, which constrains the similarity between intra-modality image-text pairs to be consistent with that between inter-modality text-text pairs, indirectly aligning visual features with cross-modality semantics. Furthermore, we design a Cross-View Semantics Compensation (CVSC) module, which integrates multi-view texts and improves the image-text matching of one-to-one in ESFA and CSIA to one-to-many, further strengthening the alignment of visual features within the semantic space. Extensive experimental results on three public datasets demonstrate the effectiveness and superiority of our proposed RSBA. Neng Dong, Shuanglin Yan, Liyan Zhang 0001, Jinhui Tang 0001 |
IJCAI | 1 |
| 2025 | Cross-modal Collaborative Representation Learning for Text-to-Image Person RetrievalabstractText-to-image person retrieval (TIPR) aims to find images of the same identity that match a given text description. Current TIPR methods mainly focus on mining the association between images and texts, ignoring their potential complementarity. Besides, existing matching losses treat all positive pairs from the same identity equally, leading to noisy correspondences. In this paper, we propose CoRL: a cross-modal Collaborative Representation Learning framework designed to improve TIPR by effectively leveraging the complementarity between modalities. The text typically contains identity details with less noise, which helps distinguish visually similar pedestrians. This inspires us to integrate it into the corresponding image to emphasize identity-related and modality-shared visual information. However, corresponding text for each image is not always available, especially during inference. Accordingly, we introduce a Virtual-text Embedding Synthesizer that generates high-quality virtual-text features for cross-modal collaboration, eliminating the need for actual texts. We then design a Cross-Modal Collaboration learning process, incorporating a Cross-modal Relation Consistency loss to promote interaction and fusion between image and virtual-text features for mutual enhancement. Additionally, an Identity-bounded Matching loss is proposed to handle different types of image-text pairs distinctly, leading to more accurate cross-modal correspondences. Extensive experiments on multiple benchmarks demonstrate the superiority of CoRL over existing TIPR methods. Shuanglin Yan, Jun Liu 0036, Neng Dong, Jinhui Tang 0001 |
IJCAI | 3 |
| 2025 | DINOv2 Driven Gait Representation Learning for Video-Based Visible-Infrared Person Re-identificationabstractVideo-based Visible-Infrared person re-identification (VVI-ReID) aims to retrieve the same pedestrian across visible and infrared modalities from video sequences. Existing methods tend to exploit modality-invariant visual features but largely overlook gait features, which are not only modality-invariant but also rich in temporal dynamics, thus limiting their ability to model the spatiotemporal consistency essential for cross-modal video matching. To address these challenges, we propose a DINOv2-Driven Gait Representation Learning (DinoGRL) framework that leverages the rich visual priors of DINOv2 to learn gait features complementary to appearance cues, facilitating robust sequence-level representations for cross-modal retrieval. Specifically, we introduce a Semantic-Aware Silhouette and Gait Learning (SASGL) model, which generates and enhances silhouette representations with general-purpose semantic priors from DINOv2 and jointly optimizes them with the ReID objective to achieve semantically enriched and task-adaptive gait feature learning. Furthermore, we develop a Progressive Bidirectional Multi-Granularity Enhancement (PBMGE) module, which progressively refines feature representations by enabling bidirectional interactions between gait and appearance streams across multiple spatial granularities, fully leveraging their complementarity to enhance global representations with rich local details and produce highly discriminative features. Extensive experiments on HITSZ-VCM and BUPT datasets demonstrate the superiority of our approach, significantly outperforming existing state-of-the-art methods. Neng Dong, Fan Li 0006, Huafeng Li 0001 |
ACM Multimedia | 4 |
| 2025 | TriMatch: Triple Matching for Text-to-Image Person Re-IdentificationabstractText-to-image person re-identification (TIReID) is a cross-modal retrieval task that aims to retrieve target person images based on a given text description. Existing methods primarily focus on mining the semantic associations across modalities, relying on the matching between heterogeneous features for retrieval. However, due to the inherent heterogeneous gaps between modalities, it is challenging to establish precise semantic associations, particularly in fine-grained correspondences, often leading to incorrect retrieval results. To address this issue, this letter proposes an innovative Triple Matching (TriMatch) framework that integrates cross-modal (image-text) matching and unimodal (image-image, text-text) matching for high-precision person retrieval. The framework introduces a generation task that performs cross-modal (image-to-text and text-to-image) feature generation and intra-modal feature alig achieve unimodal matching. By incorporating the generation task, TriMatch considers not only the semantic correlations between modalities but also the semantic consistency within single modalities, thereby effectively enhancing the accuracy of target person retrieval. Extensive experiments on multiple datasets demonstrate the superiority of TriMatch over existing methods. Shuanglin Yan, Neng Dong, Huafeng Li 0001 |
IEEE Signal Process. Lett. | 2 |
| 2025 | Diverse Semantics-Guided Feature Alignment and Decoupling for Visible-Infrared Person Re-IdentificationabstractVisible-Infrared Person Re-Identification (VI-ReID) is a challenging task due to the large modality discrepancy between visible and infrared images, which complicates the alignment of their features into a suitable common space. Moreover, style noise, such as illumination and color contrast, reduces the identity discriminability and modality invariance of features. To address these challenges, we propose a novel Diverse Semantics-guided Feature Alignment and Decoupling (DSFAD) network to align identity-relevant features from different modalities into a textual embedding space and disentangle identity-irrelevant features within each modality. Specifically, we develop a Diverse Semantics-guided Feature Alignment (DSFA) module, which generates pedestrian descriptions with diverse sentence structures to guide the cross-modality alignment of visual features. Furthermore, to filter out style information, we propose a Semantic Margin-guided Feature Decoupling (SMFD) module, which decomposes visual features into pedestrian-related and style-related components, and then constrains the similarity between the former and the textual embeddings to be at least a margin higher than that between the latter and the textual embeddings. Additionally, to prevent the loss of pedestrian semantics during feature decoupling, we design a Semantic Consistency-guided Feature Restitution (SCFR) module, which further excavates useful information for identification from the style-related features and restores it back into the pedestrian-related features, and then constrains the similarity between the features after restitution and the textual embeddings to be consistent with that between the features before decoupling and the textual embeddings. Extensive experiments on three VI-ReID datasets demonstrate the superiority of our DSFAD. The code will be made publicly available at https://github.com/nengdong96/DSFAD. Neng Dong, Shuanglin Yan, Liyan Zhang 0001, Jinhui Tang 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | CLIP-Driven Semantic Discovery Network for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VIReID) primarily deals with matching identities across person images from different modalities. Due to the modality gap between visible and infrared images, cross-modality identity matching poses significant challenges. Recognizing that high-level semantics of pedestrian appearance, such as gender, shape, and clothing style, remain consistent across modalities, this paper intends to bridge the modality gap by infusing visual features with high-level semantics. Given the capability of Contrastive Language-Image Pre-training (CLIP) to sense high-level semantic information corresponding to visual representations, we explore the application of CLIP within the domain of VIReID. Consequently, we propose a CLIP-Driven Semantic Discovery Network (CSDN) that consists of Modality-specific Prompt Learner, Semantic Information Integration (SII), and High-level Semantic Embedding (HSE). Specifically, considering the diversity stemming from modality discrepancies in language descriptions, we devise bimodal learnable text tokens to capture modality-private semantic information for visible and infrared images, respectively. Additionally, acknowledging the complementary nature of semantic details across different modalities, we integrate text features from the bimodal language descriptions to achieve comprehensive semantics. Finally, we establish a connection between the integrated text features and the visual features across modalities. This process embed rich high-level semantic information into visual representations, thereby promoting the modality invariance of visual representations. The effectiveness and superiority of our proposed CSDN over existing methods have been substantiated through experimental evaluations on multiple widely used benchmarks. Neng Dong, Liehuang Zhu, Hao Peng 0001, Dapeng Tao |
IEEE Trans. Multim. | 2 |
| 2024 | Prototypical Prompting for Text-to-image Person Re-identificationabstractIn this paper, we study the problem of Text-to-Image Person Re-identification (TIReID), which aims to find images of the same identity described by a text sentence from a pool of candidate images. Benefiting from Vision-Language Pre-training, such as CLIP (Contrastive Language-Image Pretraining), the TIReID techniques have achieved remarkable progress recently. However, most existing methods only focus on instance-level matching and ignore identity-level matching, which involves associating multiple images and texts belonging to the same person. In this paper, we propose a novel prototypical prompting framework (Propot) designed to simultaneously model instance-level and identity-level matching for TIReID. Our Propot transforms the identity-level matching problem into a prototype learning problem, aiming to learn identity-enriched prototypes. Specifically, Propot works by 'initialize, adapt, enrich, then aggregate'. We first use CLIP to generate high-quality initial prototypes. Then, we propose a domain-conditional prototypical prompting (DPP) module to adapt the prototypes to the TIReID task using task-related information. Further, we propose an instance-conditional prototypical prompting (IPP) module to update prototypes conditioned on intra-modal and inter-modal instances to ensure prototype diversity. Finally, we design an adaptive prototype aggregation module to aggregate these prototypes, generating final identity-enriched prototypes. With identity-enriched prototypes, we diffuse its rich identity information to instances through prototype-to-instance contrastive loss to facilitate identity-level matching. Extensive experiments conducted on three benchmarks demonstrate the superiority of Propot compared to existing TIReID methods. Shuanglin Yan, Jun Liu 0036, Neng Dong, Liyan Zhang 0002, Jinhui Tang 0001 |
ACM Multimedia | 3 |
| 2024 | Erasing, Transforming, and Noising Defense Network for Occluded Person Re-IdentificationabstractOcclusion perturbation presents a significant challenge in person re-identification (re-ID), and existing methods that rely on external visual cues require additional computational resources and only consider the issue of missing information caused by occlusion. In this paper, we propose a simple yet effective framework, termed Erasing, Transforming, and Noising Defense Network (ETNDNet), which treats occlusion as a noise disturbance and solves occluded person re-ID from the perspective of adversarial defense. In the proposed ETNDNet, we introduce three strategies: Firstly, we randomly erase the feature map to create an adversarial representation with incomplete information, enabling adversarial learning of identity loss to protect the re-ID system from the disturbance of missing information. Secondly, we introduce random transformations to simulate the position misalignment caused by occlusion, training the extractor and classifier adversarially to learn robust representations immune to misaligned information. Thirdly, we perturb the feature map with random values to address noisy information introduced by obstacles and non-target pedestrians, and employ adversarial gaming in the re-ID system to enhance its resistance to occlusion noise. Without bells and whistles, ETNDNet has three key highlights: (i) it does not require any external modules with parameters, (ii) it effectively handles various issues caused by occlusion from obstacles and non-target pedestrians, and (iii) it designs the first GAN-based adversarial defense paradigm for occluded person re-ID. Extensive experiments on six public datasets fully demonstrate the effectiveness, superiority, and practicality of the proposed ETNDNet. The code will be released at https://github.com/nengdong96/ETNDNet. Neng Dong, Liyan Zhang 0001, Shuanglin Yan, Hao Tang 0007, Jinhui Tang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Learning Comprehensive Representations with Richer Self for Text-to-Image Person Re-IdentificationabstractText-to-image person re-identification (TIReID) retrieves pedestrian images of the same identity based on a query text. However, existing methods typically treat it as a one-to-one image-text matching problem, only focusing on the relationship between image-text pairs within a view. The many-to-many matching between image-text pairs across views under the same identity is not taken into account, which is one of the main reasons for the poor performance of existing methods. To this end, we propose a simple yet effective framework, called LCR2S, for modeling many-to-many correspondences of the same identity by learning comprehensive representations for both modalities from a novel perspective. We construct a support set for each image (text) by using other images (texts) under the same identity and design a multi-head attentional fusion module to fuse the image (text) and its support set. The resulting enriched image and text features are aligned to train a "richer" TIReID model with many-to-many correspondences. Since the support set is unavailable during inference, we propose to distill the knowledge learned by the "richer" model into a lightweight model for inference with a single image/text as input. The lightweight model focus on semantic association and reasoning of multi-view information, which can generate a comprehensive representation containing multi-view information with only a single-view input to perform accurate text-to-image retrieval during inference. In particular, we use the intra-modal features and inter-modal semantic relations of the "richer" model to supervise the lightweight model to inherit its powerful capability. Extensive experiments demonstrate the effectiveness of LCR2S, and it also achieves new state-of-the-art performance on three popular TIReID datasets. Shuanglin Yan, Neng Dong, Jun Liu 0036, Liyan Zhang 0002, Jinhui Tang 0001 |
ACM Multimedia | 2 |
| 2023 | CLIP-Driven Fine-Grained Text-Image Person Re-IdentificationabstractText-Image Person Re-identification (TIReID) aims to retrieve the image corresponding to the given text query from a pool of candidate images. Existing methods employ prior knowledge from single-modality pre-training to facilitate learning, but lack multi-modal correspondence information. Vision-Language Pre-training, such as CLIP (Contrastive Language-Image Pretraining), can address the limitation. However, CLIP falls short in capturing fine-grained information, thereby not fully leveraging its powerful capacity in TIReID. Besides, the popular explicit local matching paradigm for mining fine-grained information heavily relies on the quality of local parts and cross-modal inter-part interaction/guidance, leading to intra-modal information distortion and ambiguity problems. Accordingly, in this paper, we propose a CLIP-driven Fine-grained information excavation framework (CFine) to fully utilize the powerful knowledge of CLIP for TIReID. To transfer the multi-modal knowledge effectively, we conduct fine-grained information excavation to mine modality-shared discriminative details for global alignment. Specifically, we propose a multi-level global feature learning (MGF) module that fully mines the discriminative local information within each modality, thereby emphasizing identity-related discriminative clues through enhanced interaction between global image (text) and informative local patches (words). MGF generates a set of enhanced global features for later inference. Furthermore, we design cross-grained feature refinement (CFR) and fine-grained correspondence discovery (FCD) modules to establish cross-modal correspondence at both coarse and fine-grained levels (image-word, sentence-patch, word-patch), ensuring the reliability of informative local patches/words. CFR and FCD are removed during inference to optimize computational efficiency. Extensive experiments on multiple benchmarks demonstrate the superior performance of our method in TIReID. Shuanglin Yan, Neng Dong, Liyan Zhang 0001, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Triple Adversarial Learning and Multi-View Imaginative Reasoning for Unsupervised Domain Adaptation Person Re-IdentificationabstractDue to the importance of practical applications, unsupervised domain adaptation (UDA) person re-identification (re-ID) has attracted increasing attention. However, most of existing methods often lack the multi-view information reasoning and ignore the domain discrepancy of the pedestrian images with the same identity, which constrain the further improvement of recognition performance. So, this paper proposes a triple adversarial learning and multi-view imaginative reasoning network (TAL-MIRN) for UDA person re-ID, which consists of a multi-view imaginative reasoning module (IRM) and a triple adversarial learning module (TALM). IRM makes the classified pedestrian identity features from a single-view image extracted by a feature encoder consistent with the classification results of the aggregated multi-view pedestrian identity features, so the strong multi-view imaginative reasoning ability of the feature encoder is obtained. TALM is composed by the adversarial learning between the camera classifier and feature encoder, adversarial learning of joint distribution alignment, and adversarial learning of the difference between two classifiers used in classification. In particular, the domain-invariant features at camera level are guaranteed by the adversarial learning between the feature extractor and camera classifier. The joint alignment of identity and domain is achieved by the competition between the feature extractor and classifier integrated with identity and domain. The discriminability and robustness of the learned features are enhanced by playing a MinMax game between two different identity classifiers. Furthermore, a simple normalization operation named as cross normalization (CN) is proposed to increase both modeling and generalization capability of the proposed TAL-MIRN across multiple domains. The proposed TAL-MIRN is applied to five benchmark datasets, and the comparative experimental results confirm its superiority over the state-of-the-art methods. The related source codes is available athttps://github.com/lhf12278/TALM-IRM. Huafeng Li 0001, Neng Dong, Zhengtao Yu 0001, Dapeng Tao, Guanqiu Qi |
IEEE Trans. Circuits Syst. Video Technol. | 2 |