EDBT 2026 Demo / reviewers in the wild / expert
Yujian Feng
dblp:195/8994
· DBLP profile ↗
16ranked-venue papers
6as first author
16since 2021 · last 2025
0000-0003-1051-7217ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Saliency Feature Learning for Multimodality Person Retrieval in Visual Internet of ThingsabstractIn the visual Internet of Things (VIoT), smart surveillance is an important component of the multimodality person retrieval (MPR) task. Capturing discriminative pedestrian information from images aids in identifying target individuals through sketch or text descriptions, improving the reliability of VIoT systems. Current granularity-level matching methods face challenges with information redundancy and modality differences in MPR tasks. In this article, we propose a novel approach for intra and intermodality saliency feature learning by multigranularity feature selection and relation (MFSR). First, to reduce redundant information within each modality, the intramodality feature selection module (IFSM) employs the adaptive weighting mechanism to enhance salient pedestrian features while suppressing irrelevant features. Second, the multigranularity feature relation module (MFRM) aligns cross-modality salient person features by increasing the similarity scores of local and global features across modalities, to reduce differences between multimodality visual-text pairs. Finally, the cross-modality similarity matching (CSM) loss is designed to enhance consistency in visual-text pairs of the same identity, ensuring compact intraclass features by minimizing discrepancies between cross-modality similarity distributions and identity-matching distributions. Experimental results show that our approach achieves state-of-the-art performance on benchmark datasets. Shuai You, Yujian Feng, Shitao Wang, Fei Wu 0004, Yuchen Sha, Yimu Ji 0001, Xiaoyuan Jing |
IEEE Internet Things J. | 2 |
| 2025 | Semi-supervised cross-modality person re-identification based on pseudo label learning
Fei Wu 0004, Ruixuan Zhou, Yang Gao 0001, Yujian Feng, Qinghua Huang, Xiaoyuan Jing |
Image Vis. Comput. | 4 |
| 2025 | Learning multi-granularity representation with transformer for visible-infrared person re-identification
Yujian Feng, Feng Chen 0047, Guozi Sun, Fei Wu 0004, Yimu Ji 0001, Tianliang Liu, Shangdong Liu, Xiaoyuan Jing, Jiebo Luo 0001 |
Pattern Recognit. | 1 |
| 2025 | Homogeneous and heterogeneous relational graph for visible-infrared person re-identification
Yujian Feng, Feng Chen 0047, Jian Yu 0007, Yimu Ji 0001, Fei Wu 0004, Shangdong Liu, Xiaoyuan Jing |
Pattern Recognit. | 1 |
| 2025 | Diverse Co-Saliency Feature Learning for Text-Based Person RetrievalabstractText-based Person Retrieval (TPR) plays a pivotal role in video surveillance systems for safeguarding public safety. As a fine-grained retrieval task, TPR faces the significant challenge of precisely capturing highly discriminative features across image and text modalities. Existing methods primarily focus on establishing modality-shared feature spaces to bridge cross-modal discrepancies. However, these methods are prone to disturbances from irrelevant information, such as background noises in the visual modality, and often over-emphasize specific local regions while neglecting the capture of diverse discriminative modal features, thereby limiting the robustness of cross-modal matching. In this paper, we introduce a novel framework, termed the Diverse Co-saliency Feature Learning Network (DCFL), which mines the co-saliency information between image and text modalities and enhances the diversity of cross-modal discriminative features while mitigating the interference of noise. Specifically, to construct cross-modal co-saliency features, we devise the Intra-modal Saliency Feature Learning (ISFL) and Cross-modal Saliency Feature Matching (CSFM) modules. ISFL employs a weighted mask mechanism to guide the model in reducing the impact of noise information in both modalities. Complementing ISFL, CSFM establishes consistent relationships between saliency features across modalities, leveraging text descriptions to align pedestrian-relevant visual regions. Furthermore, we propose the Diverse Co-saliency Feature Mining (DCFM) to bolster the diversity of discriminative co-saliency features across both image and text modalities. This module integrates a diversity regularization term, enabling the extraction of varied visual cues and capturing comprehensive features of the target individual. Extensive benchmark experiments demonstrate a substantial superiority of our approach over the state-of-the-art methods. The code will be released publicly. Shuai You, Cuiqun Chen, Yujian Feng, Hai Liu 0006, Yimu Ji 0001, Mang Ye |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Semantic Distillation and Structural Alignment Network for Fake News DetectionabstractIn recent years, the rapid proliferation of multi-modal fake news has posed potential harm across various sectors of society, making the detection of multi-modal fake news crucial. Most existing methods can not effectively reduce the redundant information and preserve both semantic and structural information. To address these problems, this paper proposes a semantic distillation and structural alignment (SDSA) network. We design an semantic distillation module for modality-specific features to preserve task-relevant semantic information and eliminate redundant information. Then, we propose a triple similarity alignment module to preserve structural information. Specifically, intra-modal similarity alignment mines intra-modal consistency by preserving the neighborhood structure within each modality, inter-modal similarity alignment explores cross-modality consistency by bringing the cross-modality feature neighborhood structures, and joint similarity alignment aims to preserve the structural information of fused features. Experiments conducted on two widely used fake news datasets demonstrate that the SDSA method outperforms state-of-the-art approaches. Shangdong Liu, Xiaofan Yue, Fei Wu 0004, Yujian Feng, Yimu Ji 0001 |
ICASSP | 5 |
| 2024 | Cycle mapping with adversarial event classification network for fake news detection
Fei Wu 0004, Yujian Feng, Guangwei Gao, Yimu Ji 0001, Xiaoyuan Jing |
Multim. Tools Appl. | 3 |
| 2024 | Cross-Modality Spatial-Temporal Transformer for Video-Based Visible-Infrared Person Re-IdentificationabstractVideo-based visible-infrared person re-identification (VVI-ReID) aims to match the identity of a person captured in video sequences from both visible and infrared cameras. The VVI-ReID task requires considering both the spatial relationship between body parts within each frame and the temporal change of appearance between successive frames. Existing VVI Re-ID methods employ Convolutional Neural Networks to extract local spatial features and Long Short-Term Memory to form temporal associations. However, these methods can not effectively capture the global spatial feature and the long-range temporal dependencies in ultra-long sequences. In this paper, we propose a Cross-modality Spatial-temporal Transformer (CST) including a Cross-frame Tube Transformer Module (CTTM) and a Multi-frame Transformer Fusion Module (MTFM) to address these challenges. Firstly, CTTM tokenizes a video clip into multiple 3D tubes, each encapsulating local spatial-temporal information of pedestrians, and then obtains global spatial-temporal representations by establishing the relationship between tubes. Secondly, we design MTFM to exchange information between multiple frames using message tokens, thus modeling the long-range temporal dependencies of features of pedestrians. In addition, to prevent the potential representation collapse caused by triplet-based loss functions, we propose a diversity-consistency (DC) loss function to preserve the diversity and consistency of cross-modality feature representations by imposing variance, invariance, and covariance constraints in feature representations. Extensive benchmark experiments demonstrate that our approach outperforms the state-of-the-art methods with large margins. Yujian Feng, Feng Chen 0047, Jian Yu 0007, Yimu Ji 0001, Fei Wu 0004, Tianliang Liu, Shangdong Liu, Xiaoyuan Jing, Jiebo Luo 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Feature Aggregation via Attention Mechanism for Visible-Thermal Person Re-IdentificationabstractVisible-thermal person re-identification (VT-ReID) is an image retrieval task that aims at matching the target pedestrian across the visible and thermal modalities. However, intra-class variations and cross-modality discrepancy degrade the performance of VT-ReID. Recent methods focus on extracting discriminative local features of each modality to alleviate the intra-class variations and cross-modality discrepancy, but these methods ignore semantic relations between the local features of two modalities,i.e., the spatial relations and channel relations. In this paper, we proposed a feature aggregation module (FAM) to enhance the correlation between local features including spatial dependencies and channel dependencies. Furthermore, FAM implements cross-modality feature aggregation on the enhanced features to reduce the cross-modality discrepancy. Moreover, we also proposed near neighbor cross-modality loss (NNCLoss) to mine feature consistency between modalities by constructing a cross-modality near neighbor set, which facilitates feature alignment between two modalities. Extensive experiments on two datasets demonstrate the superior performance of our approach over the existing state-of-the-arts. Baotai Wu, Yujian Feng, Yunfei Sun, Yimu Ji 0001 |
IEEE Signal Process. Lett. | 2 |
| 2023 | Multi-Scale Aggregation Transformers for Multispectral Object DetectionabstractMultispectral object detection for autonomous driving is multi-object localization and classification task on visible and thermal modalities. In this scenario, modality differences lead to the lack of object information in a single modality and the misalignment of cross-modality information. To alleviate these problems, most existing methods extract information based on a single scale (e.g., these methods mainly focus on detecting significant cars or pedestrians), which leads to insufficient performance in capturing multi-scale discriminative information (e.g., small bicycles and blurred pedestrians) and safety hazards in the driving process. In this paper, we propose a Multi-Scale Aggregation Network (MSANet) consisting of two parts Multi-Scale Aggregation Transformer (MSAT) and the Cross-modal Merging Fusion Mechanism (CMFM), which combined with the advantages of Transformer and CNN to extract rich image information from two modalities by mining both local and global context dependencies. Firstly, to reduce the lack of information in a single modality, we design a novel MSAT module to extract rich details and texture from multi-scale. Secondly, to alleviate feature misalignment caused by modality differences, the CMFM is utilized to aggregate complementary information on multiple levels. Comprehensive experiments on two benchmarks demonstrate that our approach shows better results than several state-of-the-art methods. The code is available athttps://github.com/ysh-strive/MSANet. Shuai You, Xuedong Xie, Yujian Feng, Chaojun Mei, Yimu Ji 0001 |
IEEE Signal Process. Lett. | 3 |
| 2023 | Occluded Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) aims to match person images between the visible and near-infrared modalities. Previous VI-ReID methods are based on holistic pedestrian images and achieve excellent performance. However, in real-world scenarios, images captured by visible and near-infrared cameras usually contain occlusions. The performance of these methods degrades significantly due to the loss of information of discriminative features from the occlusion of the images. We define visible-infrared person re-identification in this occlusion scene as Occluded VI-ReID, where only partial content information of pedestrian images can be used to match images of different modalities from different cameras. In this paper, we propose a matching framework for occlusion scenes, which contains a local feature enhance module (LFEM) and a modality information fusion module (MIFM). LFEM adopts Transformer to learn features of each modality, and adjusts the importance of patches to enhance the representation ability of local features of the non-occluded areas. MIFM utilizes a co-attention mechanism to infer the correlation between each image for reducing the difference between modalities. We construct two occluded VI-ReID datasets, namely Occluded-SYSU-MM01 and Occluded-RegDB datasets. Our approach outperforms existing state-of-the-art methods on two occlusion datasets, while remains top performance on two holistic datasets. Yujian Feng, Yimu Ji 0001, Fei Wu 0004, Guangwei Gao, Yang Gao 0001, Tianliang Liu, Shangdong Liu, Xiaoyuan Jing, Jiebo Luo 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Visible-Infrared Person Re-Identification via Cross-Modality Interaction TransformerabstractVisible-infrared person re-identification (VI Re-ID) is designed to match person images of the same identity from visible and infrared cameras. Transformer structures have been successfully applied in the field of VI Re-ID. However, previous Transformer-based methods were mainly designed to capture global content information in a single modality, and could not simultaneously perceive semantic information between two modalities from a global perspective. To solve this problem, we propose a novel framework named the cross-modality interaction Transformer (CMIT). It has strong abilities in modeling spatial and sequential features that can capture dependencies between long-range features, and explicitly improves the discriminativeness of features by exchanging information across modalities, thus contributing to obtaining modality-invariant representations. Specifically, CMIT utilizes a cross-modality attention mechanism to enrich the feature representations of each patch token by interacting with the patch tokens of the other modality, and aggregates local features of the CNN structure and global information of the Transformer structure to mine feature saliency representation. Furthermore, the modality-discriminative (MD) loss function is proposed to learn potential consistency between modalities to encourage intra-modality compactness within class and inter-modality separation between classes. Extensive experiments on two benchmarks demonstrate that our approach outperforms state-of-the-art methods. Yujian Feng, Jian Yu 0007, Feng Chen 0047, Yimu Ji 0001, Fei Wu 0004, Shangdong Liu, Xiaoyuan Jing |
IEEE Trans. Multim. | 1 |
| 2022 | Part-facial relational and modality-style attention networks for heterogeneous face recognition
Jian Yu 0007, Yujian Feng, Yang Gao 0001 |
Neurocomputing | 2 |
| 2022 | An Intrinsic Structured Graph Alignment Module With Modality-Invariant Representations for NIR-VIS Face RecognitionabstractMost existing near-infrared to visible (NIR-VIS) face recognition (FR) methods rely on global feature representations to reduce cross-modality discrepancies, but ignore the structural relationships between local features, e.g., the relative positions of eyes, nose, and mouth. Precise alignment of these local features can enhance the learning of modality-invariant face representations, thereby improving the performance of NIR-VIS FR. Therefore, in this letter, we propose an intrinsic structured graph alignment (ISGA) module that aims to obtain a graph-level local feature alignment across modalities. To this end, we first construct an intrinsic structure graph to model the inherent structural relationships of local features, and then enhance the discriminative feature representation by aligning the graphs between modalities. To jointly encourage cross-modality class consistency between semantics and structural relationships, a cross-modality class distribution (CMCD) loss is proposed to add an identity-preserving constraint to each class distribution in the embedding space between two modalities. To solve the resulting problem of suppressed class divisibility, we maximize the mutual information between inputs and class predictions. Extensive experiments on challenging NIR-VIS datasets indicate that our approach outperforms the state-of-the-arts. The code is available athttps://github.com/JianYu777/ISGA-CMCD. Jian Yu 0007, Yujian Feng |
IEEE Signal Process. Lett. | 2 |
| 2021 | Spectrum-aware discriminative deep feature learning for multi-spectral face recognition
Fei Wu 0004, Xiaoyuan Jing, Yujian Feng, Yimu Ji 0001, Ruchuan Wang 0001 |
Pattern Recognit. | 3 |
| 2021 | Efficient Cross-Modality Graph Reasoning for RGB-Infrared Person Re-IdentificationabstractThe modality and pose variance between RGB and infrared (IR) images are two key challenges for RGB-IR person re-identification. Existing methods mainly focus on leveraging pixel or feature alignment to handle the intra-class variations and cross-modality discrepancy. However, these methods are hard to keep semantic identity consistency between global and local representation, which the consistency is important for the cross-modality pedestrian re-identification task. In this work, we propose a novel cross-modality graph reasoning method (CGRNet) to globally model and reason over relations between modalities and context, and to keep semantic identity consistency between global and local representation. Specifically, we propose a local modality-similarity module to put the distribution of modality-specific features into a common subspace without losing identity information. Besides, we squeeze the input feature of RGB and IR images into a channel-wise global vector, and through graph reasoning, the identity relationship and modality relationship in each vector are inferred. Extensive experiments on two datasets demonstrate the superior performance of our approach over the existing state-of-the-art. The code is available athttps://github.com/fegnyujian/CGRNet. Yujian Feng, Feng Chen 0047, Yimu Ji 0001, Fei Wu 0004 |
IEEE Signal Process. Lett. | 1 |