Tengfei Liang

dblp:209/5413 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Federated Privacy Re-identification via Frequency Domain Splitting
Xuanwen Su, Xu Wang 0053, Tengfei Liang, Yi Jin 0001, Yidong Li
ICIG (3)3
2025 CFF: Coarse-to-Fine-to-Fusion Semantic Prototype Generation for Zero-Shot Classification
abstract
Zero-Shot Learning focuses on recognizing images from unseen classes with the model trained only on seen classes and auxiliary information. Auxiliary information represents the semantic concepts of classes and is crucial for unseen class generalization. Existing works have tried human-annotated attributes, word embedding, or texts as auxiliary information. However, these auxiliary information is semantically insufficient and visually misaligned, constraining model performance. In this work, we propose the Coarse-to-Fine-to-Fusion prototype generation network (CFF). To obtain vision-oriented text corpora, we design the Coarse-to-Fine Text Generation (CFTG) paradigm, utilizing large language models to generate coarse- and fine-grained texts. For text embedding and fusion, we propose the Semantic Prototype Generation (SPG) module, fusing learnable prompts with coarse- and fine-grained embedding, enabling fine-grained and fusion prototype generation. Moreover, we propose the Visual-Semantic Alignment (VSA) loss to narrow the domain gap. Extensive experiments on AWA2, SUN, and CUB datasets demonstrate the effectiveness of our method.
Xuanwen Su, Tengfei Liang, Yi Jin 0001, Tao Wang 0011, Yidong Li
ICME3
2025 M3-ReID: Unifying Multi-View, Granularity, and Modality for Video-Based Visible-Infrared Person Re-Identification
abstract
Video-based visible-infrared person re-identification (VVI-ReID) task focuses on cross-modality retrieval of pedestrian videos, which are captured in visible and infrared modalities by non-overlapping cameras across diverse scenes, and holds significant value for security surveillance scenarios. The challenges of this task mainly stem from three issues: the difficulty of capturing comprehensive spatio-temporal cues, intra-class variations within video sequences, and inter-modality discrepancies between visible and infrared data. Existing methods mainly try to address the modality gap or focus on one of the other aspects, but rarely do they jointly consider these key factors. Motivated by these core challenges, we propose the M3-ReID (Multi-View & Granularity & Modality) method, a unified framework that simultaneously enhances spatio-temporal feature extraction, intra-class discrimination, and cross-modality consistency. Specifically, to capture diverse spatio-temporal patterns, we design a Multi-View Learning module that leverages different spatial and temporal-spatial perspectives to adaptively emphasize diverse key regions and motion cues. To enhance intra-class modeling of each identity, we introduce a Multi-Granularity Representation strategy that optimizes features across both fine-grained frame level and coarse-grained video level by minimizing mutual information among redundant frames while enhancing identity representations. Furthermore, to bridge the visible-infrared gap, we propose a Multi-Modality Alignment mechanism that explicitly aligns metric learning and cross-modality matching goals, transforming features into a unified embedding space with modality consistency and class discrimination. Extensive experiments on benchmark VVI-ReID datasets demonstrate the superiority of our proposed M3-ReID framework against existing methods.
Tengfei Liang, Yi Jin 0001, Zhun Zhong, Xin Chen 0003, Xianjia Meng, Tao Wang 0011, Yidong Li
IEEE Trans. Inf. Forensics Secur.1
2024 Bridging the Gap: Multi-Level Cross-Modality Joint Alignment for Visible-Infrared Person Re-Identification
abstract
Visible-Infrared person Re-IDentification (VI-ReID) is a challenging cross-modality image retrieval task that aims to match pedestrians’ images across visible and infrared cameras. To solve the modality gap, existing mainstream methods adopt a learning paradigm converting the image retrieval task into an image classification task with cross-entropy loss and auxiliary metric learning losses. These losses follow the strategy of adjusting the distribution of extracted embeddings to reduce the intra-class distance and increase the inter-class distance. However, such objectives do not precisely correspond to the final test setting of the retrieval task, resulting in a new gap at the optimization level. By rethinking these keys of VI-ReID, we propose a simple and effective method, the Multi-level Cross-modality Joint Alignment (MCJA), bridging both the modality and objective-level gap. For the former, we design the Visible-Infrared Modality Coordinator in the image space and propose the Modality Distribution Adapter in the feature space, effectively reducing modality discrepancy of the feature extraction process. For the latter, we introduce a new Cross-Modality Retrieval loss. It is the first work to constrain from the perspective of the ranking list in the VI-ReID, aligning with the goal of the testing stage. Moreover, to strengthen the robustness and cross-modality retrieval ability, we further introduce a Multi-Spectral Enhanced Ranking strategy for the testing phase. Based on the global feature only, our method outperforms existing methods by a large margin, achieving the remarkable rank-1 of 89.51% and mAP of 87.58% on the most challenging single-shot setting and all-search mode of the SYSU-MM01 dataset.
Tengfei Liang, Yi Jin 0001, Wu Liu 0005, Tao Wang 0011, Songhe Feng, Yidong Li
IEEE Trans. Circuits Syst. Video Technol.1
2024 A Generative-Based Image Fusion Strategy for Visible-Infrared Person Re-Identification
abstract
Cross-modality person re-identification task is a challenging task aiming to recognize images of the same identity between different modalities. To alleviate the cross-modality discrepancies between images, existing approaches mainly guide models to mine modality invariant features. Although those approaches are effective, they lose the modality-specific features that include important information beneficial to VI-ReID. Therefore, some approaches are using generative adversarial networks to compensate for modality information. However, the quality of images generated by these methods is usually poor, and most of them focus only on the learning of modality-sharable features. To solve these problems, this paper proposes a generative-based cross-modality image fusion strategy (GC-IFS), which can generate high-quality cross-modality paired images and fuse the information of the two modalities. Firstly, considering the importance of the identity discriminative information of the generated image, we propose a contrastive-learning image generation (CLIG) network to generate cross-modality paired images. Meanwhile, to fully integrate and utilize the information of the two modalities and eliminate the influence of cross-modality discrepancies, we design a part-based dual multi-modality feature fusion (P-DMFF) module to extract the unified feature representation. Extensive experiments on SYSU-MM01 and RegDB datasets demonstrate that our strategy outperforms the state-of-the-art methods for the VI-ReID task.
Jia Qi, Tengfei Liang, Wu Liu 0005, Yidong Li, Yi Jin 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 MMI-Det: Exploring Multi-Modal Integration for Visible and Infrared Object Detection
abstract
The Visible-Infrared (VIS-IR) object detection is a challenging detection task, which combines visible and infrared data to give information on the category and location of objects in the scene. Therefore, the core of this task is to combine complementary information in the visible and infrared modalities to provide more object detection results for detection. The existing methods mainly face the problem of insufficient ability to perceive and combine visible-infrared modal information and have difficulty in balancing the optimization directions of the fusion and detection tasks. To solve these problem, we propose the MMI-Det which is a multi-modal fusion method for visible and infrared object detection. The method can provide a good combination of complementary information in the visible-infrared modalities and output accurate and robust object information. Specifically, to improve the ability of the model to perceive environment at the visible-infrared image level, we designed the Contour Enhancement Module. Furthermore, to extract complementary information from VIS and IR modalities, we design the Fusion Focus Module. It can extract different frequency spectral features of the visible and infrared modalities and focus on the key information of the object at different spatial locations. Moreover, we design the Contrast Bridge Module to improve the ability to extract modal invariant features in the visible-infrared scene. Finally, to ensure that our model can balance the optimization directions of image fusion and object detection, we design the Info Guided Module as a way to improve the effectiveness of the model’s training optimization. We implement extensive experiments on the public FLIR, M3FD, LLVIP, TNO and MSRS datasets, and compared with previous methods, our method achieves better performance with powerful multi-modal information perception capabilities.
Yuqiao Zeng, Tengfei Liang, Yi Jin 0001, Yidong Li
IEEE Trans. Circuits Syst. Video Technol.2
2023 MetaZSCIL: A Meta-Learning Approach for Generalized Zero-Shot Class Incremental Learning
abstract
Generalized zero-shot learning (GZSL) aims to recognize samples whose categories may not have been seen at training. Standard GZSL cannot handle dynamic addition of new seen and unseen classes. In order to address this limitation, some recent attempts have been made to develop continual GZSL methods. However, these methods require end-users to continuously collect and annotate numerous seen class samples, which is unrealistic and hampers the applicability in the real-world. Accordingly, in this paper, we propose a more practical and challenging setting named Generalized Zero-Shot Class Incremental Learning (CI-GZSL). Our setting aims to incrementally learn unseen classes without any training samples, while recognizing all classes previously encountered. We further propose a bi-level meta-learning based method called MetaZSCIL to directly optimize the network to learn how to incrementally learn. Specifically, we sample sequential tasks from seen classes during the offline training to simulate the incremental learning process. For each task, the model is learned using a meta-objective such that it is capable to perform fast adaptation without forgetting. Note that our optimization can be flexibly equipped with most existing generative methods to tackle CI-GZSL. This work introduces a feature generative framework that leverages visual feature distribution alignment to produce replayed samples of previously seen classes to reduce catastrophic forgetting. Extensive experiments conducted on five widely used benchmarks demonstrate the superiority of our proposed method.
Tengfei Liang, Songhe Feng, Yi Jin 0001, Gengyu Lyu, Haojun Fei, Yang Wang 0003
AAAI2
2023 Cross-Modality Transformer With Modality Mining for Visible-Infrared Person Re-Identification
abstract
The visible-infrared person re-identification (VI-ReID) is a challenging ReID task, which aims to retrieve and match the same identity's images between the heterogeneous visible and infrared modalities. Thus, the core of this task is to bridge the huge gap between these two modalities. The existing methods mainly face the problem of insufficient perception of modality information, and can not learn good discriminative modality-invariant embeddings for identities, which limits their performance. To solve these problems, we propose a new cross-modality transformer-based method (CMTR) for this visible-infrared person re-identification task, which can explicitly mine the information of each modality and generate better discriminative features based on it. Specifically, to capture inherent characteristics of modalities, we design the novel modality embeddings, which are fused with token embeddings to encode modality information directly. Moreover, to enhance representation of modality embeddings and adjust the distribution of embeddings, we further propose a modality-aware enhancement loss based on the learned modality information, which contains two components to reduce intra-class distance and enlarging inter-class distance simultaneously. To our knowledge, this is the first exploration of applying pure transformer network to the cross-modality re-identification task. We implement extensive experiments on the public SYSU-MM01 and RegDB datasets, and compared with previous methods, our method achieves good performance with more compact and powerful embeddings for the cross-modality retrieval.
Tengfei Liang, Yi Jin 0001, Wu Liu 0005, Yidong Li
IEEE Trans. Multim.1
2022 Boundary Corrected Multi-Scale Fusion Network for Real-Time Semantic Segmentation
abstract
Image semantic segmentation aims at the pixel-level classification of images, which has requirements for both accuracy and speed in practical application. Existing semantic segmentation methods mainly rely on the high-resolution input to achieve high accuracy and do not meet the requirements of inference time. Although some methods focus on high-speed scene parsing with lightweight architectures, they can not fully mine semantic features under low computation with relatively low performance. To realize the real-time and high-precision segmentation, we propose a new method named Boundary Corrected Multi-scale Fusion Network, which uses the designed Low-resolution Multi-scale Fusion Module to extract semantic information. Moreover, to deal with boundary errors caused by low-resolution feature map fusion, we further design an additional Boundary Corrected Loss to constrain overly smooth features. Extensive experiments show that our method achieves a state-of-the-art balance of accuracy and speed for the real-time semantic segmentation.
Tianjiao Jiang, Yi Jin 0001, Tengfei Liang, Xu Wang 0053, Yidong Li
ICIP3
2022 Camera-Aware Style Separation and Contrastive Learning for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification (ReID) is a challenging task without data annotation to guide discriminative learning. Existing methods attempt to solve this problem by clustering extracted embeddings to generate pseudo labels. However, most methods ignore the intra-class gap caused by camera style variance, and some methods are relatively complex and indirect although they try to solve the negative impact of the camera style on feature distribution. To solve this problem, we propose a camera-aware style separation and contrastive learning method (CA-UReID), which directly separates camera styles in the feature space with the designed camera-aware attention module. It can explicitly divide the learnable feature into camera-specific and camera-agnostic parts, reducing the influence of different cameras. Moreover, to further narrow the gap across cameras, we design a camera-aware contrastive center loss to learn more discriminative embeddings for each identity. Extensive experiments demonstrate the superiority of our method over the state-of-the-art methods on the unsupervised person ReID task.
Tengfei Liang, Yi Jin 0001, Tao Wang 0011, Yidong Li
ICME2
2022 Keypoint-Guided Modality-Invariant Discriminative Learning for Visible-Infrared Person Re-identification
abstract
The visible-infrared person re-identification (VI-ReID) task aims to retrieve images of pedestrians across cameras with different modalities. In this task, the major challenges arise from two aspects: intra-class variations among images of the same identity, and cross-modality discrepancies between visible and infrared images. Existing methods mainly focus on the latter, attempting to alleviate the impact of modality discrepancy, which ignore the former issue of identity variations and achieve limited discrimination. To address both aspects, we propose a Keypoint-guided Modality-invariant Discriminative Learning (KMDL) method, which can simultaneously adapt to intra-ID variations and bridge the cross-modality gap. By introducing human keypoints, our method makes further exploration in the image space, feature space and loss constraints to solve the above issues. Specifically, considering the modality discrepancy in original images, we first design a Hue Jitter Augmentation (HJA) strategy, introducing the hue disturbance to alleviate color dependence in the input stage. To obtain discriminative fine-grained representation for retrieval, we design the Global-Keypoint Graph Module (GKGM) in feature space, which can directly extract keypoint-aligned features and mine relationships within global and keypoint embeddings. Based on these semantic local embeddings, we further propose the Keypoint-Aware Center (KAC) loss that can effectively adjust the feature distribution under the supervision of ID and keypoint to learn discriminative representation for the matching. Extensive experiments on SYSU-MM01 and RegDB datasets demonstrate the effectiveness of our KMDL method.
Tengfei Liang, Yi Jin 0001, Wu Liu 0005, Songhe Feng, Tao Wang 0011, Yidong Li
ACM Multimedia1
2021 MSO: Multi-Feature Space Joint Optimization Network for RGB-Infrared Person Re-Identification
abstract
The RGB-infrared cross-modality person re-identification (ReID) task aims to recognize the images of the same identity between the visible modality and the infrared modality. Existing methods mainly use a two-stream architecture to eliminate the discrepancy between the two modalities in the final common feature space, which ignore the single space of each modality in the shallow layers. To solve it, in this paper, we present a novel multi-feature space joint optimization (MSO) network, which can learn modality-sharable features in both the single-modality space and the common space. Firstly, based on the observation that edge information is modality-invariant, we propose an edge features enhancement module to enhance the modality-sharable features in each single-modality space. Specifically, we design a perceptual edge features (PEF) loss after the edge fusion strategy analysis. According to our knowledge, this is the first work that proposes explicit optimization in the single-modality feature space on cross-modality ReID task. Moreover, to increase the difference between cross-modality distance and class distance, we introduce a novel cross-modality contrastive-center (CMCC) loss into the modality-joint constraints in the common feature space. The PEF loss and CMCC loss jointly optimize the model in an end-to-end manner, which markedly improves the network's performance. Extensive experiments demonstrate that the proposed model significantly outperforms state-of-the-art methods on both the SYSU-MM01 and RegDB datasets.
Yajun Gao, Tengfei Liang, Yi Jin 0001, Xiaoyan Gu 0001, Wu Liu 0005, Yidong Li, Congyan Lang
ACM Multimedia2