Zhan-Xiang Feng

dblp:168/2151 · also Zhanxiang Feng · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Learning boost-inhibition for weakly supervised visible-infrared group re-identification
Ling Mei 0001, Zhan-Xiang Feng, Jian-Huang Lai, Yiwei Cheng, Peiying Zhang 0001, Jian Wang 0010
Pattern Recognit.2
2025 Perspective Driven Prototype Alignment for Aerial-Ground Person Re-identification
Yuli Huang, Zhan-Xiang Feng, Jian-Huang Lai
ICIG (2)3
2024 Salient Part-Aligned and Keypoint Disentangling Transformer for Person Re-Identification in Aerial Imagery
abstract
Person re-identification (Re-ID) in aerial imagery aims to retrieve individuals utilizing the UAV surveillance platform. However, the unpredictable changing views of UAVs result in failing to attend consistent foreground regions and achieve semantic alignment. Existing works are limited by coarse-grained part-aligned division and disturbance-sensitive self-attention mechanism. To address these issues, we propose a transformer-based Salient Part-Aligned and Keypoint Disentangling (SPAKD) framework to focus on salient human body regions and align semantic parts with the assistance of keypoints, which consists of a Salient Part-Aware Cross-Attention (SPACA) module and a Keypoint-Assisted Decoder (KAD) module. Specifically, SPACA enhances relationships between salient foreground regions and the whole image, thereby obtaining discriminative part features. KAD employs learnable part prototypes to disentangle human keypoints and aligns decoupled keypoint-assisted features with salient foreground parts based on their affinity information. Extensive experiments demonstrate that our method achieves state-of-the-art performance on aerial-based person Re-ID datasets.
Junyang Qiu, Zhan-Xiang Feng, Jian-Huang Lai
ICME2
2024 Uncertainty Modeling for Group Re-Identification
Jian-Huang Lai, Zhan-Xiang Feng, Xiaohua Xie
Int. J. Comput. Vis.3
2023 TANet: Adversarial Network via Tokens Transformer for Universal Domain Adaptation
Zhan-Xiang Feng, Jian-Huang Lai
ICIG (1)2
2023 Patch-based camera-aware person-to-group learning and group similarity strategy for unsupervised group re-identification
Lisha Yu, Sien Huang, Jian-Huang Lai, Zhan-Xiang Feng
Neurocomputing4
2022 Uncertainty Modeling with Second-Order Transformer for Group Re-identification
abstract
Group re-identification (G-ReID) focuses on associating the group images containing the same persons under different cameras. The key challenge of G-ReID is that all the cases of the intra-group member and layout variations are hard to exhaust. To this end, we propose a novel uncertainty modeling, which treats each image as a distribution depending on the current member and layout, then digs out potential group features by random samplings. Based on potential and original group features, uncertainty modeling can learn better decision boundaries, which is implemented by two modules, member variation module (MVM) and layout variation module (LVM). Furthermore, we propose a novel second-order transformer framework (SOT), which is inspired by the fact that the position modeling in the transformer is coped with the G-ReID task. SOT is composed of the intra-member module and inter-member module. Specifically, the intra-member module extracts the first-order token for each member, and then the inter-member module learns a second-order token as a group feature by the above first-order tokens, which can be regarded as the token of tokens. A large number of experiments have been conducted on three available datasets, including CSG, DukeGroup and RoadGroup. The experimental results show that the proposed SOT outperforms all previous state-of-the-art methods.
Jian-Huang Lai, Zhan-Xiang Feng, Xiaohua Xie
AAAI3
2022 Modeling 3D Layout For Group Re-Identification
abstract
Group re-identification (GReID) attempts to correctly associate groups with the same members under different cameras. The main challenge is how to resist the membership and layout variations. Existing works attempt to incorporate layout modeling on the basis of appearance features to achieve robust group representations. However, layout ambiguity is introduced because these methods only consider the 2D layout on the imaging plane. In this paper, we overcome the above limitations by 3D layout modeling. Specifically, we propose a novel 3D transformer (3DT) that reconstructs the relative 3D layout relationship among members, then applies sampling and quantification to preset a series of layout tokens along three dimensions, and selects the corresponding tokens as layout features for each member. Furthermore, we build a synthetic GReID dataset, City1M, including 1.84M images, 45K persons and 11.5K groups with 3D annotations to alleviate data shortages and poor annotations. To the best of our knowledge, 3DT is the first work to address GReID with 3D perspective, and the City1M is the currently largest dataset. Several experiments show the superiority of our 3DT and City1M. Our project has been released on https://github.com/LinlyAC/City1M-dataset.
Kaiheng Dang, Jian-Huang Lai, Zhan-Xiang Feng, Xiaohua Xie
CVPR4
2022 Learning Adaptive Progressive Representation for Group Re-identification
Kuoyu Deng, Zhan-Xiang Feng, Jian-Huang Lai
PRCV (1)2
2022 Seeing Like a Human: Asynchronous Learning With Dynamic Progressive Refinement for Person Re-Identification
abstract
Learning discriminative and rich features is an important research task for person re-identification. Previous studies have attempted to capture global and local features at the same time and layer of the model in a non-interactive manner, which are called synchronous learning. However, synchronous learning leads to high similarity, and further defects in model performance. To this end, we propose asynchronous learning based on the human visual perception mechanism. Asynchronous learning emphasizes the time asynchrony and space asynchrony of feature learning and achieves mutual promotion and cyclical interaction for feature learning. Furthermore, we design a dynamic progressive refinement module to improve local features with the guidance of global features. The dynamic property allows this module to adaptively adjust the network parameters according to the input image, in both the training and testing stage. The progressive property narrows the semantic gap between the global and local features, which is due to the guidance of global features. Finally, we have conducted several experiments on four datasets, including Market1501, CUHK03, DukeMTMC-ReID, and MSMT17. The experimental results show that asynchronous learning can effectively improve feature discrimination and achieve strong performance.
Jian-Huang Lai, Zhan-Xiang Feng, Xiaohua Xie
IEEE Trans. Image Process.3
2021 Attention-Guided Siamese Network for Clothes-Changing Person Re-identification
Zhan-Xiang Feng, Sien Huang, Jian-Huang Lai
ICIG (2)1
2021 Training Person Re-identification Networks with Transferred Images
Junkai Deng, Zhan-Xiang Feng, Peijia Chen, Jian-Huang Lai
PRCV (1)2
2021 Resolution-Aware Knowledge Distillation for Efficient Inference
abstract
Minimizing the computation complexity is essential for the popularization of deep networks in practical applications. Nowadays, most researches attempt to accelerate deep networks by designing new network structure or compressing the network parameters. Meanwhile, transfer learning techniques such as knowledge distillation are utilized to keep the performance of deep models. In this paper, we focus on accelerating deep models and relieving the computation burden by using low-resolution (LR) images as inputs while maintaining competitive performance, which is rarely researched in the current literature. Deep networks may encounter serious performance degradation when using LR inputs because many details are unavailable from LR images. Besides, the existing approaches may fail to learn discriminative features for LR images because of the dramatic appearance variations between LR and high-resolution (HR) images. To tackle with the above problems, we propose a resolution-aware knowledge distillation (RKD) framework to narrow the cross-resolution variations by transferring knowledge from HR domain to LR domain. The proposed framework consists of a HR teacher network and a LR student network. First, we introduce a discriminator and propose an adversarial learning strategy to shrink the variations between inputs with changing resolution. Then we design a cross-resolution knowledge distillation (CRKD) loss to train discriminative student network by exploiting the knowledge of the teacher network. The CRKD loss is consisted of a resolution-aware distillation loss, a pair-wise constraint, and a maximum mean discrepancy loss. Experimental results on person re-identification, image classification, face recognition, and defect segmentation tasks demonstrate that RKD outperforms traditional knowledge distillation method by achieving better performance with lower computation complexities. Furthermore, CRKD surpasses the state-of-the-art knowledge distillation methods in transferring knowledge across different resolutions under RKD framework, especially when coping with large resolution differences.
Zhan-Xiang Feng, Jian-Huang Lai, Xiaohua Xie
IEEE Trans. Image Process.1
2020 Open-World Group Retrieval with Ambiguity Removal: A Benchmark
abstract
Group retrieval has attracted plenty of attention in artificial intelligence, traditional group retrieval researches assume that members in a group are unique and do not change under different cameras. However, the assumption may not be met for practical situations such as open-world and group-ambiguity scenarios. This paper tackles an important yet non-studied problem: re-identifying changing groups of people under the open-world and group-ambiguity scenarios in different camera fields. The open-world scenario considers that there are probably non-target people for the probe set appear in the searching gallery, while the group-ambiguity scenario means the group members may change. The open-world and group-ambiguity issue is very challenging for the existing methods because the changing of group members results in dramatic visual variations. Nevertheless, as far as we know, the existing literature lacks benchmarks which target on coping with this issue. In this paper, we propose a new group retrieval dataset named OWGA-Campus to consider these challenges. Moreover, we propose a person-to-group similarity matching based ambiguity removal (P2GSM-AR) method to solve these problems and realize the intention of group retrieval. Experimental results on OWGA-Campus dataset demonstrate the effectiveness and robustness of the proposed P2GSM-AR approach in improving the performance of the state-of-the-art feature extraction methods of person re-id towards the open-world and ambiguous group retrieval task.
Ling Mei 0001, Jian-Huang Lai, Zhan-Xiang Feng, Xiaohua Xie
ICPR3
2020 Deep Face Recognition Based on Penalty Cosface
Shuoyan Lin, Jianxiong Tang, Zhan-Xiang Feng, Jian-Huang Lai
PRCV (2)3
2020 From pedestrian to group retrieval via siamese network and correlation
Ling Mei 0001, Jian-Huang Lai, Zhan-Xiang Feng, Xiaohua Xie
Neurocomputing3
2020 Learning Modality-Specific Representations for Visible-Infrared Person Re-Identification
abstract
Traditional person re-identification (re-id) methods perform poorly under changing illuminations. This situation can be addressed by using dual-cameras that capture visible images in a bright environment and infrared images in a dark environment. Yet, this scheme needs to solve the visible-infrared matching issue, which is largely under-studied. Matching pedestrians across heterogeneous modalities is extremely challenging because of different visual characteristics. In this paper, we propose a novel framework that employ modality-specific networks to tackle with the heterogeneous matching problem. The proposed framework utilizes the modality-related information and extracts modality-specific representations (MSR) by constructing an individual network for each modality. In addition, a cross-modality Euclidean constraint is introduced to narrow the gap between different networks. We also integrate the modality-shared layers into modality-specific networks to extract shareable information and use a modality-shared identity loss to facilitate the extraction of modality-invariant features. Then a modality-specific discriminant metric is learned for each domain to strengthen the discriminative power of MSR. Eventually, we use a view classifier to learn view information. The experiments demonstrate that the MSR effectively improves the performance of deep networks on VI-REID and remarkably outperforms the state-of-the-art methods.
Zhan-Xiang Feng, Jian-Huang Lai, Xiaohua Xie
IEEE Trans. Image Process.1
2019 Low Resolution Person Re-identification by an Adaptive Dual-Branch Network
Zhan-Xiang Feng, Jian-Huang Lai, Xiaohua Xie
ICIG (1)1
2018 Conditional Face Synthesis for Data Augmentation
Xiaohua Xie, Jian-Huang Lai, Zhan-Xiang Feng
PRCV (3)4
2018 Image super-resolution via a densely connected recursive network
Zhan-Xiang Feng, Jian-Huang Lai, Xiaohua Xie, Jun-Yong Zhu
Neurocomputing1
2018 Learning View-Specific Deep Networks for Person Re-Identification
abstract
In recent years, a growing body of research has focused on the problem of person re-identification (re-id). The re-id techniques attempt to match the images of pedestrians from disjoint non-overlapping camera views. A major challenge of the re-id is the serious intra-class variations caused by changing viewpoints. To overcome this challenge, we propose a deep neural network-based framework which utilizes the view information in the feature extraction stage. The proposed framework learns a view-specific network for each camera view with a cross-view Euclidean constraint (CV-EC) and a cross-view center loss. We utilize the CV-EC to decrease the margin of the features between diverse views and extend the center loss metric to a view-specific version to better adapt the re-id problem. Moreover, we propose an iterative algorithm to optimize the parameters of the view-specific networks from coarse to fine. The experiments demonstrate that our approach significantly improves the performance of the existing deep networks and outperforms the state-of-the-art methods on the VIPeR, CUHK01, CUHK03, SYSU-mReId, and Market-1501 benchmarks.
Zhan-Xiang Feng, Jian-Huang Lai, Xiaohua Xie
IEEE Trans. Image Process.1
2017 Face recognition by landmark pooling-based CNN with concentrate loss
abstract
Face recognition has been a hot research topic in recent years, convolutional neural network (CNN) based methods have achieved state of the art results and significantly improve the performance. Along with the CNN framework, we propose a novel loss function called concentrate loss which focuses on the class centers in the mini-batch. The concentrate loss aims to push the samples towards corresponding class centers and simultaneously enlarge the gap between different class centers. Additionally, we ultilize facial landmark pooling technique to take full advantage of facial structure information. Experiment results on Labeled Faces in the Wild (LFW), YouTube Faces (YTF), and the BluFR benchmark demonstrate the efficiency of our proposal.
Xiaohua Xie, Zhan-Xiang Feng, Jian-Huang Lai
ICIP3
2016 Face hallucination by deep traversal network
abstract
In this paper, we propose a novel patch-based face hallucination method that consists of two patch-based sparse autoencoder (SAE) networks and a deep fully connected network (namely traversal network). The SAE networks are used to capture the intrinsic features of low-resolution (LR) images and high-resolution (HR) images in the hidden layers, while the traversal network is used to map features from the LR hidden layer to the HR hidden layer. In the training stage, these three networks are jointly optimized. Compared with previous network-based methods that learn an end-to-end mapping from LR images to HR images, our method learns the mapping between hidden layers, which can better alleviate the over-fitting problem. Experimental results demonstrate that our method is efficient and robust for hallucinating face images from both lab environment and the wild. The proposal achieves state-of-the-art performance when conducting face hallucination in CAS-PEAL-R1 database, CMU-PIE database and Casia database.
Zhan-Xiang Feng, Jian-Huang Lai, Xiaohua Xie, Dakun Yang, Ling Mei 0001
ICPR1