EDBT 2026 Demo / reviewers in the wild / expert
Zhan-Xiang Feng
dblp:168/2151 · also Zhanxiang Feng
· DBLP profile ↗
23ranked-venue papers
7as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning boost-inhibition for weakly supervised visible-infrared group re-identification
Ling Mei 0001, Zhan-Xiang Feng, Jian-Huang Lai, Yiwei Cheng, Peiying Zhang 0001, Jian Wang 0010 |
Pattern Recognit. | 2 |
| 2025 | Perspective Driven Prototype Alignment for Aerial-Ground Person Re-identification
Yuli Huang, Zhan-Xiang Feng, Jian-Huang Lai |
ICIG (2) | 3 |
| 2024 | Salient Part-Aligned and Keypoint Disentangling Transformer for Person Re-Identification in Aerial ImageryabstractPerson re-identification (Re-ID) in aerial imagery aims to retrieve individuals utilizing the UAV surveillance platform. However, the unpredictable changing views of UAVs result in failing to attend consistent foreground regions and achieve semantic alignment. Existing works are limited by coarse-grained part-aligned division and disturbance-sensitive self-attention mechanism. To address these issues, we propose a transformer-based Salient Part-Aligned and Keypoint Disentangling (SPAKD) framework to focus on salient human body regions and align semantic parts with the assistance of keypoints, which consists of a Salient Part-Aware Cross-Attention (SPACA) module and a Keypoint-Assisted Decoder (KAD) module. Specifically, SPACA enhances relationships between salient foreground regions and the whole image, thereby obtaining discriminative part features. KAD employs learnable part prototypes to disentangle human keypoints and aligns decoupled keypoint-assisted features with salient foreground parts based on their affinity information. Extensive experiments demonstrate that our method achieves state-of-the-art performance on aerial-based person Re-ID datasets. Junyang Qiu, Zhan-Xiang Feng, Jian-Huang Lai |
ICME | 2 |
| 2024 | Uncertainty Modeling for Group Re-Identification
Jian-Huang Lai, Zhan-Xiang Feng, Xiaohua Xie |
Int. J. Comput. Vis. | 3 |
| 2023 | TANet: Adversarial Network via Tokens Transformer for Universal Domain Adaptation
Zhan-Xiang Feng, Jian-Huang Lai |
ICIG (1) | 2 |
| 2023 | Patch-based camera-aware person-to-group learning and group similarity strategy for unsupervised group re-identification
Lisha Yu, Sien Huang, Jian-Huang Lai, Zhan-Xiang Feng |
Neurocomputing | 4 |
| 2022 | Uncertainty Modeling with Second-Order Transformer for Group Re-identificationabstractGroup re-identification (G-ReID) focuses on associating the group images containing the same persons under different cameras. The key challenge of G-ReID is that all the cases of the intra-group member and layout variations are hard to exhaust. To this end, we propose a novel uncertainty modeling, which treats each image as a distribution depending on the current member and layout, then digs out potential group features by random samplings. Based on potential and original group features, uncertainty modeling can learn better decision boundaries, which is implemented by two modules, member variation module (MVM) and layout variation module (LVM). Furthermore, we propose a novel second-order transformer framework (SOT), which is inspired by the fact that the position modeling in the transformer is coped with the G-ReID task. SOT is composed of the intra-member module and inter-member module. Specifically, the intra-member module extracts the first-order token for each member, and then the inter-member module learns a second-order token as a group feature by the above first-order tokens, which can be regarded as the token of tokens. A large number of experiments have been conducted on three available datasets, including CSG, DukeGroup and RoadGroup. The experimental results show that the proposed SOT outperforms all previous state-of-the-art methods. Jian-Huang Lai, Zhan-Xiang Feng, Xiaohua Xie |
AAAI | 3 |
| 2022 | Modeling 3D Layout For Group Re-IdentificationabstractGroup re-identification (GReID) attempts to correctly associate groups with the same members under different cameras. The main challenge is how to resist the membership and layout variations. Existing works attempt to incorporate layout modeling on the basis of appearance features to achieve robust group representations. However, layout ambiguity is introduced because these methods only consider the 2D layout on the imaging plane. In this paper, we overcome the above limitations by 3D layout modeling. Specifically, we propose a novel 3D transformer (3DT) that reconstructs the relative 3D layout relationship among members, then applies sampling and quantification to preset a series of layout tokens along three dimensions, and selects the corresponding tokens as layout features for each member. Furthermore, we build a synthetic GReID dataset, City1M, including 1.84M images, 45K persons and 11.5K groups with 3D annotations to alleviate data shortages and poor annotations. To the best of our knowledge, 3DT is the first work to address GReID with 3D perspective, and the City1M is the currently largest dataset. Several experiments show the superiority of our 3DT and City1M. Our project has been released on https://github.com/LinlyAC/City1M-dataset. Kaiheng Dang, Jian-Huang Lai, Zhan-Xiang Feng, Xiaohua Xie |
CVPR | 4 |
| 2022 | Learning Adaptive Progressive Representation for Group Re-identification
Kuoyu Deng, Zhan-Xiang Feng, Jian-Huang Lai |
PRCV (1) | 2 |
| 2022 | Seeing Like a Human: Asynchronous Learning With Dynamic Progressive Refinement for Person Re-IdentificationabstractLearning discriminative and rich features is an important research task for person re-identification. Previous studies have attempted to capture global and local features at the same time and layer of the model in a non-interactive manner, which are called synchronous learning. However, synchronous learning leads to high similarity, and further defects in model performance. To this end, we propose asynchronous learning based on the human visual perception mechanism. Asynchronous learning emphasizes the time asynchrony and space asynchrony of feature learning and achieves mutual promotion and cyclical interaction for feature learning. Furthermore, we design a dynamic progressive refinement module to improve local features with the guidance of global features. The dynamic property allows this module to adaptively adjust the network parameters according to the input image, in both the training and testing stage. The progressive property narrows the semantic gap between the global and local features, which is due to the guidance of global features. Finally, we have conducted several experiments on four datasets, including Market1501, CUHK03, DukeMTMC-ReID, and MSMT17. The experimental results show that asynchronous learning can effectively improve feature discrimination and achieve strong performance. Jian-Huang Lai, Zhan-Xiang Feng, Xiaohua Xie |
IEEE Trans. Image Process. | 3 |
| 2021 | Attention-Guided Siamese Network for Clothes-Changing Person Re-identification
Zhan-Xiang Feng, Sien Huang, Jian-Huang Lai |
ICIG (2) | 1 |
| 2021 | Training Person Re-identification Networks with Transferred Images
Junkai Deng, Zhan-Xiang Feng, Peijia Chen, Jian-Huang Lai |
PRCV (1) | 2 |
| 2021 | Resolution-Aware Knowledge Distillation for Efficient InferenceabstractMinimizing the computation complexity is essential for the popularization of deep networks in practical applications. Nowadays, most researches attempt to accelerate deep networks by designing new network structure or compressing the network parameters. Meanwhile, transfer learning techniques such as knowledge distillation are utilized to keep the performance of deep models. In this paper, we focus on accelerating deep models and relieving the computation burden by using low-resolution (LR) images as inputs while maintaining competitive performance, which is rarely researched in the current literature. Deep networks may encounter serious performance degradation when using LR inputs because many details are unavailable from LR images. Besides, the existing approaches may fail to learn discriminative features for LR images because of the dramatic appearance variations between LR and high-resolution (HR) images. To tackle with the above problems, we propose a resolution-aware knowledge distillation (RKD) framework to narrow the cross-resolution variations by transferring knowledge from HR domain to LR domain. The proposed framework consists of a HR teacher network and a LR student network. First, we introduce a discriminator and propose an adversarial learning strategy to shrink the variations between inputs with changing resolution. Then we design a cross-resolution knowledge distillation (CRKD) loss to train discriminative student network by exploiting the knowledge of the teacher network. The CRKD loss is consisted of a resolution-aware distillation loss, a pair-wise constraint, and a maximum mean discrepancy loss. Experimental results on person re-identification, image classification, face recognition, and defect segmentation tasks demonstrate that RKD outperforms traditional knowledge distillation method by achieving better performance with lower computation complexities. Furthermore, CRKD surpasses the state-of-the-art knowledge distillation methods in transferring knowledge across different resolutions under RKD framework, especially when coping with large resolution differences. Zhan-Xiang Feng, Jian-Huang Lai, Xiaohua Xie |
IEEE Trans. Image Process. | 1 |
| 2020 | Open-World Group Retrieval with Ambiguity Removal: A BenchmarkabstractGroup retrieval has attracted plenty of attention in artificial intelligence, traditional group retrieval researches assume that members in a group are unique and do not change under different cameras. However, the assumption may not be met for practical situations such as open-world and group-ambiguity scenarios. This paper tackles an important yet non-studied problem: re-identifying changing groups of people under the open-world and group-ambiguity scenarios in different camera fields. The open-world scenario considers that there are probably non-target people for the probe set appear in the searching gallery, while the group-ambiguity scenario means the group members may change. The open-world and group-ambiguity issue is very challenging for the existing methods because the changing of group members results in dramatic visual variations. Nevertheless, as far as we know, the existing literature lacks benchmarks which target on coping with this issue. In this paper, we propose a new group retrieval dataset named OWGA-Campus to consider these challenges. Moreover, we propose a person-to-group similarity matching based ambiguity removal (P2GSM-AR) method to solve these problems and realize the intention of group retrieval. Experimental results on OWGA-Campus dataset demonstrate the effectiveness and robustness of the proposed P2GSM-AR approach in improving the performance of the state-of-the-art feature extraction methods of person re-id towards the open-world and ambiguous group retrieval task. Ling Mei 0001, Jian-Huang Lai, Zhan-Xiang Feng, Xiaohua Xie |
ICPR | 3 |
| 2020 | Deep Face Recognition Based on Penalty Cosface
Shuoyan Lin, Jianxiong Tang, Zhan-Xiang Feng, Jian-Huang Lai |
PRCV (2) | 3 |
| 2020 | From pedestrian to group retrieval via siamese network and correlation
Ling Mei 0001, Jian-Huang Lai, Zhan-Xiang Feng, Xiaohua Xie |
Neurocomputing | 3 |
| 2020 | Learning Modality-Specific Representations for Visible-Infrared Person Re-IdentificationabstractTraditional person re-identification (re-id) methods perform poorly under changing illuminations. This situation can be addressed by using dual-cameras that capture visible images in a bright environment and infrared images in a dark environment. Yet, this scheme needs to solve the visible-infrared matching issue, which is largely under-studied. Matching pedestrians across heterogeneous modalities is extremely challenging because of different visual characteristics. In this paper, we propose a novel framework that employ modality-specific networks to tackle with the heterogeneous matching problem. The proposed framework utilizes the modality-related information and extracts modality-specific representations (MSR) by constructing an individual network for each modality. In addition, a cross-modality Euclidean constraint is introduced to narrow the gap between different networks. We also integrate the modality-shared layers into modality-specific networks to extract shareable information and use a modality-shared identity loss to facilitate the extraction of modality-invariant features. Then a modality-specific discriminant metric is learned for each domain to strengthen the discriminative power of MSR. Eventually, we use a view classifier to learn view information. The experiments demonstrate that the MSR effectively improves the performance of deep networks on VI-REID and remarkably outperforms the state-of-the-art methods. Zhan-Xiang Feng, Jian-Huang Lai, Xiaohua Xie |
IEEE Trans. Image Process. | 1 |
| 2019 | Low Resolution Person Re-identification by an Adaptive Dual-Branch Network
Zhan-Xiang Feng, Jian-Huang Lai, Xiaohua Xie |
ICIG (1) | 1 |
| 2018 | Conditional Face Synthesis for Data Augmentation
Xiaohua Xie, Jian-Huang Lai, Zhan-Xiang Feng |
PRCV (3) | 4 |
| 2018 | Image super-resolution via a densely connected recursive network
Zhan-Xiang Feng, Jian-Huang Lai, Xiaohua Xie, Jun-Yong Zhu |
Neurocomputing | 1 |
| 2018 | Learning View-Specific Deep Networks for Person Re-IdentificationabstractIn recent years, a growing body of research has focused on the problem of person re-identification (re-id). The re-id techniques attempt to match the images of pedestrians from disjoint non-overlapping camera views. A major challenge of the re-id is the serious intra-class variations caused by changing viewpoints. To overcome this challenge, we propose a deep neural network-based framework which utilizes the view information in the feature extraction stage. The proposed framework learns a view-specific network for each camera view with a cross-view Euclidean constraint (CV-EC) and a cross-view center loss. We utilize the CV-EC to decrease the margin of the features between diverse views and extend the center loss metric to a view-specific version to better adapt the re-id problem. Moreover, we propose an iterative algorithm to optimize the parameters of the view-specific networks from coarse to fine. The experiments demonstrate that our approach significantly improves the performance of the existing deep networks and outperforms the state-of-the-art methods on the VIPeR, CUHK01, CUHK03, SYSU-mReId, and Market-1501 benchmarks. Zhan-Xiang Feng, Jian-Huang Lai, Xiaohua Xie |
IEEE Trans. Image Process. | 1 |
| 2017 | Face recognition by landmark pooling-based CNN with concentrate lossabstractFace recognition has been a hot research topic in recent years, convolutional neural network (CNN) based methods have achieved state of the art results and significantly improve the performance. Along with the CNN framework, we propose a novel loss function called concentrate loss which focuses on the class centers in the mini-batch. The concentrate loss aims to push the samples towards corresponding class centers and simultaneously enlarge the gap between different class centers. Additionally, we ultilize facial landmark pooling technique to take full advantage of facial structure information. Experiment results on Labeled Faces in the Wild (LFW), YouTube Faces (YTF), and the BluFR benchmark demonstrate the efficiency of our proposal. Xiaohua Xie, Zhan-Xiang Feng, Jian-Huang Lai |
ICIP | 3 |
| 2016 | Face hallucination by deep traversal networkabstractIn this paper, we propose a novel patch-based face hallucination method that consists of two patch-based sparse autoencoder (SAE) networks and a deep fully connected network (namely traversal network). The SAE networks are used to capture the intrinsic features of low-resolution (LR) images and high-resolution (HR) images in the hidden layers, while the traversal network is used to map features from the LR hidden layer to the HR hidden layer. In the training stage, these three networks are jointly optimized. Compared with previous network-based methods that learn an end-to-end mapping from LR images to HR images, our method learns the mapping between hidden layers, which can better alleviate the over-fitting problem. Experimental results demonstrate that our method is efficient and robust for hallucinating face images from both lab environment and the wild. The proposal achieves state-of-the-art performance when conducting face hallucination in CAS-PEAL-R1 database, CMU-PIE database and Casia database. Zhan-Xiang Feng, Jian-Huang Lai, Xiaohua Xie, Dakun Yang, Ling Mei 0001 |
ICPR | 1 |