VLDB 2026 Research / reviewers in the wild / expert
Zhen-Xiang Ma
dblp:371/4730
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2025
0009-0002-4697-5033ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Few-Shot Fine-Grained Image Classification with Progressively Feature Refinement and Continuous Relationship ModelingabstractRecently, a number of effective methods have been proposed to tackle the challenging task of Few-Shot Fine-Grained Image Classification (FS-FGIC). However, how to fully leverage the backbone network to discover and extract detailed features to generate more discriminative class prototypes, as well as how to accurately model the similarity relationship between query samples and the class prototypes, are still issues to be further considered. Therefore, we propose a novel progreSsively featUre refInement and conTinuous rElationship moDeling method, SUITED for short, to address these two issues existing in the State-of-the-Art FS-FGIC methods. Specifically, we design the Progressive Feature Refinement Module (PFRM) to fully exploit the backbone network's progressive feature extraction capabilities, forming multi-scale feature representations to further enhance discriminative features. Then, the Continuous Relationship Modeling Module (CRMM) is proposed to capture the dependencies between query samples and the corresponding class prototypes, achieving precise optimization of the distances among corresponding sample points in the feature space. We conducted extensive experiments on five fine-grained benchmark datasets, and the experimental results demonstrate that the proposed method is comprehensively ahead of the existing State-of-the-Art methods. Zhen-Xiang Ma, Zhen-Duo Chen 0001, Tai Zheng, Xin Luo 0006, Zixia Jia, Xin-Shun Xu |
AAAI | 1 |
| 2025 | DGPrompt: Dual-guidance prompts generation for vision-language models
Tai Zheng, Zhen-Duo Chen 0001, Zi-Chao Zhang 0002, Zhen-Xiang Ma, Li-Jun Zhao 0005, Chong-Yu Zhang, Xin Luo 0006, Xin-Shun Xu |
Neural Networks | 4 |
| 2025 | BTG-Net++: Enhanced Bi-Directional Task-Guided Network for Few-Shot Fine-Grained Image ClassificationabstractIn recent years, a number of effective Few-Shot Fine-Grained Image Classification (FS-FGIC) methods have been proposed, which mainly focus on extracting discriminative information within high-level features in a single episode/task. However, this is insufficient for addressing the cross-task challenges of FS-FGIC, which is represented in two aspects. On the one hand, from the perspective of the Fine-Grained Image Classification (FGIC) task, there is a need to supplement the model with mid-level features containing rich fine-grained information. On the other hand, from the perspective of the Few-Shot Learning (FSL) task, explicit modeling of cross-task general knowledge is required. In this paper, we propose a novel Enhanced Bi-directional Task-Guided Network (BTG-Net++) to tackle these issues. Specifically, from the FGIC task perspective, we design the Semantic-Guided Noise Filtering (SGNF) module to filter noise on mid-level features rich in detailed information with the assistance of high-level features. Further, from the FSL task perspective, the General Knowledge Prompt Modeling (GKPM) module is proposed to retain the cross-task general knowledge by utilizing the prompting mechanism, thereby enhancing the model’s generalization performance on unseen novel classes. We have conducted extensive experiments on five fine-grained benchmark datasets, and the results demonstrate that BTG-Net++ shows considerable improvements compared with state-of-the-art methods. Zhen-Xiang Ma, Zhen-Duo Chen 0001, Tai Zheng, Xin Luo 0006, Xin-Shun Xu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Distributed Learning for Privacy-Preserving Semi-Supervised Video Anomaly DetectionabstractSemi-supervised video anomaly detection (SS-VAD) is essential for intelligent monitoring. However, collecting large-scale surveillance videos from various organizations raises significant privacy concerns regarding sensitive information. Federated learning offers a promising solution by enabling distributed learning among multiple participants while safeguarding privacy. Despite its potential, research on applying federated learning to SS-VAD remains unexplored due to the inherent challenges of this task. In this paper, we solve this task via proposing DLPP, a novel distributed learning framework for privacy-preserving SS-VAD. It addresses the issue of statistical heterogeneity among data from different participants in real-world federated SS-VAD applications, particularly focusing on non-independent and identically distributed (non-IID) data and imbalanced data volumes. In specific, it addresses these challenges in two key innovations: 1) For the non-IID data challenge, it dynamically updates the client model based on the overall gradient at the client of the previous training round and the degree of divergence between the server model and the client model. In this way, it can better adapt the server model to each client and promote convergence. 2) For the imbalanced data volumes challenge, it adaptively allocates client aggregation weights by comprehensively considering the data volumes, model quality, and learning efficiency of clients. This means a more robust server model can be obtained, and model bias reduced. We conduct extensive experiments to evaluate the performance of DLPP on benchmark datasets by partitioning data to simulate various degrees of non-IID environments. The results show that DLPP significantly outperforms both Baseline and SOTA methods, achieving up to a 3.89% improvement, and its communication efficiency is 3x better than FedAvg. Xiao-Dong Xie, Yu-Wei Zhan, Zhen-Xiang Ma, Hong-Mei Liu, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Cross-Layer and Cross-Sample Feature Optimization Network for Few-Shot Fine-Grained Image ClassificationabstractRecently, a number of Few-Shot Fine-Grained Image Classification (FS-FGIC) methods have been proposed, but they primarily focus on better fine-grained feature extraction while overlooking two important issues. The first one is how to extract discriminative features for Fine-Grained Image Classification tasks while reducing trivial and non-generalizable sample level noise introduced in this procedure, to overcome the over-fitting problem under the setting of Few-Shot Learning. The second one is how to achieve satisfying feature matching between limited support and query samples with variable spatial positions and angles. To address these issues, we propose a novel Cross-layer and Cross-sample feature optimization Network for FS-FGIC, C2-Net for short. The proposed method consists of two main modules: Cross-Layer Feature Refinement (CLFR) module and Cross-Sample Feature Adjustment (CSFA) module. The CLFR module further refines the extracted features while integrating outputs from multiple layers to suppress sample-level feature noise interference. Additionally, the CSFA module addresses the feature mismatch between query and support samples through both channel activation and position matching operations. Extensive experiments have been conducted on five fine-grained benchmark datasets, and the results show that the C2-Net outperforms other state-of-the-art methods by a significant margin in most cases. Our code is available at: https://github.com/zenith0923/C2-Net. Zhen-Xiang Ma, Zhen-Duo Chen 0001, Li-Jun Zhao 0005, Zi-Chao Zhang 0002, Xin Luo 0006, Xin-Shun Xu |
AAAI | 1 |
| 2024 | Bi-directional Task-Guided Network for Few-Shot Fine-Grained Image ClassificationabstractIn recent years, the Few-Shot Fine-Grained Image Classification (FS-FGIC) problem has gained widespread attention. A number of effective methods have been proposed that focus on extracting discriminative information within high-level features in a single episode/task. However, this is insufficient for addressing the cross-task challenges of FS-FGIC, which is represented in two aspects. On the one hand, from the perspective of the Fine-Grained Image Classification (FGIC) task, there is a need to supplement the model with mid-level features containing rich fine-grained information. On the other hand, from the perspective of the Few-Shot Learning (FSL) task, explicit modeling of cross-task general knowledge is required. In this paper, we propose a novel Bi-directional Task-Guided Network (BTG-Net) to tackle these issues. Specifically, from the FGIC task perspective, we design the Semantic-Guided Noise Filtering (SGNF) module to filter noise on mid-level features rich in detailed information. Further, from the FSL task perspective, the General Knowledge Prompt Modeling (GKPM) module is proposed to retain the cross-task general knowledge by utilizing the prompting mechanism, thereby enhancing the model's generalization performance on novel classes. We have conducted extensive experiments on five fine-grained benchmark datasets, and the results demonstrate that BTG-Net outperforms state-of-the-art methods comprehensively. Zhen-Xiang Ma, Zhen-Duo Chen 0001, Li-Jun Zhao 0005, Zi-Chao Zhang 0002, Tai Zheng, Xin Luo 0006, Xin-Shun Xu |
ACM Multimedia | 1 |
| 2024 | Angular Isotonic Loss Guided Multi-Layer Integration for Few-Shot Fine-Grained Image ClassificationabstractRecent research on few-shot fine-grained image classification (FSFG) has predominantly focused on extracting discriminative features. The limited attention paid to the role of loss functions has resulted in weaker preservation of similarity relationships between query and support instances, thereby potentially limiting the performance of FSFG. In this regard, we analyze the limitations of widely adopted cross-entropy loss and introduce a novel Angular ISotonic (AIS) loss. The AIS loss introduces an angular margin to constrain the prototypes to maintain a certain distance from a pre-set threshold. It guides the model to converge more stably, learn clearer boundaries among highly similar classes, and achieve higher accuracy faster with limited instances. Moreover, to better accommodate the feature requirements of the AIS loss and fully exploit its potential in FSFG, we propose a Multi-Layer Integration (MLI) network that captures object features from multiple perspectives to provide more comprehensive and informative representations of the input images. Extensive experiments demonstrate the effectiveness of our proposed method on four standard fine-grained benchmarks. Codes are available at: https://github.com/Legenddddd/AIS-MLI. Li-Jun Zhao 0005, Zhen-Duo Chen 0001, Zhen-Xiang Ma, Xin Luo 0006, Xin-Shun Xu |
IEEE Trans. Image Process. | 3 |