Zizhu Fan

dblp:88/1897 · DBLP profile ↗
← Back
39ranked-venue papers
14as first author
17since 2021 · last 2026
0000-0001-5354-4827ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 11 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SimpleDiffusion: A Lightweight and Efficient Conditional Diffusion Model for Multi-Modal Salient Object Detection
abstract
Multi-modal salient object detection (MSOD), which integrates complementary modalities such as depth or thermal data, primarily faces two challenges: accurately preserving salient object details and effectively aligning cross-modal features. Recent advances in using Stable Diffusion to generate images with fine edge details have inspired researchers to reformulate MSOD as a conditional mask generation process guided by salient features, which has achieved excellent visual results. However, these approaches often overlook the high computational cost and large-scale architecture of Stable Diffusion, both of which render it unsuitable for real-world MSOD applications. Therefore, we propose SimpleDiffusion, the first lightweight and efficient conditional diffusion model for MSOD that does not rely on Stable Diffusion. Specifically, we propose an Adaptive Cross-Modal Fusion Conditional Network and a Latent Denoising Network to reduce the complexity of diffusion models. Furthermore, we design a Multi-modal Feature Rectification and Fusion Module to enhance the representational capacity of cross-modal salient features. Customized training and sampling strategies are also developed to improve inference efficiency and reduce erroneous object segmentations. Experiments on multiple MSOD datasets demonstrate that SimpleDiffusion reduces model size by over tenfold and improves inference speed by more than fivefold compared to other diffusion-based methods, while maintaining comparable or superior performance.
Shuo Zhang 0013, Wenbing Tang 0001, Jing Liu 0012, Li Han 0001, Jiandun Li, Hongchun Yuan, Zizhu Fan
AAAI8
2025 HCETrack: Visual Tracking with Historical Context Information and Feature Enhancement
abstract
Currently, most deep trackers follow the Siamese paradigm, where target tracking is achieved through similarity matching between the template and search region features. This approach often processes each image pair independently, making it difficult to establish sufficient temporal correlation. This leads to poor performance of trackers when dealing with scenes where target appearance changes drastically, such as partial occlusion, rapid motion, scale variations, and target deformation. To address the above issues, this paper proposes a Transformer-based tracker that can utilize accurate and updated historical context information. Specifically, by comprehensively fusing information from the template frame, historical context information(including the target’s location and surrounding state information), and search frame, our method is able to more accurately locate the target and significantly enhance the tracker’s robustness and regression accuracy. Furthermore, this paper introduces a feature enhancement network to strengthen the extracted target features of the search region, thereby suppressing the impact of background and noise on tracking performance. Based on the above design, the proposed method is referred to as HCETrack, which effectively leverages historical context information and a feature enhancement network. Our method can achieve remarkable performance improvements on the LaSOT, LaSOText, and GOT-10k datasets.
Zizhu Fan
SMC3
2025 Mixed Modality Generation and Hierarchical Feature Aggregation for Visible-Infrared Person Re-Identification
abstract
ABSTRACT The main challenge in visible‐infrared person re‐identification (VI‐ReID), which involves matching images of target persons across different modalities, is the significant modality gap between visible and infrared images. Current approaches generally fall into two categories: network architectures that map features from different modalities into a shared feature space, and methods that focus on modality generation and transformation. However, these methods often fail to fully account for contextual relationships, and the generated modalities may lack semantic interpretability. To address these limitations, this paper proposes a mixed modality generator that aligns the visible and infrared modalities as closely as possible within a unified feature space. To effectively leverage multi‐layered information, we introduce a hierarchical feature aggregation module, which establishes connections between features across different layers. Comprehensive experiments on the SYSU‐MM01, RegDB, and LLCM datasets demonstrate that our method significantly outperforms existing state‐of‐the‐art approaches, setting a new benchmark for VI‐ReID performance.
Keming Wei, Zizhu Fan
IET Image Process.5
2025 Low-rank kernel consistent multi-view subspace clustering
Wei Zhang 0079, Shiqi Wang 0001, Zizhu Fan
Neurocomputing7
2025 Contrastive Graph Semantic Learning via prototype for recommendation
Mi Wen, Weiwei Li 0007, Zizhu Fan, Xiaoqing Yu
Inf. Sci.4
2025 A novel heterogeneous data classification approach combining gradient boosting decision trees and hybrid structure model
Yuting Huang 0001, Zizhu Fan
Pattern Recognit.4
2024 Real-Time Text Detection with Multi-level Feature Fusion and Pixel Clustering
Zhufeng Jiang, Xingyu Han, Zizhu Fan
PRCV (7)5
2024 Semi-supervised multiview fuzzy broad learning
Chao Xi, Zizhu Fan, Cheng Peng 0016
Inf. Sci.2
2024 Efficient license plate recognition in unconstrained scenarios
Zizhu Fan, Linrui Shi
J. Vis. Commun. Image Represent.3
2024 Semi-supervised fuzzy broad learning system based on mean-teacher model
Zizhu Fan, Chao Xi
Pattern Anal. Appl.1
2024 Pedestrian Intrusion Detection in Railway Station Based on Mirror Translation Attention and Feature Pooling Enhancement
abstract
Pedestrian intrusion detection is crucial to ensuring safe railway operation. Current pedestrian detection algorithms lack consideration for real-world railway scenarios, such as the reflective properties of screen doors and train windows, may mistakenly trigger pedestrian intrusion alerts. Scale variability and pedestrian overlap often lead to detection inaccuracy, making them inadequate for addressing the specific requirements of railway perimeter security. This letter introduces an innovative pedestrian detection algorithm that incorporates Mirror Translation Attention (MTA) and Feature Pooling Enhancement (FPE). MTA, including mirror flipping and offsetting the feature mapping, could significantly mitigate missed detection caused by reflective surfaces. Additionally, we introduce sparsity to the inputs of the self-attention, which significantly enhancing the model's inference speed. A multi-scale approach is adopted to accommodate the diversity in pedestrian sizes, while the FPE addresses occlusion issues across various scales. Compared to the advanced YOLOv8 model, the proposed method improves AP50 by 1.6% to 92.11% and reduces model parameters by 63.55% in our self-built railway pedestrian intrusion dataset.
Zhufeng Jiang, Guoliang Luo, Zizhu Fan
IEEE Signal Process. Lett.4
2023 Kernel Fisher Dictionary Transfer Learning
abstract
Dictionary learning is an efficient knowledge representation method that can learn the essential features of data. Traditional dictionary learning methods are difficult to obtain nonlinear information when processing large-scale and high-dimensional datasets. While most dictionary learning algorithms are based on the assumption that the training data and test data have the same feature distribution, which is not always true in practical applications. To address the above problems, we propose the Kernel Fisher Dictionary Transfer Learning (KFDTL) algorithm. First, we map each sample to high-dimensional space through kernel mapping and use any dictionary learning algorithm to learn the essential features. Then, the feature-based transfer learning method is performed to predict the labels of the target samples. This method includes three main contributions: (1) KFDTL constructs a discriminative Fisher embedding model to make the same class samples have similar coding coefficients; (2) Based on the relationship between profiles and atoms, KFDTL constructs an adaptive model that adapts source domain samples to target domain samples; (3) The kernel method is used to efficiently solve nonlinear problems. Experiments on a large number of public image datasets have proved the effectiveness of the proposed method. The source code of the proposed method is available at https://github.com/zzfan3/KFDTL .
Linrui Shi, Zheng Zhang 0006, Zizhu Fan, Chao Xi, Gaochang Wu
ACM Trans. Knowl. Discov. Data3
2023 Discriminative Fisher Embedding Dictionary Transfer Learning for Object Recognition
abstract
In transfer learning model, the source domain samples and target domain samples usually share the same class labels but have different distributions. In general, the existing transfer learning algorithms ignore the interclass differences and intraclass similarities across domains. To address these problems, this article proposes a transfer learning algorithm based on discriminative Fisher embedding and adaptive maximum mean discrepancy (AMMD) constraints, called discriminative Fisher embedding dictionary transfer learning (DFEDTL). First, combining the label information of source domain and part of target domain, we construct the discriminative Fisher embedding model to preserve the interclass differences and intraclass similarities of training samples in transfer learning. Second, an AMMD model is constructed using atoms and profiles, which can adaptively minimize the distribution differences between source domain and target domain. The proposed method has three advantages: 1) using the Fisher criterion, we construct the discriminative Fisher embedding model between source domain samples and target domain samples, which encourages the samples from the same class to have similar coding coefficients; 2) instead of using the training samples to design the maximum mean discrepancy (MMD), we construct the AMMD model based on the relationship between the dictionary atoms and profiles; thus, the source domain samples can be adaptive to the target domain samples; and 3) the dictionary learning is based on the combination of source and target samples which can avoid the classification error caused by the difference among samples and reduce the tedious and expensive data annotation. A large number of experiments on five public image classification datasets show that the proposed method obtains better classification performance than some state-of-the-art dictionary and transfer learning methods. The code has been available at https://github.com/shilinrui/DFEDTL.
Zizhu Fan, Linrui Shi, Qiang Liu 0018, Zheng Zhang 0006
IEEE Trans. Neural Networks Learn. Syst.1
2022 A survey of crowd counting and density estimation based on convolutional neural network
Zizhu Fan, Zheng Zhang 0006, Guangming Lu 0002, Yudong Zhang 0001, Yaowei Wang 0001
Neurocomputing1
2022 Incomplete multi-modal brain image fusion for epilepsy classification
Qi Zhu 0001, Huijie Li, Haizhou Ye, Ran Wang 0004, Zizhu Fan, Daoqiang Zhang
Inf. Sci.6
2022 NMFLRR: Clustering scRNA-Seq Data by Integrating Nonnegative Matrix Factorization With Low Rank Representation
abstract
Fast-developing single-cell technologies create unprecedented opportunities to reveal cell heterogeneity and diversity. Accurate classification of single cells is a critical prerequisite for recovering the mechanisms of heterogeneity. However, the scRNA-seq profiles we obtained at present have high dimensionality, sparsity, and noise, which pose challenges for existing clustering methods in grouping cells that belong to the same subpopulation based on transcriptomic profiles. Although many computational methods have been proposed developing novel and effective computational methods to accurately identify cell types remains a considerable challenge. We present a new computational framework to identify cell types by integrating low-rank representation (LRR) and nonnegative matrix factorization (NMF); this framework is named NMFLRR. The LRR captures the global properties of original data by using nuclear norms, and a locality constrained graph regularization term is introduced to characterize the data's local geometric information. The similarity matrix and low-dimensional features of data can be simultaneously obtained by applying the alternating direction method of multipliers (ADMM) algorithm to handle each variable alternatively in an iterative way. We finally obtained the predicted cell types by using a spectral algorithm based on the optimized similarity matrix. Nine real scRNA-seq datasets were used to test the performance of NMFLRR and fifteen other competitive methods, and the accuracy and robustness of the simulation results suggest the NMFLRR is a promising algorithm for the classification of single cells. The simulation code is freely available at: https://github.com/wzhangwhu/NMFLRR_code.
Wei Zhang 0079, Xiaoli Xue, Zizhu Fan
IEEE J. Biomed. Health Informatics4
2021 Improving decomposition-based multiobjective evolutionary algorithm with local reference point aided search
Jing Jiang 0021, Fei Han 0001, Jie Wang 0050, Henry Han, Zizhu Fan
Inf. Sci.6
2020 Fast kernel sparse representation based classification for Undersampling problem in face recognition
Zizhu Fan
Multim. Tools Appl.1
2018 An interactively constrained discriminative dictionary learning algorithm for image classification
Zheng Zhang 0006, Zizhu Fan, Jie Wen 0001
Eng. Appl. Artif. Intell.3
2018 Virtual dictionary based kernel sparse representation for face recognition
Zizhu Fan, Da Zhang 0001, Xin Wang 0061, Qi Zhu 0001, Yuan-Fang Wang
Pattern Recognit.1
2016 An efficient sparse representation based classification for undersampled face recognition
abstract
The typical sparse representation for classification (SRC) can obtain desirable recognition result when the training samples in each class are sufficient. Nevertheless, if the training sample set is small scale, i.e., each class has a few training samples, even single sample, the traditional SRC cannot perform well. Although one of the variants of the traditional SRC, the extended SRC(ESRC), can effectively address the above small-scale training set (SSTS) problem, its computational efficiency is very low and consequently constrains the application of the ESRC algorithm. In order to improve the computational efficiency of the ESRC algorithm, we propose a new algorithm based on coordinate descent scheme in this work. Our proposed algorithm is referred as to the fast extended SRC (FESRC) algorithm. Experiments on popular face datasets show that the FESRC algorithm can obtain the high computational efficiency without significantly degrading the recognition results.
Zizhu Fan, Lipan Kang
CEC1
2016 Individualized learning for improving kernel Fisher discriminant analysis
Zizhu Fan, Yong Xu 0001, Xiaozhao Fang, David Zhang 0001
Pattern Recognit.1
2015 Weighted sparse representation for face recognition
Zizhu Fan, Qi Zhu 0001, Ergen Liu
Neurocomputing1
2015 Manifold discriminant regression learning for image classification
Yuwu Lu, Zhihui Lai 0001, Zizhu Fan, Jinrong Cui, Qi Zhu 0001
Neurocomputing3
2015 L0-norm sparse representation based on modified genetic algorithm for face recognition
Zizhu Fan, Qi Zhu 0001, Chengli Sun, Lipan Kang
J. Vis. Commun. Image Represent.1
2015 Appearance-based bidirectional representation for palmprint recognition
Jinrong Cui, Jiajun Wen 0001, Zizhu Fan
Multim. Tools Appl.3
2015 Erratum to: Appearance-based bidirectional representation for palmprint recognition
Jinrong Cui, Jiajun Wen 0001, Zizhu Fan
Multim. Tools Appl.3
2014 Kernel sparse representation based classification for undersampled problem
abstract
Sparse representation for classification (SRC) has attracted much attention in recent years. It usually performs well under the following assumptions. The first assumption is that each class has sufficient training samples. In other words, SRC is not good at dealing with the undersampled problem, i.e., each class has few training samples, even single sample. The second one is that the sample vectors belonging to different classes should not distribute on the same vector direction. However, the above two assumptions are not always satisfied in real-world problems. In this paper, we propose a novel SRC based algorithm, i.e., kernel sparse representation based classifier for undersampled problem (KSRC-UP) to perform classification. It does not need the above assumptions in principle. KSRC-UP can deal well with the small scale and high dimensional real world data sets. Experiments on the popular face databases show that our KSRC-UP method can perform better than other SRC methods.
Zizhu Fan, Qi Zhu 0001, Yuwu Lu
SMARTCOMP1
2014 Locality and similarity preserving embedding for feature selection
Xiaozhao Fang, Yong Xu 0001, Xuelong Li 0001, Zizhu Fan, Hong Liu 0008, Yan Chen 0018
Neurocomputing4
2014 An efficient KPCA algorithm based on feature correlation evaluation
Zizhu Fan, Baogen Xu, Pengzhi Tang
Neural Comput. Appl.1
2014 Modified Principal Component Analysis: An Integration of Multiple Similarity Subspace Models
abstract
We modify the conventional principal component analysis (PCA) and propose a novel subspace learning framework, modified PCA (MPCA), using multiple similarity measurements. MPCA computes three similarity matrices exploiting the similarity measurements: 1) mutual information; 2) angle information; and 3) Gaussian kernel similarity. We employ the eigenvectors of similarity matrices to produce new subspaces, referred to as similarity subspaces. A new integrated similarity subspace is then generated using a novel feature selection approach. This approach needs to construct a kind of vector set, termed weak machine cell (WMC), which contains an appropriate number of the eigenvectors spanning the similarity subspaces. Combining the wrapper method and the forward selection scheme, MPCA selects a WMC at a time that has a powerful discriminative capability to classify samples. MPCA is very suitable for the application scenarios in which the number of the training samples is less than the data dimensionality. MPCA outperforms the other state-of-the-art PCA-based methods in terms of both classification accuracy and clustering result. In addition, MPCA can be applied to face image reconstruction. MPCA can use other types of similarity measurements. Extensive experiments on many popular real-world data sets, such as face databases, show that MPCA achieves desirable classification results, as well as has a powerful capability to represent data.
Zizhu Fan, Yong Xu 0001, Wangmeng Zuo, Jian Yang 0003, Jinhui Tang 0001, Zhihui Lai 0001, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2013 A sparse representation method of bimodal biometrics and palmprint recognition experiments
Yong Xu 0001, Zizhu Fan, Minna Qiu, David Zhang 0001, Jing-Yu Yang 0001
Neurocomputing2
2013 From the idea of "sparse representation" to a representation-based transformation method for feature extraction
Yong Xu 0001, Qi Zhu 0001, Zizhu Fan, Yaowu Wang, Jeng-Shyang Pan 0001
Neurocomputing3
2013 Using the idea of the sparse representation to perform coarse-to-fine face recognition
Yong Xu 0001, Qi Zhu 0001, Zizhu Fan, David Zhang 0001, Jian-Xun Mi, Zhihui Lai 0001
Inf. Sci.3
2013 Coarse to fine K nearest neighbor classifier
Yong Xu 0001, Qi Zhu 0001, Zizhu Fan, Minna Qiu, Yan Chen 0018, Hong Liu 0008
Pattern Recognit. Lett.3
2012 Kernel based sparse representation for face recognition
Qi Zhu 0001, Yong Xu 0001, Zizhu Fan
ICPR4
2012 Supervised sparse representation method with a heuristic strategy and face recognition experiments
Yong Xu 0001, Wangmeng Zuo, Zizhu Fan
Neurocomputing3
2011 Local Linear Discriminant Analysis Framework Using Sample Neighbors
abstract
The linear discriminant analysis (LDA) is a very popular linear feature extraction approach. The algorithms of LDA usually perform well under the following two assumptions. The first assumption is that the global data structure is consistent with the local data structure. The second assumption is that the input data classes are Gaussian distributions. However, in real-world applications, these assumptions are not always satisfied. In this paper, we propose an improved LDA framework, the local LDA (LLDA), which can perform well without needing to satisfy the above two assumptions. Our LLDA framework can effectively capture the local structure of samples. According to different types of local data structure, our LLDA framework incorporates several different forms of linear feature extraction approaches, such as the classical LDA and principal component analysis. The proposed framework includes two LLDA algorithms: a vector-based LLDA algorithm and a matrix-based LLDA (MLLDA) algorithm. MLLDA is directly applicable to image recognition, such as face recognition. Our algorithms need to train only a small portion of the whole training set before testing a sample. They are suitable for learning large-scale databases especially when the input data dimensions are very high and can achieve high classification accuracy. Extensive experiments show that the proposed algorithms can obtain good classification results.
Zizhu Fan, Yong Xu 0001, David Zhang 0001
IEEE Trans. Neural Networks1
2008 A New Unsupervised Approach to Face Recognition
Zizhu Fan, Ergen Liu
ICIC (1)1