Fanglin Chen 0001

dblp:85/7057-1 · DBLP profile ↗
← Back
38ranked-venue papers
7as first author
27since 2021 · last 2025
0000-0002-9193-5412ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 3 first-author · 16 since 2021Artificial intelligence and machine learning · 19 · 3 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 ParseCaps: An Interpretable Parsing Capsule Network for Medical Image Diagnosis
abstract
Deep learning has excelled in medical image classification, but its clinical application is limited by poor interpretability. Capsule networks, known for encoding hierarchical relationships and spatial features, show potential in addressing this issue. Nevertheless, traditional capsule networks often underperform due to their shallow structures, and deeper variants lack hierarchical architectures, thereby compromising interpretability. This paper introduces a novel capsule network, ParseCaps, which utilizes the sparse axial attention routing and parse convolutional capsule layer to form a parse-tree-like structure, enhancing both depth and interpretability. Firstly, sparse axial attention routing optimizes connections between child and parent capsules, as well as emphasizes the weight distribution across instantiation parameters of parent capsules. Secondly, the parse convolutional capsule layer generates capsule predictions aligning with the parse tree. Finally, based on the loss design that is effective whether concept ground truth exists or not, ParseCaps advances interpretability by associating each dimension of the global capsule with a comprehensible concept, thereby facilitating clinician trust and understanding of the model's classification results. Experimental results on three medical datasets show that ParseCaps not only outperforms other capsule network variants in classification accuracy and robustness, but also provides interpretable explanations, regardless of the availability of concept labels.
Xinyu Geng, Xiaolin Huang, Fanglin Chen 0001, Jun Xu 0008
AAAI4
2025 Dynamic VAEs via semantic-aligned matching for continual zero-shot learning
Junbo Yang, Borui Hu, Yang Liu 0069, Xinbo Gao 0001, Jungong Han, Fanglin Chen 0001, Xuangou Wu
Pattern Recognit.7
2025 Interpretable Multi-Agent Reinforcement Learning for Traffic Signal Control: Influence Mechanism and Piecewise Linear Approximation
abstract
Traffic signal control plays a crucial role in intelligent transportation systems, with cooperative control being challenging to implement but essential for its effectiveness. Many methods model multi-intersection traffic networks as grids and address the problem using multi-agent reinforcement learning (RL). Despite these existing studies, there is an opportunity to further enhance our understanding of the connectivity and globality of the traffic networks by capturing the spatiotemporal traffic information with efficient neural networks in deep RL. In this paper, we propose a novel multi-agent actor-critic framework based on an interpretable influence mechanism with a centralized learning and decentralized execution method. Specifically, we first construct an actor-critic framework, for which the piecewise linear neural network (PWLNN), named biased ReLU (BReLU), is used as the function approximator to obtain a more accurate and theoretically grounded approximation, and exhibits interpretability. Then, to model the relationships among agents in multi-intersection scenarios, we introduce an interpretable influence mechanism based on efficient hinging hyperplanes neural network (EHHNN), which derives weights by analysis of variance (ANOVA) decomposition among agents and extracts spatiotemporal dependencies of the traffic features. Finally, our proposed framework is validated on two synthetic traffic networks and a real road network to coordinate signal control between intersections, achieving lower traffic delays across the entire traffic network compared with benchmark performance.
Zhiyue Luo, Jun Xu 0008, Fanglin Chen 0001
IEEE Trans Autom. Sci. Eng.4
2024 SA²VP: Spatially Aligned-and-Adapted Visual Prompt
abstract
As a prominent parameter-efficient fine-tuning technique in NLP, prompt tuning is being explored its potential in computer vision. Typical methods for visual prompt tuning follow the sequential modeling paradigm stemming from NLP, which represents an input image as a flattened sequence of token embeddings and then learns a set of unordered parameterized tokens prefixed to the sequence representation as the visual prompts for task adaptation of large vision models. While such sequential modeling paradigm of visual prompt has shown great promise, there are two potential limitations. First, the learned visual prompts cannot model the underlying spatial relations in the input image, which is crucial for image encoding. Second, since all prompt tokens play the same role of prompting for all image tokens without distinction, it lacks the fine-grained prompting capability, i.e., individual prompting for different image tokens. In this work, we propose the Spatially Aligned-and-Adapted Visual Prompt model (SA^2VP), which learns a two-dimensional prompt token map with equal (or scaled) size to the image token map, thereby being able to spatially align with the image map. Each prompt token is designated to prompt knowledge only for the spatially corresponding image tokens. As a result, our model can conduct individual prompting for different image tokens in a fine-grained manner. Moreover, benefiting from the capability of preserving the spatial structure by the learned prompt token map, our SA^2VP is able to model the spatial relations in the input image, leading to more effective prompting. Extensive experiments on three challenging benchmarks for image classification demonstrate the superiority of our model over other state-of-the-art methods for visual prompt tuning. Code is available at https://github.com/tommy-xq/SA2VP.
Wenjie Pei, Tongqi Xia, Fanglin Chen 0001, Jiandong Tian, Guangming Lu 0002
AAAI3
2024 OrthCaps: An Orthogonal CapsNet with Sparse Attention Routing and Pruning
abstract
Redundancy is a persistent challenge in Capsule Networks (CapsNet), leading to high computational costs and parameter counts. Although previous studies have introduced pruning after the initial capsule layer, dynamic routing's fully connected nature and non-orthogonal weight matrices reintroduce redundancy in deeper layers. Besides, dynamic routing requires iterating to converge, further increasing computational demands. In this paper, we propose an Orthogonal Capsule Network (OrthCaps) to reduce redundancy, improve routing performance and decrease parameter counts. Firstly, an efficient pruned capsule layer is introduced to discard redundant capsules. Secondly, dynamic routing is replaced with orthogonal sparse attention routing, eliminating the need for iterations and fully connected structures. Lastly, weight matrices during routing are orthogonalized to sustain low capsule similarity, which is the first approach to use Householder orthogonal decomposition to enforce orthogonality in CapsNet. Our experiments on baseline datasets affirm the efficiency and robustness of OrthCaps in classification tasks, in which ablation studies validate the criticality of each component. OrthCaps-Shallow outperforms other Capsule Network benchmarks on four datasets, utilizing only 110k parameters - a mere 1.25% of a standard Capsule Network's total. To the best of our knowledge,$it$achieves the smallest parameter count among existing Capsule Networks. Similarly, OrthCaps-Deep demonstrates competitive performance across four datasets, utilizing only 1.2% of the parameters required by its counterparts.
Xinyu Geng, Jiawei Gong, Yuerong Xue, Jun Xu 0008, Fanglin Chen 0001, Xiaolin Huang
CVPR6
2024 Domain-Rectifying Adapter for Cross-Domain Few-Shot Segmentation
abstract
Few-shot semantic segmentation (FSS) has achieved great success on segmenting objects of novel classes, supported by only a few annotated samples. However, existing FSS methods often underperform in the presence of domain shifts, especially when encountering new domain styles that are unseen during training. It is suboptimal to directly adapt or generalize the entire model to new domains in the few-shot scenario. Instead, our key idea is to adapt a small adapter for rectifying diverse target domain styles to the source domain. Consequently, the rectified target domain features can fittingly benefit from the well-optimized source domain segmentation model, which is intently trained on sufficient source domain data. Training domain-rectifying adapter requires sufficiently diverse target domains. We thus propose a novel local-global style perturbation method to simulate diverse potential target domains by perturbating the feature channel statistics of the individual images and collective statistics of the entire source domain, respectively. Additionally, we propose a cyclic domain alignment module to facilitate the adapter effectively rectifying domains using a reverse domain rectification supervision. The adapter is trained to rectify the image features from diverse synthesized target domains to align with the source domain. During testing on target domains, we start by rectifying the image features and then conduct few-shot segmentation on the domain-rectified features. Extensive experiments demonstrate the effectiveness of our method, achieving promising results on cross-domain few-shot semantic segmentation tasks. Our code is available at https://github.com/Matt-Su/DR-Adapter.
Jiapeng Su, Wenjie Pei, Guangming Lu 0002, Fanglin Chen 0001
CVPR5
2024 WeCromCL: Weakly Supervised Cross-Modality Contrastive Learning for Transcription-Only Supervised Text Spotting
Zhengyao Fang, Pengyuan Lv, Chengquan Zhang, Fanglin Chen 0001, Guangming Lu 0002, Wenjie Pei
ECCV (31)5
2024 Universal Object Detection with Large Vision Model
Feng Lin 0009, Wenze Hu, Yaowei Wang 0001, Yonghong Tian 0001, Guangming Lu 0002, Fanglin Chen 0001, Yong Xu 0007, Xiaoyu Wang 0002
Int. J. Comput. Vis.6
2024 Exploring the complementarity between convolution and transformer matching for visual tracking
Zheng'ao Wang, Ming Li 0028, Wenjie Pei, Guangming Lu 0002, Fanglin Chen 0001
Knowl. Based Syst.5
2024 Robust Tracking via Fully Exploring Background Prior Knowledge
abstract
Typical Siamese-based trackers focus on the target region and pay less attention to the background area. However, the background area can provide the tracker with prior knowledge about the target surroundings. Nonetheless, since the tracker can naturally utilize the target template for localization, importing additional background knowledge requires proper design so that the background area prior knowledge can be fully explored. Furthermore, the introduction of the entire background regions is redundant. Instead, the part background distractors in the regions are more meaningful for the discrimination of the tracker. In this work, we propose a background prior knowledge fully explored tracker for robust tracking. Firstly, we present a Transformer-based explicitly and fully background-utilizing scheme by boosting the tracker to independently exploit the background for localization. Specifically, a target-distractor independent decoder explicitly utilizes the background knowledge by making the target and the distractors independently perform fusion with the search feature. Secondly, we design a simple yet efficient discriminative distractors mining module to refine the background prior knowledge by replacing the whole background region with the mined background distractors. Extensive experiments demonstrate that the proposed method performs favorably against state-of-the-art trackers on nine benchmarks.
Zheng'ao Wang, Zikun Zhou, Fanglin Chen 0001, Jun Xu 0008, Wenjie Pei, Guangming Lu 0002
IEEE Trans. Circuits Syst. Video Technol.3
2023 Deep adaptive hiding network for image hiding using attentive frequency extraction and gradual depth extraction
Le Zhang 0016, Yao Lu 0008, Jinxing Li 0003, Fanglin Chen 0001, Guangming Lu 0002, David Zhang 0001
Neural Comput. Appl.4
2023 Image-Text Retrieval With Cross-Modal Semantic Importance Consistency
abstract
Cross-modal image-text retrieval is an important area of Vision-and-Language task that models the similarity of image-text pairs by embedding features into a shared space for alignment. To bridge the heterogeneous gap between the two modalities, current approaches achieve inter-modal alignment and intra-modal semantic relationship modeling through complex weighted combinations between items. In the intra-modal association and inter-modal interaction processes, the higher-weight items have a higher contribution to the global semantics. However, the same item always produces different contributions in the two processes, since most traditional approaches only focus on the alignment. This usually results in semantic changes and misalignment. To address this issue, this paper proposes Cross-modal Semantic Importance Consistency (CSIC) which achieves invariance in the semantic of items during aligning. The proposed technique measures the semantic importance of items obtained from intra-modal and inter-modal self-attention and learns a more reasonable representation vector by inter-calibrating the importance distribution to improve performance. We conducted extensive experiments on the Flickr30K and MS COCO datasets. The results show that our approach can significantly improve retrieval performance, proving the proposed approach’s superiority and rationality.
Zejun Liu, Fanglin Chen 0001, Jun Xu 0008, Wenjie Pei, Guangming Lu 0002
IEEE Trans. Circuits Syst. Video Technol.2
2023 Pedestrian Detection by Exemplar-Guided Contrastive Learning
abstract
Typical methods for pedestrian detection focus on either tackling mutual occlusions between crowded pedestrians, or dealing with the various scales of pedestrians. Detecting pedestrians with substantial appearance diversities such as different pedestrian silhouettes, different viewpoints or different dressing, remains a crucial challenge. Instead of learning each of these diverse pedestrian appearance features individually as most existing methods do, we propose to perform contrastive learning to guide the feature learning in such a way that the semantic distance between pedestrians with different appearances in the learned feature space is minimized to eliminate the appearance diversities, whilst the distance between pedestrians and background is maximized. To facilitate the efficiency and effectiveness of contrastive learning, we construct an exemplar dictionary with representative pedestrian appearances as prior knowledge to construct effective contrastive training pairs and thus guide contrastive learning. Besides, the constructed exemplar dictionary is further leveraged to evaluate the quality of pedestrian proposals during inference by measuring the semantic distance between the proposal and the exemplar dictionary. Extensive experiments on both daytime and nighttime pedestrian detection validate the effectiveness of the proposed method.
Zebin Lin, Wenjie Pei, Fanglin Chen 0001, David Zhang 0001, Guangming Lu 0002
IEEE Trans. Image Process.3
2022 Few-Shot Object Detection by Knowledge Distillation Using Bag-of-Visual-Words Representations
Wenjie Pei, Dianwen Mei, Fanglin Chen 0001, Jiandong Tian, Guangming Lu 0002
ECCV (10)4
2022 Multi-faceted Distillation of Base-Novel Commonality for Few-Shot Object Detection
Wenjie Pei, Dianwen Mei, Fanglin Chen 0001, Jiandong Tian, Guangming Lu 0002
ECCV (9)4
2022 Correlation-Based Transformer Tracking
Minghan Zhong, Fanglin Chen 0001, Jun Xu 0008, Guangming Lu 0002
ICANN (1)2
2022 Pruning Based Training-Free Neural Architecture Search
abstract
Neural Architecture Search (NAS) plays an important role in searching for high-performance neural networks. How-ever, NAS algorithms are slow and require a terrific amount of computing resources, because they need to be trained on supernet or dense candidate networks to obtain information for evaluation. If the high-performance network architecture could be selected without training, it would eliminate a signif-icant part of the computational cost. Therefore, we propose a zero-cost metric called EX-score, which can represent the ex-pressivity of the network and rank the untrained architectures. To further reduce cost, we design a pruning based zero-cost neural architecture search framework (PZ-NAS) using EX-score. PZ-NAS can prune the initialised supernet rapidly and obtains hundreds of times faster speed performance, whilst archieving comparable accuracy property on CIFAR-IO and ImageNet.
Jiawang Zhou, Fanglin Chen 0001, Guangming Lu 0002
ICME2
2022 Learning Generalizable Latent Representations for Novel Degradations in Super-Resolution
abstract
Typical methods for blind image super-resolution (SR) focus on dealing with unknown degradations by directly estimating them or learning the degradation representations in a latent space. A potential limitation of these methods is that they assume the unknown degradations can be simulated by the integration of various handcrafted degradations (e.g., bicubic downsampling), which is not necessarily true. The real-world degradations can be beyond the simulation scope by the handcrafted degradations, which are referred to as novel degradations. In this work, we propose to learn a latent representation space for degradations, which can be generalized from handcrafted (base) degradations to novel degradations. Furthermore, we perform variational inference to match the posterior of degradations in latent representation space with a prior distribution (e.g., Gaussian distribution). Consequently, we are able to sample more high-quality representations for a novel degradation to augment the training data for SR model. We conduct extensive experiments on both synthetic and real-world datasets to validate the effectiveness and advantages of our method for blind super-resolution with novel degradations.
Fengjun Li, Xin Feng 0005, Fanglin Chen 0001, Guangming Lu 0002, Wenjie Pei
ACM Multimedia3
2022 Multi-modal Finger Feature Fusion Algorithms on Large-Scale Dataset
Chuhao Zhou, Yuanrong Xu, Fanglin Chen 0001, Guangming Lu 0002
PRCV (2)3
2022 Multiscale feature fusion for surveillance video diagnosis
Fanglin Chen 0001, Weihang Wang 0005, Huiyuan Yang, Wenjie Pei, Guangming Lu 0002
Knowl. Based Syst.1
2022 Neighborhood-Exact Nearest Neighbor Search for face retrieval
Fanglin Chen 0001, Wenjie Pei, Guangming Lu 0002
Knowl. Based Syst.1
2022 Stepwise-Refining Speech Separation Network via Fine-Grained Encoding in High-Order Latent Domain
abstract
The crux of single-channel speech separation is how to encode the mixture of signals into such a latent embedding space that the signals from different speakers can be precisely separated. Existing methods for speech separation either transform the speech signals into frequency domain to perform separation or seek to learn a separable embedding space by constructing a latent domain based on convolutional filters. While the latter type of methods learning an embedding space achieves substantial improvement for speech separation, we argue that the embedding space defined by only one latent domain does not suffice to provide a thoroughly separable encoding space for speech separation. In this paper, we propose the Stepwise-Refining Speech Separation Network (SRSSN), which follows a coarse-to-fine separation framework. It first learns a 1-order latent domain to define an encoding space and thereby performs a rough separation in the coarse phase. Then the proposedSRSSNlearns a new latent domain along each basis function of the existing latent domain to obtain a high-order latent domain in the refining phase, which enables our model to perform a refining separation to achieve a more precise speech separation. We demonstrate the effectiveness of ourSRSSNby conducting extensive experiments, including speech separation in a clean (noise-free) setting on WSJ0-2/3mix datasets as well as in noisy/reverberant settings on WHAM!/WHAMR! datasets. Furthermore, we also perform experiments of speech recognition on separated speech signals by our model to evaluate the performance of speech separation indirectly.
Zengwei Yao, Wenjie Pei, Fanglin Chen 0001, Guangming Lu 0002, David Zhang 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Generative Memory-Guided Semantic Reasoning Model for Image Inpainting
abstract
The critical challenge of single image inpainting stems from accurate semantic inference via limited information while maintaining image quality. Typical methods for semantic image inpainting train an encoder-decoder network by learning a one-to-one mapping from the corrupted image to the inpainted version. While such methods perform well on images with small corrupted regions, it is challenging for these methods to deal with images with large corrupted area due to two potential limitations. 1) Such one-to-one mapping paradigm tends to overfit each single training pair of images; 2) The inter-image prior knowledge about the general distribution patterns of visual semantics, which can be transferred across images sharing similar semantics, is not explicitly exploited. In this paper, we propose the Generative Memory-guided Semantic Reasoning Model (GM-SRM), which infers the content of corrupted regions based on not only the known regions of the corrupted image, but also the learned inter-image reasoning priors characterizing the generalizable semantic distribution patterns between similar images. In particular, the proposed GM-SRM first pre-learns a generative memory from the whole training data to explicitly learn the distribution of different semantic patterns. Then the learned memory are leveraged to retrieve the matching semantics for the current corrupted image to perform semantic reasoning during image inpainting. While the encoder-decoder network is used for guaranteeing the pixel-level content consistency, our generative priors are favorable for performing high-level semantic reasoning, which is particularly effective for inferring semantic content for large corrupted area. Extensive experiments on Paris Street View, CelebA-HQ, and Places2 benchmarks demonstrate that our GM-SRM outperforms the state-of-the-art methods for image inpainting in terms of both visual quality and quantitative metrics.
Xin Feng 0005, Wenjie Pei, Fengjun Li, Fanglin Chen 0001, David Zhang 0001, Guangming Lu 0002
IEEE Trans. Circuits Syst. Video Technol.4
2022 High Resolution Fingerprint Retrieval Based on Pore Indexing and Graph Comparison
abstract
Fingerprint retrieval aims to identify a query fingerprint image in a large database using indexing algorithms. Because of the abundant level 3 pore features within high-resolution fingerprint images, pore-based fingerprint retrieval algorithms have been rapidly developed. These retrieval algorithms, however, suffer from severe calculation-consuming problems with the pores increasing. This paper proposes a pore-based fingerprint retrieval method for high-resolution fingerprint images. The proposed method consists of two main steps. 1) In the pore indexing step, an indexing space is constructed using the binary codes of pores in enrolled images. Then, a designed graph-based searching algorithm searches the nearest neighbors of pores from the query image to construct one-to-many correspondences. 2) In the refinement step, the one-to-many correspondences are refined by a random walker-based graph comparison algorithm to remove the false correspondences. The remained nearest neighbors are used to calculate the similarities between the query image and the enrolled images. The proposed method is evaluated on two databases, showing that our method achieves better retrieval accuracies with a higher speed than the existing pore-based retrieval algorithms.
Yuanrong Xu, Yao Lu 0008, Fanglin Chen 0001, Guangming Lu 0002, David Zhang 0001
IEEE Trans. Inf. Forensics Secur.3
2022 Semantic-Interactive Graph Convolutional Network for Multilabel Image Recognition
abstract
Multilabel image recognition, a critically practical task in computer vision, aims to predict multiple objects present in each image. The existing studies mainly focus on conceptual visual cues but fail to reconcile the visual information with their semantic guidance. Intuitively, humans can not only associate extra topological concepts but also imagine other approximate scenes based on a semantic description. Inspired by such semantic-interactive capability, two different types of semantic priors, i.e., the concept correlations of the same scene and semantic similarities among different scenes, should be further explored for the recognition decisions. To efficiently interact with these semantic relationships, in this article, we propose a novel semantic-interactive graph convolutional network (SI-GCN), which can leverage the topological information learned from knowledge graphs to boost the performance of multilabel recognition. Specifically, the proposed SI-GCN framework consists of two different GCN-based branches in parallel, i.e., concept correlations learning (CCL) branch and semantic similarity learning (SSL) branch. Inputting the semantic-embedding vectors of all the concepts, the CCL branch maps the label co-occurrence graph into a set of interdependent concept classifiers. Recalibrating the image feature embedding with the standardized supervision of the semantic similarity graph, the SSL branch learns the semantically consistent in-batch visual representations. Finally, a well-established interactive learning scheme is formulated to concurrently optimize the obtained concept classifiers and the visual representation learning in an end-to-end manner. Extensive experiments on the MS-COCO and Pascal VOC 2007 & 2012 benchmarks demonstrate the superiorities of the proposed SI-GCN method compared to the state-of-the-art baselines.
Bingzhi Chen, Zheng Zhang 0006, Yao Lu 0008, Fanglin Chen 0001, Guangming Lu 0002, David Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.4
2021 Contrastive Feature Decomposition for Image Reflection Removal
abstract
The crux of image reflection removal stems from the difficulty of recognizing the diverse reflection patterns. Typical methods optimize the modeling of background restoration by performing low-level supervision on the restored image to minimize its per-pixel difference from the groundtruth, which re-lies on substantial training samples to learn diverse reflection patterns robustly and avoid overfitting spurious reflection patterns. In this work, we perform supervision on the contrastive distribution between the predicted background and the reflection image. Specifically, our proposed method restores the background and the reflection images in parallel, and seeks to maximize the distribution consistency between the predicted background-reflection contrast and the groundtruth contrast in the latent space. Such supervision pushes the model to focus on contrastive modeling between the background and reflection image. Extensive experiments on four real-world bench-marks demonstrate that our method consistently outperforms state-of-the-art methods.
Xin Feng 0005, Haobo Ji, Bo Jiang 0017, Wenjie Pei, Fanglin Chen 0001, Guangming Lu 0002
ICME5
2021 Deep-Masking Generative Network: A Unified Framework for Background Restoration From Superimposed Images
abstract
Restoring the clean background from the superimposed images containing a noisy layer is the common crux of a classical category of tasks on image restoration such as image reflection removal, image deraining and image dehazing. These tasks are typically formulated and tackled individually due to diverse and complicated appearance patterns of noise layers within the image. In this work we present the Deep-Masking Generative Network (DMGN), which is a unified framework for background restoration from the superimposed images and is able to cope with different types of noise. Our proposed DMGN follows a coarse-to-fine generative process: a coarse background image and a noise image are first generated in parallel, then the noise image is further leveraged to refine the background image to achieve a higher-quality background image. In particular, we design the novel Residual Deep-Masking Cell as the core operating unit for our DMGN to enhance the effective information and suppress the negative information during image generation via learning a gating mask to control the information flow. By iteratively employing this Residual Deep-Masking Cell, our proposed DMGN is able to generate both high-quality background image and noisy image progressively. Furthermore, we propose a two-pronged strategy to effectively leverage the generated noise image as contrasting cues to facilitate the refinement of the background image. Extensive experiments across three typical tasks for image background restoration, including image reflection removal, image rain steak removal and image dehazing, show that our DMGN consistently outperforms state-of-the-art methods specifically designed for each single task.
Xin Feng 0005, Wenjie Pei, Zihui Jia, Fanglin Chen 0001, David Zhang 0001, Guangming Lu 0002
IEEE Trans. Image Process.4
2018 Gender Identification of Human Brain Image with A Novel 3D Descriptor
abstract
Determining gender by examining the human brain is not a simple task because the spatial structure of the human brain is complex, and no obvious differences can be seen by the naked eyes. In this paper, we propose a novel three-dimensional feature descriptor, the three-dimensional weighted histogram of gradient orientation (3D WHGO) to describe this complex spatial structure. The descriptor combines local information for signal intensity and global three-dimensional spatial information for the whole brain. We also improve a framework to address the classification of three-dimensional images based on MRI. This framework, three-dimensional spatial pyramid, uses additional information regarding the spatial relationship between features. The proposed method can be used to distinguish gender at the individual level. We examine our method by using the gender identification of individual magnetic resonance imaging (MRI) scans of a large sample of healthy adults across four research sites, resulting in up to individual-level accuracies under the optimized parameters for distinguishing between females and males. Compared with previous methods, the proposed method obtains higher accuracy, which suggests that this technology has higher discriminative power. With its improved performance in gender identification, the proposed method may have the potential to inform clinical practice and aid in research on neurological and psychiatric disorders.
Fanglin Chen 0001, Dewen Hu
IEEE ACM Trans. Comput. Biol. Bioinform.2
2015 Improve scene classification by using feature and kernel combination
Fanglin Chen 0001, Li Zhou 0003, Dewen Hu
Neurocomputing2
2015 Including Signal Intensity Increases the Performance of Blind Source Separation on Brain Imaging Data
abstract
When analyzing brain imaging data, blind source separation (BSS) techniques critically depend on the level of dimensional reduction. If the reduction level is too slight, the BSS model would be overfitted and become unavailable. Thus, the reduction level must be set relatively heavy. This approach risks discarding useful information and crucially limits the performance of BSS techniques. In this study, a new BSS method that can work well even at a slight reduction level is presented. We proposed the concept of "signal intensity" which measures the significance of the source. Only picking the sources with significant intensity, the new method can avoid the overfitted solutions which are nonexistent artifacts. This approach enables the reduction level to be set slight and retains more useful dimensions in the preliminary reduction. Comparisons between the new and conventional algorithms were performed on both simulated and real data.
Ming Li 0028, Yadong Liu 0001, Fanglin Chen 0001, Dewen Hu
IEEE Trans. Medical Imaging3
2014 Action recognition by hidden temporal models
Jianzhai Wu, Dewen Hu, Fanglin Chen 0001
Vis. Comput.3
2013 A Fusion Method for Partial Fingerprint Recognition
abstract
Conventional algorithms for fingerprint recognition are mainly based on minutiae information. However, the small number of minutiae in partial fingerprints is still a challenge in fingerprint matching. In this paper, a novel algorithm is proposed to improve the performance of partial fingerprint matching. A simulation scheme was firstly proposed to construct a serial of partial fingerprints with different area. Then, the influence of the fingerprint area in partial fingerprint recognition is studied. By comparing the performance of partial fingerprint recognition with different fingerprint area, some useful conclusions can be drawn: (1) The decrease of the fingerprint area degrades the performance of partial fingerprint recognition; (2) When the fingerprint area decreases, the genuine matching scores will decrease, whereas the imposter matching scores will increase. Based on these observations, we proposed a fusion scheme based on modified support vector machine (SVM) to combine the area information for fingerprint recognition. Experimental result illustrates the effectiveness of the proposed method.
Fanglin Chen 0001, Ming Li 0028, Yi Zhang 0104
Int. J. Pattern Recognit. Artif. Intell.1
2013 Hierarchical Minutiae Matching for Fingerprint and Palmprint Identification
abstract
Fingerprints and palmprints are the most common authentic biometrics for personal identification, especially for forensic security. Previous research have been proposed to speed up the searching process in fingerprint and palmprint identification systems, such as those based on classification or indexing, in which the deterioration of identification accuracy is hard to avert. In this paper, a novel hierarchical minutiae matching algorithm for fingerprint and palmprint identification systems is proposed. This method decomposes the matching step into several stages and rejects many false fingerprints or palmprints on different stages, thus it can save much time while preserving a high identification rate. Experimental results show that the proposed algorithm can save almost 50% searching time compared with traditional methods and illustrate its effectiveness.
Fanglin Chen 0001, Xiaolin Huang, Jie Zhou 0001
IEEE Trans. Image Process.1
2011 Separating Overlapped Fingerprints
abstract
Fingerprint images generally contain either a single fingerprint (e.g., rolled images) or a set of nonoverlapped fingerprints (e.g., slap fingerprints). However, there are situations where several fingerprints overlap on top of each other. Such situations are frequently encountered when latent (partial) fingerprints are lifted from crime scenes or residue fingerprints are left on fingerprint sensors. Overlapped fingerprints constitute a serious challenge to existing fingerprint recognition algorithms, since these algorithms are designed under the assumption that fingerprints have been properly segmented. In this paper, a novel algorithm is proposed to separate overlapped fingerprints into component or individual fingerprints. The basic idea is to first estimate the orientation field of the given image with overlapped fingerprints and then separate it into component orientation fields using a relaxation labeling technique. We also propose an algorithm to utilize fingerprint singularity information to further improve the separation performance. Experimental results indicate that the algorithm leads to good separation of overlapped fingerprints that leads to a significant improvement in the matching accuracy.
Fanglin Chen 0001, Jianjiang Feng, Anil K. Jain 0001, Jie Zhou 0001
IEEE Trans. Inf. Forensics Secur.1
2010 A hierarchical algorithm for multi-feature based fingerprint identification
abstract
There are different features, such as minutiae, orientation-based minutia descriptor, FingerCode, ridge feature map, orientation map, and density map, to represent fingerprints. Previous studies showed that the performance can be improved by combining these features through a fusion strategy. However, the more features are used, the more time is consumed. In fingerprint verification application, it is tolerated for the system is one-to-one matching. But in automatic fingerprint identification systems (AFIS) which is one-to-N matching (N is the number of fingerprint in the database and it is usually very large), it is not acceptable. In this paper, a fast algorithm for multi-feature based fingerprint identification is proposed. A hierarchical strategy is utilized to quickly discard most false matchings in different stages. Experimental results show that the proposed algorithm can reduce the searching time a lot while preserves the searching ratio.
Fanglin Chen 0001, Jie Zhou 0001
ICIP1
2009 A Novel Algorithm for Detecting Singular Points from Fingerprint Images
abstract
Fingerprint analysis is typically based on the location and pattern of detected singular points in the images. These singular points (cores and deltas) not only represent the characteristics of local ridge patterns but also determine the topological structure (i.e., fingerprint type) and largely influence the orientation field. In this paper, we propose a novel algorithm for singular points detection. After an initial detection using the conventional Poincaré Index method, a so-called DORIC feature is used to remove spurious singular points. Then, the optimal combination of singular points is selected to minimize the difference between the original orientation field and the model-based orientation field reconstructed using the singular points. A core-delta relation is used as a global constraint for the final selection of singular points. Experimental results show that our algorithm is accurate and robust, giving better results than competing approaches. The proposed detection algorithm can also be used for more general 2D oriented patterns, such as fluid flow motion, and so forth.
Jie Zhou 0001, Fanglin Chen 0001, Jinwei Gu
IEEE Trans. Pattern Anal. Mach. Intell.2
2009 Crease detection from fingerprint images and its applications in elderly people
Jie Zhou 0001, Fanglin Chen 0001
Pattern Recognit.2
2009 Reconstructing Orientation Field From Fingerprint Minutiae to Improve Minutiae-Matching Accuracy
abstract
Minutiae are very important features for fingerprint representation, and most practical fingerprint recognition systems only store the minutiae template in the database for further usage. The conventional methods to utilize minutiae information are treating it as a point set and finding the matched points from different minutiae sets. In this paper, we propose a novel algorithm to use minutiae for fingerprint recognition, in which the fingerprint's orientation field is reconstructed from minutiae and further utilized in the matching stage to enhance the system's performance. First, we produce "virtual" minutiae by using interpolation in the sparse area, and then use an orientation model to reconstruct the orientation field from all "real" and "virtual" minutiae. A decision fusion scheme is used to combine the reconstructed orientation field matching with conventional minutiae-based matching. Since orientation field is an important global feature of fingerprints, the proposed method can obtain better results than conventional methods. Experimental results illustrate its effectiveness.
Fanglin Chen 0001, Jie Zhou 0001, Chunyu Yang 0005
IEEE Trans. Image Process.1