Bin Fang 0001

dblp:94/4033-1 · DBLP profile ↗
← Back
130ranked-venue papers
6as first author
52since 2021 · last 2026
0000-0003-1955-6626ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 81 · 5 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 5 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-authorSystems, architecture and hardware · 3 · 2 since 2021Computer networks · 3 · 2 since 2021Software engineering, systems software and programming languages · 2
YearPublicationVenuePosition
2026 Symmetric Positive Definite manifold deep metric learning for bearing fault diagnosis
Junshi Cheng, Ruisheng Ran, Bin Fang 0001, Benchao Li
Eng. Appl. Artif. Intell.3
2026 SPD-DANN: An SPD manifold unsupervised domain adaptation method for cross subject motor imagery EEG decoding
Junshi Cheng, Ruisheng Ran, Bin Fang 0001
Neural Networks3
2026 FG-MoE: Heterogeneous mixture of experts model for fine-grained visual classification
Songming Yang, Bin Fang 0001
Pattern Recognit.3
2025 SPD-DFNet: A SPD Manifold Dual Flow Unsupervised Domain Adaptation Network for EEG Decoding
abstract
EEG signals contain valuable physiological and psychological information, essential for brain-computer interfaces (BCIs) and neurorehabilitation. However, the non-stationarity and subject-specific variability of EEG make cross-subject generalization difficult, limiting practical deployment without costly recalibration. Unsupervised domain adaptation (UDA) aims to reduce domain discrepancies for better generalization. While many UDA methods use discrepancy-based or adversarial strategies to extract domain-invariant features, they are limited by Euclidean space, which doesn't capture EEG's nonlinear relationships. To address this, we propose a novel UDA framework on the Riemannian manifold of SPD matrices, using adversarial training. Our approach features a dual flow architecture with separate feature extractors for the source and target domains, incorporating a manifold soft parameter sharing mechanism and multi-level alignment loss for better domain alignment and feature separability. Unlike traditional methods, our model preserves domain-specific structures while aligning the feature extractors' outputs, capturing more effective domain-invariant information. Experiments on three BCI datasets show that our method outperforms several state-of-the-art UDA approaches in cross-subject EEG classification.
Junshi Cheng, Ruisheng Ran, Bin Fang 0001
BIBM4
2025 A Local Structure-based Convolutional Neural Network and Its Application to Fault Diagnosis
abstract
In contemporary industry, fault diagnosis is essential for detecting equipment health status and ensuring normal equipment operation. Currently, the application of the feedforward convolutional neural network method has been embraced in fault diagnosis, revealing substantial potential. A notable example of this is the PCANet. However, PCA is a linear method, which assumes the data has a global linear structure or is linearly separable. But the signal data from mechanical equipment is more of a nonlinear structure in the high-dimensional space and PCA cannot well represent the complex data and even lose some information. To address this issue, based on PCANet, a local structure-based Convolutional neural network, named LPPNet, is proposed by using locality preserving projection (LPP) to learn the convolutional filter of CNN. Because LPP can deal with complex nonlinear data and maintain the local structure of the data, more representative features can be obtained when it is used to calculate the convolution filter. This method is evaluated with the comprehensive experiments about fault diagnosis. The experiments show that LPPNet achieves a lower detection error rate compared to PCANet and some representative methods, which can be beneficial for practical applications.
Ruisheng Ran, Weizhe Ding, Bin Fang 0001
IJCNN4
2025 Incremental Information-Aware: Mine Abundant and Accurate Information for Video Captioning
abstract
This paper proposes the Incremental Information-Aware (IIA) framework to enhance video captioning by generating semantically abundant and accurate descriptions. The Semantic Incremental Information Aware model (Semantic-IIA) aims to capture detailed semantic information, while the Structural Incremental Information Aware model (Structural-IIA) focuses on identifying key structural content. These models are designed in parallel to complement each other, ensuring both semantic abundance and structural accuracy in generated captions. Experiments on the MSVD and MSR-VTT datasets demonstrate superior performance, with CIDEr scores of 108.4 and 59.8, and BLEU-4 scores of 61.0 and 47.3, respectively. Our code is available at https://github.com/Zhongnibug/IIA.
Ningkai Zhong, Bin Fang 0001, Mengdi Li 0005, Langping Wang
ICMR2
2025 SEDS: Semantic Emphasis and Detail Supplementation for Accurate and Comprehensive Video Captioning
Mengdi Li 0005, Bin Fang 0001, Langping Wang, Ningkai Zhong
PRCV (11)2
2025 Advancing Video Captioning via Visual-Linguistic Feature Fusion
abstract
Video captioning is a challenging multimodal task that combines computer vision and natural language processing. Previous methods primarily extract useful information from visual features by designing sophisticated encoders. Natural language descriptions generated solely from visual features often fall short of expected performance due to the semantic gap between the visual modality and the textual modality. In this paper, we propose a novel encoder–decoder-based approach for video captioning, which enhances video feature representations by incorporating object and action-centric linguistic features from upstream encoders. Specially, we leverage a cross-attention mechanism to incorporate textual information into visual features, thereby enhancing their expressive capability. Additionally, we introduce a fusion layer to facilitate the interaction between heterogeneous representations and mitigating irrelevant noise. Extensive experiments on the MSVD, MSR-VTT and VATEX datasets demonstrate the superiority of our method and achieve significantly superior performance across standard evaluation metrics.
Bin Fang 0001
Int. J. Pattern Recognit. Artif. Intell.2
2025 MSFM-UNET: enhancing medical image segmentation with multi-scale and multi-view frequency fusion
Qiang Gao 0017, Yi Wang 0074, Feiyan Zhou, Yong Li 0023, Bin Fang 0001, Lan Du 0002, Cunjian Chen
Pattern Anal. Appl.6
2025 Image-set classification using Discriminant Neighborhood Preserving Embedding on Grassmann manifold
Benchao Li, Yuanyuan Zheng, Ruisheng Ran, Bin Fang 0001
Signal Process.4
2025 Sparse Reduced-Rank Fully Connected Layers with Its Applications in Detection and Classification
abstract
Fully connected (FC) layers play a significant role in deep neural networks (DNNs) models. Owing to the complexity of its parameters, an FC layer has sufficient capacity to manage high-dimensional tasks, so a large amount of memory and powerful computing capabilities become essential requirements. However, the large number of parameters in an FC layer greatly limits the practical application of this model. To address this problem, we apply matrix optimization to an FC layer. First, an added penalty term properly maintains the sparsity of the imposed weights. Second, a rank constraint is applied to the two components of the factorized weight matrix. Our compression algorithm can effectively reduce the number of required network parameters, which not only reduces the computational complexity of the network but also results in better generalizability on a test dataset. Finally, the effectiveness of the proposed method is verified in two different computer vision task domains. Experiments show that our sparse reduced-rank method achieves a better compression ratio with a lower accuracy loss relative to the competing approaches. The code is available at https://github.com/cheer79/Compress_FC .
Mingliang Zhou 0001, Xuekai Wei, Yong Feng 0002, Tao Xiang 0001, Bin Fang 0001, Zhaowei Shang, Fan Jia 0005, Xu Zhuang, Huayan Pu, Jun Luo 0003
ACM Trans. Multim. Comput. Commun. Appl.6
2024 Cardiac MR Image Semantic Segmentation based on Joint Dual-stream CNNs
abstract
Image semantic segmentation is the pixel-level understanding and classification of visual regions of interest. In this paper, a novel cardiac magnetic resonance image semantic segmentation architecture is presented based on a joint dual-stream convolutional neural network (CNN) called JDSCNN. The method effectively integrates multi-level, multi-layer and multi-scale context and detailed information to perform semantic segmentation. First, a semantic stream and a detail stream are used in the proposed network to extract the context and edge features respectively. Specifically, for the semantic stream, a gated fully fusion module is integrated as the skip connection to extract multi-level features. Besides, a feature pyramid module is proposed to concatenate multi-layer features in the decoder path. A feature aggregation module is also used to obtain multi-scale context. For the detail stream, a gated shape CNN exploiting multi-level context is proposed to extract edge features of the objects. Second, we integrate an inter-stream fusion module to fuse and propagate the features in each stream bidirectionally. Finally, a multi-task loss function combining the edge information and segmentation information is applied. JDSCNN is evaluated on the public cardiac image dataset, which achieves significant results for the left ventricle, right ventricle and myocardium compared with the state-of-the-art algorithm. Our code is available at: https://github.com/csathu/JDSCNN.
Hengqi Hu, Bin Fang 0001, Mingliang Zhou 0001
BIBM2
2024 Perceiving Multi-Layer Representations for No-reference Image Quality Assessment
abstract
In this paper, we propose an end-to-end no-reference image quality assessment (NR-IQA) method that perceives multilayer representations from low-level to high-level stages. First, multi-layer representations (MR) of the distorted images are extracted from different layers of multiple feature extraction networks to obtain fine-grained information. Second, a gated recurrent unit (GRU)-based fusion encoder (GFE) is presented to model the interrelationships between multi-layer representations, thereby generating the global feature. Finally, we construct a perception-oriented quality regression network (PQRN) to generate the quality scores. Experimental results on commonly used benchmark datasets verify the effectiveness of our proposed method over existing state-of-the-art approaches by a large margin.
Qunyue Huang, Bin Fang 0001, Xi Ai, Tianyu Nie
ICASSP2
2024 OSFENet: Object Spatiotemporal Feature Enhanced Network for Surgical Phase Recognition
Pingjie You, Hengqi Hu, Yi Wang 0074, Bin Fang 0001
ICIC (12)5
2024 Locality Preserving Projections with Autoencoder
Ruisheng Ran, Ji Feng, Bin Fang 0001
Expert Syst. Appl.5
2024 Representation Learning Based on Vision Transformer
abstract
In recent years, with the rapid development of information technology, the volume of image data has grown exponentially. However, these datasets typically contain a large amount of redundant information. To extract effective features and reduce redundancy from images, a representation learning method based on the Vision Transformer (ViT) has been proposed, and to our best knowledge, Transformer was first applied to zero-shot learning (ZSL). The method adopts a symmetric encoder–decoder structure, where the encoder incorporates Multi-Head Self-Attention (MSA) mechanism of ViT to reduce the dimensionality of image features, eliminate redundant information, and decrease computational burden. Consequently, it effectively extracts features, and the decoder is utilized for reconstructing image data. We evaluated the representation learning capability of the proposed method in various tasks, including data visualization, image reconstruction, face recognition, and ZSL. By comparing with state-of-the-art representation learning methods, the outstanding results obtained validate the effectiveness of this method in the field of representation learning.
Ruisheng Ran, Qianwei Hu, Wenfeng Zhang, Shunshun Peng, Bin Fang 0001
Int. J. Pattern Recognit. Artif. Intell.6
2024 Saliency and Depth-Aware Full Reference 360-Degree Image Quality Assessment
abstract
With the widespread adoption of virtual reality and 360-degree video, there is a pressing need for objective metrics to assess quality in this immersive panoramic format reliably. However, existing image quality assessment models developed for traditional fixed-viewpoint content do not fully consider the specific perceptual issues involved in 360-degree viewing. This paper proposes a 360-degree image full-reference quality assessment (FR-IQA) methodology based on a multi-channel architecture. The proposed 360-degree FR-IQA method further optimizes and identifies the distorted image quality using two easily obtained useful saliency and depth-aware image features. The convolutional neural network (CNN) is designed for training. Furthermore, the proposed method accounts for predicting user viewing behaviors within 360-degree images, which will further benefit the multi-channel CNN architecture and enable the weighted average pooling of the predicted FR-IQA scores. The performance is evaluated on publicly available databases to demonstrate the advantages brought by the proposed multi-channel model in performance evaluation and cross-database evaluation experiments, where it outperforms other state-of-the-art ones. Moreover, an ablation study exhibits good generalization ability and robustness.
Xuekai Wei, Qunyue Huang, Bin Fang 0001, Lei Ouyang, Weizhi Xian, Jun Luo 0003, Huayan Pu, Xueyong Xu, Chang Lu 0005, Hao Nan, Xu Liu 0006, Yachao Li 0001, Mingliang Zhou 0001
Int. J. Pattern Recognit. Artif. Intell.3
2024 Deep Dual-Stream Convolutional Neural Networks for Cardiac Image Semantic Segmentation
abstract
Cardiac image segmentation is essential when applying biomedical informatics to improve industrial healthcare applications. To extract context and detailed information more efficiently and further improve cardiac image segmentation accuracy, we present a novel deep dual-stream convolutional neural network (CNN) for cardiac image semantic segmentation in this article. We use a body stream and a shape stream, respectively, in this method. First, in the body stream we propose integrating a gated fully fusion module to fuse multilevel features in the encoder and decoder paths. In addition, we integrate a feature aggregation module to extract the multiscale context. Second, in the shape stream, we propose using a gated shape CNN exploiting multilevel context to extract detailed information, such as boundary and shape features. Finally, we apply a multitask loss function to align the predicted masks with the ground truth labels. Our experiments on the public cardiac magnetic resonance image dataset show significant performance in the left and right ventricular cavities and myocardium compared to the state-of-the-art algorithms.
Hengqi Hu, Bin Fang 0001, Yuting Ran, Xuekai Wei, Weizhi Xian, Mingliang Zhou 0001, Sam Kwong
IEEE Trans. Ind. Informatics2
2024 Polynomial linear discriminant analysis
Ruisheng Ran, Ting Wang 0040, Bin Fang 0001
J. Supercomput.4
2024 Perceptual Quality Analysis in Deep Domains Using Structure Separation and High-Order Moments
abstract
Images are composed of “things” (i.e., structured objects) and “stuff” (i.e., textured surfaces), which have completely different effects on the human visual system (HVS). A good image quality assessment (IQA) method should fully consider the visual salience effects of image structures and the masking effects of image textures. In this article, we propose a perceptual quality analysis model using structure separation and high-order moments (SSHMPQA) in the deep domain. First, we use a total variation (TV) model to separate the perceptual structures in images from their deep feature maps, thereby maintaining meaningful object shapes with texture suppression and defining perceptual structure-aware distances in the deep domain. Then, we use the first- to fourth-order moments to calculate the mean, skewness and kurtosis of the probability distributions of the deep features. On this basis, we define a perceptual texture-aware distance in the deep domain. We then formulate the final model by solving a well-defined perceptual optimization problem. The proposed SSHMPQA model has good interpretability and is data-driven; moreover, the model does not require a complex and long training process because the optimization problem is convex and has an exact analytical solution. To verify the effectiveness of our model, comprehensive experiments are conducted. The experimental results show that the proposed model is superior to other state-of-the-art traditional and deep learning-based full-reference (FR) IQA methods.
Weizhi Xian, Mingliang Zhou 0001, Bin Fang 0001, Tao Xiang 0001, Weijia Jia 0001, Bin Chen 0022
IEEE Trans. Multim.3
2024 Low-Light Enhancement Method Based on a Retinex Model for Structure Preservation
abstract
Enhancing low-light image visibility is a critical task in computer vision since it helps to improve input for high-level algorithms. High-quality images typically have clear structural information. In previous studies, due to the lack of proper structural guidance, restored images had some problems, such as unclear structural areas and overexposed or underexposed local areas. To address the above problems, in this paper, we introduce a coefficient of variation (COV) with excellent performance in maintaining structural information, and then we propose a low-light image enhancement method that utilizes the COV to extract structural information from images. First, we apply a traditional retinex model to estimate both reflectance and illumination. Second, we use the COV to indicate the degree of dispersion of the input sample, which enables us to obtain a robust structure-distinguishing weight map for low-light images. The weight map is adaptively divided to obtain a structural weight map, which is then used to enhance the gradient image. This process is applied before the reflectance layer of the retinex model. Finally, the result is obtained by using the block coordinate descent method. According to extensive experiments, outstanding results can be achieved by our proposed method in terms of both subjective and objective evaluation metrics in comparison with other state-of-the-art methods. The source code is available at our website.
Mingliang Zhou 0001, Xingtai Wu, Xuekai Wei, Tao Xiang 0001, Bin Fang 0001, Sam Kwong
IEEE Trans. Multim.5
2023 On-the-fly Cross-lingual Masking for Multilingual Pre-training
abstract
In multilingual pre-training with the objective of MLM (masked language modeling) on multiple monolingual corpora, multilingual models only learn cross-linguality implicitly from isomorphic spaces formed by overlapping different language spaces due to the lack of explicit cross-lingual forward pass.In this work, we present CLPM (Cross-lingual Prototype Masking), a dynamic and token-wise masking scheme, for multilingual pre-training, using a special token [C] x to replace a random token x in the input sentence.[C] x is a cross-lingual prototype for x and then forms an explicit crosslingual forward pass.We instantiate CLPM for the multilingual pre-training phase of UNMT (unsupervised neural machine translation), and experiments show that CLPM can consistently improve the performance of UNMT models on {De, Ro, N e} ↔ En.Beyond UNMT or bilingual tasks, we show that CLPM can consistently improve the performance of multilingual models on cross-lingual classification.
Xi Ai, Bin Fang 0001
ACL (1)2
2023 Advancing Pancreas Segmentation through the Patch-Adjust Fusion Framework
abstract
Pancreas segmentation is pivotal for the timely diagnosis of pancreatic diseases, but it remains challenging due to organ’s small size, large spatial variations, and unclear boundaries. To enhance segmentation accuracy, researchers have explored various multi-modal coarse-to-fine (MMCF) methods, integrating distinct data-related factors such as stage, view, scale, and patch size. However, these methods often ask for multiple trained models, leading to significant storage and computational expenses. Moreover, these methods frequently treat different data-related factors separately, overlooking their potential relationships and constraining the feature representation capability of convolutional neural networks. To tackle these issues, we propose a data-efficient framework known as Patch-Adjust Fusion (PAF). PAF framework addresses these factors concurrently through dynamic patch flow adjustments, seamlessly integrating diverse elements while accounting for their interactions and dependencies. Consequently, the PAF framework demonstrates superior segmentation performance, diminished storage requisites, and expedited inference speeds compared to the MMCF framework. This is exemplified across two public CT pancreas datasets (NIH and MSD) and a private MR pancreas dataset (MR300). These results underscore its suitability for clinical applications, positioning the PAF framework as a promising solution for accurate and efficient pancreas segmentation in medical settings.
Yi Wang 0074, Bin Fang 0001
BIBM3
2023 A Global-Local Contrastive Learning Framework for Video Captioning
abstract
In this paper, a global-local contrastive learning framework is proposed to leverage global contextual information from different modalities and then effectively fuse them with the supervision of contrastive learning. First, a global-local encoder is proposed to sufficiently explore the salient contextual information from different modalities, which generates the global contextual information. Second, contrastive learning is used to minimize the semantic distance between the paired modalities, which can improve the content matching between videos and the predicted captions. Finally, an attention-based multimodal encoder is presented to effectively fuse different modalities, thereby generating the multimodal representations that include global contextual information from different modalities. Extensive experimental results on benchmark datasets indicate that our proposed method is superior to the state-of-the-art approaches.
Qunyue Huang, Bin Fang 0001, Xi Ai
ICIP2
2023 Low-Light Image Enhancement with Contrast Increase and Illumination Smooth
abstract
In image enhancement, maintaining the texture and attenuating noise are worth discussing. To address these problems, we propose a low-light image enhancement method with contrast increase and illumination smooth. First, we calculate the maximum map and the minimum map of RGB channels, and then we set maximum map as the initial value for illumination and introduce minimum map to smooth illumination. Second, we use the histogram-equalized version of the input image to construct the weight for the illumination map. Third, we propose an optimization problem to obtain the smooth illumination and refined reflectance. Experimental results show that our method can achieve better performance compared to the state-of-the-art methods.
Hongyue Leng, Bin Fang 0001, Mingliang Zhou 0001, Qin Mao
Int. J. Pattern Recognit. Artif. Intell.2
2023 Fast 3D Object Measurement Based on Point Cloud Modeling
abstract
Automated object measurement is becoming increasingly important due to its ability to reduce manual costs, increase production efficiency, and minimize errors in various fields. In this paper, we present a novel approach to three-dimensional (3D) object measurement based on point cloud modeling. Our method introduces a fast point cloud modeling computation framework consisting of five stages: coordinate centralization, rotation and translation, noise filtering, plane projection, and geometric computation. Furthermore, we propose a fast convex hull optimization algorithm to reduce the high complexity problem of traditional convex hull calculation. Our extensive experiments demonstrate that our approach outperforms existing methods in terms of measurement error rate and time savings, with a maximum time saving of 31.03% under certain error conditions.
Gang Wang 0023, Mingliang Zhou 0001, Bin Fang 0001, Yugui Zhang, Shouqin Guan, Bin Ruan
Int. J. Pattern Recognit. Artif. Intell.3
2023 Improve Session-Based Recommendation with Triplet Mining and Dynamic Perturbations Graph Neural Networks
abstract
Session-based recommendation (SBR) emphasizes mining user interests to predict the next click based on recent interactions within sessions. Most current SBR methods suffer from insufficient interactive information problems and fail to distinguish session representations with high similarities, which can neglect the inherent features within sessions. To fill the gap, we propose a triplet mining enhanced graph neural networks (TME-GNN) approach to enhance the recommendation systems by mining structural and inherent information. Technically, we first generate anchor, positive and negative embeddings based on the given session and set a triplet mining task to improve the recommendation task with subtle features by pushing positive pairs close and pulling negative pairs away. Second, to robust the model, we employ a self-supervised auxiliary task by adding dynamic perturbations to the embedding space. We conduct extensive experiments to demonstrate the superiority of our method against other state-of-the-art algorithms. Our implementations are available on the following site https://github.com/Info4Rec/TME-GNN .
Yong Feng 0002, Mingliang Zhou 0001, Xiancai Xiong, Yongheng Wang, Baohua Qiang, Qin Mao, Bin Fang 0001
Int. J. Pattern Recognit. Artif. Intell.9
2023 Cross-Modal Language Modeling in Multi-Motion-Informed Context for Lip Reading
abstract
We observe that for lip reading, the language is locally transformed, instead of globally transformed, i.e., speaking and writing follow the same basic grammar rules. In this work, we present a cross-modal language model to tackle the lip-reading challenge on silent videos. Compared to previous works, we consider multi-motion-informed contexts composed of multiple lip-motion representations from different subspaces to guide decoding via the source-target attention mechanism. We present a piece-wise pre-training strategy inspired by multi-task learning to pre-train a visual module to generate multi-motioninformed contexts for cross-modality and pre-train a decoder to generate texts for language modeling. Our final large-scale model outperforms baseline models on four datasets: LRS2, LRS3, LRW, and GRID. We will open our source code on GitHub.
Xi Ai, Bin Fang 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Isometric projection with reconstruction
Ruisheng Ran, Qianghui Zeng, Xiaopeng Jiang, Bin Fang 0001
J. Supercomput.4
2023 Low-light Image Enhancement via a Frequency-based Model with Structure and Texture Decomposition
abstract
This article proposes a frequency-based structure and texture decomposition model in a Retinex-based framework for low-light image enhancement and noise suppression. First, we utilize the total variation-based noise estimation to decompose the observed image into low-frequency and high-frequency components. Second, we use a Gaussian kernel for noise suppression in the high-frequency layer. Third, we propose a frequency-based structure and texture decomposition method to achieve low-light enhancement. We extract texture and structure priors by using the high-frequency layer and a low-frequency layer, respectively. We present an optimization problem and solve it with the augmented Lagrange multiplier to generate a balance between structure and texture in the reflectance map. Our experimental results reveal that the proposed method can achieve superior performance in naturalness preservation and detail retention compared with state-of-the-art algorithms for low-light image enhancement. Our code is available on the following website. 1
Mingliang Zhou 0001, Hongyue Leng, Bin Fang 0001, Tao Xiang 0001, Xuekai Wei, Weijia Jia 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2022 Leveraging Relaxed Equilibrium by Lazy Transition for Sequence Modeling
abstract
In sequence modeling, certain tokens are usually less ambiguous than others, and representations of these tokens require fewer refinements for disambiguation.However, given the nature of attention-based models like Transformer and UT (universal transformer), all tokens are equally processed towards depth.Inspired by the equilibrium phenomenon, we present a lazy transition, a mechanism to adjust the significance of iterative refinements for each token representation.Our lazy transition is deployed on top of UT to build LT (lazy transformer), where all tokens are processed unequally towards depth.Eventually, LT is encouraged to oscillate around a relaxed equilibrium.Our experiments show that LT outperforms baseline models on several tasks of machine translation, pre-training, Learning to Execute, and LAMBADA.
Xi Ai, Bin Fang 0001
ACL (1)2
2022 Vocabulary-informed Language Encoding
abstract
A Multilingual model relies on language encodings to identify input languages because the multilingual model has to distinguish between the input and output languages or among all the languages for cross-lingual tasks. Furthermore, we find that language encodings potentially refine multiple morphologies of different languages to form a better isomorphic space for multilinguality. To leverage this observation, we present a method to compute a vocabulary-informed language encoding as the language representation, for a required language, considering a local vocabulary covering an acceptable amount of the most frequent word embeddings in this language. In our experiments, our method can consistently improve the performance of multilingual models on unsupervised neural machine translation and cross-lingual embedding.
Xi Ai, Bin Fang 0001
COLING2
2022 An Accelerated and Flexible SIFT Parallel-Computing Approach Based on the General Multi-Core Platform
abstract
Visual retrieval has been a significant technology in the computer vision task. Visual feature descriptors are the key to the visual retrieval. The famous local feature descriptor is called the Scale Invariant Feature Transform (SIFT), which can keep invariant mapping for the scale, rotate and simulate images. To utilize effectively the SIFT feature descriptor for visual matching on different hardware platforms, this paper proposes an accelerated SIFT algorithm based on the SIFT feature computing principle of the general multi-core platform. First, our multi-core task allocation method introduces the WFM theory into task assignment for each core to improve the core computing resource utilization for high-efficient parallel computing. Then, to improve the efficiency of picture matching, we introduce global geometric constraints condition to optimal picture matching for the multi-core parallelization approach. Experimental results show that the proposed approach can save on average 87.31% on the Intel X86 platform, compared to the single-core time. Also, our approach can save on average 33.79% on the Raspberry Pi platform, compared to the single-core time.
Gang Wang 0023, Mingliang Zhou 0001, Bin Fang 0001, Haichao Huang, Zhenyu Shu, Xueshu Chen
Int. J. Pattern Recognit. Artif. Intell.3
2022 Multi-Modal Sparse Tracking by Jointing Timing and Modal Consistency
abstract
In this paper, we propose a multi-modal sparse tracking by jointing timing and modal consistency to locate the target location with the similarity of multiple local appearances. First, we propose an alignable patching strategy for red-green-blue (RGB) color mode and thermal infrared mode to adapt to the local changes of the target. Second, we propose a consistency expression of the corresponding aligned patches between the modes and the correlation of the gaussian mapping within mode to reconstruct the target judgment likelihood function. Finally, we propose an updating scenario based on timing correlation and mode sparsity to fit with the target changes. According to the experimental results, significant improvement in terms of tracking accuracy can be achieved on average compared with the state-of-the-art algorithms. The source code of our algorithm is available on https://github.com/Liincq/tracker .
Bin Fang 0001, Mingliang Zhou 0001
Int. J. Pattern Recognit. Artif. Intell.2
2022 VLSs: A Local Search Algorithm for Distributed Constraint Optimization Problems
abstract
Local search algorithms are widely applied in solving large-scale distributed constraint optimization problem (DCOP). Distributed stochastic algorithm (DSA) is a typical local search algorithm to solve DCOP. However, DSA has some drawbacks including easily falling into local optima and the unfairness of assignment choice. This paper presents a novel local search algorithm named VLSs to solve the issues. In VLSs, sampling according to the probability corresponding to assignment is introduced to enable each agent to choose other promising values. Besides, each agent alternately performs a greedy choice among multiple parallel solutions to reduce the chance of falling into local optima and a variance adjustment mechanism to guide the search into a relatively good initial solution in a periodic manner. We give the proof of variance adjustment mechanism rationality and theoretical explanation of impact of greed among multiple parallel solutions. The experimental results show the superiority of VLSs over state-of-the-art DCOP algorithms.
Fukui Li, Jingyuan He, Mingliang Zhou 0001, Bin Fang 0001
Int. J. Pattern Recognit. Artif. Intell.4
2022 A Structure Preservation and Denoising Low-Light Enhancement Model via Coefficient of Variation
abstract
In this paper, we propose a structure-preserving and denoising low-light enhancement method that uses the coefficient of variation. First, we use the coefficient of variation to process the original low-light image, which is used to obtain the enhanced illumination gradient reference map. Second, we use the total variation (TV) norm to regularize the reflectance gradient, which is used to maintain the smoothness of the image and eliminate the artifacts in the reflectance estimation. Finally, we combine the above two constraint terms with the Retinex theory, which contains the denoising regular term. The final enhanced and denoised low-light image is obtained by iterative solution. Experimental results show that our method can achieve superior performance in both subjective and objective assessments compared with other state-of-the-art methods (the source code is available at: https://github.com/bbxavi/SPDLEM .).
Xingtai Wu, Jingyuan He, Bin Fang 0001, Zhaowei Shang, Mingliang Zhou 0001
Int. J. Pattern Recognit. Artif. Intell.4
2022 3D Human Pose Estimation via Spatio-Temporal Matching from Monocular RGB Images
abstract
Three-dimensional (3D) human pose estimation aims to locate 3D keypoints of individuals from given input RGB images. For two-dimensional (2D) human pose estimation problems, majority methods inferring 2D poses are from 2D heatmaps. However, it is hard to extend this method to 3D poses inferring area which makes computational loads increase sharply. To address the above problem, we propose STM-CNN method to estimate reconstruction coefficient matrix to calculate the final 3D pose instead of estimating 3D heatmaps to decrease the computational loads. First, STM-CNN does a preprocessing procedure to calculate a set of shape and weight bases. Second, STM-CNN infers a 2D matrix called reconstruction coefficient from the STM-CNN architecture. Third, STM-CNN utilizes the preprocessing shape and weight bases and estimated reconstruction coefficient matrix to calculate the final 3D pose. Meanwhile, STM-CNN method achieves better performances compared with the state-of-the-art methods on Human3.6M.
Jielu Yan, Mingliang Zhou 0001, Bin Fang 0001
Int. J. Pattern Recognit. Artif. Intell.3
2022 Recent Advances in 3D Human Pose Estimation: From Optimization to Implementation and Beyond
abstract
3D human pose estimation describes estimating 3D articulation structure of a person from an image or a video. The technology has massive potential because it can enable tracking people and analyzing motion in real time. Recently, much research has been conducted to optimize human pose estimation, but few works have focused on reviewing 3D human pose estimation. In this paper, we offer a comprehensive survey of the state-of-the-art methods for 3D human pose estimation, referred to as pose estimation solutions, implementations on images or videos that contain different numbers of people and advanced 3D human pose estimation techniques. Furthermore, different kinds of algorithms are further subdivided into sub-categories and compared in light of different methodologies. To the best of our knowledge, this is the first such comprehensive survey of the recent progress of 3D human pose estimation and will hopefully facilitate the completion, refinement and applications of 3D human pose estimation.
Jielu Yan, Mingliang Zhou 0001, Jinli Pan, Meng Yin, Bin Fang 0001
Int. J. Pattern Recognit. Artif. Intell.5
2022 Module Against Power Consumption Attacks for Trustworthiness of Vehicular AI Chips in Wide Temperature Range
abstract
Power consumption attacks monitoring on artificial intelligence (AI) chips play a critical role in the vehicular AI systems. However, most of the current monitoring and management methods focus on the trustworthiness of industrial equipment instead of resource-constrained edge devices. To address the above problem, a closed-loop module for monitoring and management of vehicular AI chips based on fitting and filtering to resist power consumption attacks is proposed in this paper. First, considering the characteristics of power, we propose a raw data correction approach for power monitoring to monitor abnormal power consumption. Second, we address the challenging problem of precision temperature monitoring to monitor the abnormal temperature of the chip, especially in a wide temperature range. Finally, the established method is applied to attack surveillance and transformed into a power consumption management problem solved by dynamic voltage and frequency scaling (DVFS) technology. As the experimental results reveal, compared with existing methods of power and temperature monitoring and power consumption control in wide temperature, our method can achieve significantly improved monitoring and managing performance.
Zongwei Zhu, Jiawei Geng, Mingliang Zhou 0001, Bin Fang 0001
Int. J. Pattern Recognit. Artif. Intell.4
2022 A content-oriented no-reference perceptual video quality assessment method for computer graphics animation videos
Weizhi Xian, Mingliang Zhou 0001, Bin Fang 0001, Sam Kwong
Inf. Sci.3
2022 An active contour model based on adaptively variable exponent combining Legendre polynomial for image segmentation
Jiajie Zhu 0001, Bin Fang 0001, Mingliang Zhou 0001, Futing Luo, Weizhi Xian, Gang Wang 0023
Multim. Tools Appl.2
2022 Simple and Robust Locality Preserving Projections Based on Maximum Difference Criterion
Ruisheng Ran, Shougui Zhang, Bin Fang 0001
Neural Process. Lett.4
2022 A General Matrix Function Dimensionality Reduction Framework and Extension for Manifold Learning
abstract
Many dimensionality reduction methods in the manifold learning field have the so-called small-sample-size (SSS) problem. Starting from solving the SSS problem, we first summarize the existing dimensionality reduction methods and construct a unified criterion function of these methods. Then, combining the unified criterion with the matrix function, we propose a general matrix function dimensionality reduction framework. This framework is configurable, that is, one can select suitable functions to construct such a matrix transformation framework, and then a series of new dimensionality reduction methods can be derived from this framework. In this article, we discuss how to choose suitable functions from two aspects: 1) solving the SSS problem and 2) improving pattern classification ability. As an extension, with the inverse hyperbolic tangent function and linear function, we propose a new matrix function dimensionality reduction framework. Compared with the existing methods to solve the SSS problem, these new methods can obtain better pattern classification ability and have less computational complexity. The experimental results on handwritten digit, letters databases, and two face databases show the superiority of the new methods.
Ruisheng Ran, Ji Feng, Shougui Zhang, Bin Fang 0001
IEEE Trans. Cybern.4
2021 Empirical Regularization for Synthetic Sentence Pairs in Unsupervised Neural Machine Translation
abstract
UNMT tackles translation on monolingual corpora in two required languages. Since there is no explicitly cross-lingual signal, pre-training and synthetic sentence pairs are significant to the success of UNMT. In this work, we empirically study the core training procedure of UNMT to analyze the synthetic sentence pairs obtained from back-translation. We introduce new losses to UNMT to regularize the synthetic sentence pairs by jointly training the UNMT objective and the regularization objective. Our comprehensive experiments support that our method can generally improve the performance of currently successful models on three similar pairs {French, German, Romanian} English and one dissimilar pair Russian English with acceptably additional cost.
Xi Ai, Bin Fang 0001
AAAI2
2021 Almost Free Semantic Draft for Neural Machine Translation
abstract
Translation quality can be improved by global information from the required target sentence because the decoder can understand both past and future information.However, the model needs additional cost to produce and consider such global information.In this work, to inject global information but also save cost, we present an efficient method to sample and consider a semantic draft as global information from semantic space for decoding with almost free of cost.Unlike other successful adaptations, we do not have to perform an EM-like process that repeatedly samples a possible semantic from the semantic space.Empirical experiments show that the presented method can achieve competitive performance in common language pairs with a clear advantage in inference efficiency.We will open all our source code on GitHub.
Xi Ai, Bin Fang 0001
NAACL-HLT2
2021 Cascaded Deeply Supervised Convolutional Networks for Liver Lesion Segmentation
abstract
Liver lesion segmentation from abdomen computed tomography (CT) with deep neural networks remains challenging due to the small volume and the unclear boundary. To effectively tackle these problems, in this paper, we propose a cascaded deeply supervised convolutional networks (CDS-Net). The cascaded deep supervision (CDS) mechanism uses auxiliary losses to construct a cascaded segmentation method in a single network, focusing the network attention on pixels that are more difficult to classify, so that the network can segment the lesion more effectively. CDS mechanism can be easily integrated into standard CNN models and it helps to increase the model sensitivity and prediction accuracy. Based on CDS mechanism, we propose a cascaded deep supervised ResUNet, which is an end-to-end liver lesion segmentation network. We conduct experiments on LiTS and 3DIRCADb dataset. Our method has achieved competitive results compared with other state-of-the-art ones.
Kaiyi Peng, Bin Fang 0001, Mingliang Zhou 0001
Int. J. Pattern Recognit. Artif. Intell.2
2021 A Fast Perceptual Surveillance Video Coding (PSVC) Based on Background Model-Driven JND Estimation
abstract
Perceptual video coding (PVC) optimization has been an important video coding technique, which can be consistent with the perception characteristics of the human visual system (HVS). Currently, PVC schemes incorporating the just noticeable distortion (JND) model can obtain better performance gain in all PVC schemes. To further accelerate the JND computation for real-time video coding applications (e.g. surveillance video coding and conference video coding), this paper proposes a fast perceptual surveillance video coding (PSVC) scheme based on background model-driven JND estimation method. First, to utilize the surveillance scene characteristics, the computation complexity of JND estimation can be significantly decreased by reusing the content complexity of background regions. Then we apply the perceptive video coding scheme into the background modeling-based surveillance video codec. The proposed scheme adopts background modeling frame as background anchor. Experimental results show that the proposed scheme can yield remarkable time saving of 42.33% maximum and on average 34.76% with approximate bitrate reductions and similar subjective quality, compared to HEVC and other state-of-the-art schemes.
Gang Wang 0023, Mingliang Zhou 0001, Haiheng Cao, Bin Fang 0001, Shiting Wen
Int. J. Pattern Recognit. Artif. Intell.4
2021 Surveillance Video Coding for Traffic Scene based on Vehicle Knowledge and Shared Library by Cloud-Edge Computing in Cyber-Physical-Social Systems
abstract
With rapid development of intelligent video surveillance systems based on cloud computing devices and edge computing devices in Cyber-Physical-Social Systems, massive surveillance video data has brings enormous challenge for video storage and transmission. However, existing surveillance video coding approaches hardly utilize intelligent video analysis results for improving video coding. This paper proposed a surveillance video coding scheme for traffic scene based on vehicle knowledge and shared library by cloud-edge computing in Cyber-Physical-Social Systems. Firstly, in order to provide the object library for synchronous application at the encode and decode side offline, a generation method of shared long-term foreground reference object library is proposed by using the existing large-scale monitoring vehicle object datasets. Then, to meet the requirement of low complexity and high-performance coding, a virtual foreground reference picture generation method with coding-oriented object retrieval is proposed. Experimental results show that the proposed scheme can obtain the satisfactory effect of the virtual foreground reference picture. Also, it can yield remarkable bit rate reductions, compared to HEVC.
Gang Wang 0023, Mingliang Zhou 0001, Bin Fang 0001, Shiting Wen
Int. J. Pattern Recognit. Artif. Intell.3
2021 Blood Vessel Segmentation Based on the 3D Residual U-Net
abstract
In this paper, we propose blood vessel segmentation based on the 3D residual U-Net method. First, we integrate the residual block structure into the 3D U-Net. By exploring the influence of adding residual blocks at different positions in the 3D U-Net, we establish a novel and effective 3D residual U-Net. In addition, to address the challenges of pixel imbalance, vessel boundary segmentation, and small vessel segmentation, we develop a new weighted Dice loss function with a better effect than the weighted cross-entropy loss function. When training the model, we adopted a two-stage method from coarse-to-fine. In the fine stage, a local segmentation method of 3D sliding window is added. In the model testing phase, we used the 3D fixed-point method. Furthermore, we employ the 3D morphological closed operation to smooth the surfaces of vessels and volume analysis to remove noise blocks. To verify the accuracy and stability of our method, we compare our method with FCN, 3D DenseNet, and 3D U-Net. The experimental results indicate that our method has higher accuracy and better stability than the other studied methods and that the average Dice coefficients for hepatic veins and portal veins reach 71.7% and 76.5% in the coarse stage and 72.5% and 77.2% in the fine stage, respectively. In order to verify the robustness of the model, we conducted the same comparative experiment on the brain vessel datasets, and the average Dice coefficient reached 87.2%.
Mulin Xin, Yi Wang 0074, Bin Fang 0001, Yongmei Xu, Chunhong Linghu
Int. J. Pattern Recognit. Artif. Intell.5
2021 Prototype-Based Discriminative Feature Representation for Class-incremental Cross-modal Retrieval
abstract
Cross-modal retrieval aims to retrieve the related items from various modalities with respect to a query from any type. The key challenge of cross-modal retrieval is to learn more discriminative representations between different category, as well as expand to an unseen class retrieval in the open world retrieval task. To tackle the above problem, in this paper, we propose a prototype learning-based discriminative feature learning (PLDFL) to learn more discriminative representations in a common space. First, we utilize a prototype learning algorithm to cluster these samples labeled with the same semantic class, by jointly taking into consideration the intra-class compactness and inter-class sparsity without discriminative treatments. Second, we use the weight-sharing strategy to model the correlations of cross-modal samples to narrow down the modality gap. Finally, we apply the prototype to achieve class-incremental learning to prove the robustness of our proposed approach. According to our experimental results, significant retrieval performance in terms of mAP can be achieved on average compared to several state-of-the-art approaches.
Shaoquan Zhu, Yong Feng 0002, Mingliang Zhou 0001, Baohua Qiang, Bin Fang 0001
Int. J. Pattern Recognit. Artif. Intell.5
2021 Active contour driven by adaptively weighted signed pressure force combined with Legendre polynomial for image segmentation
Bin Fang 0001, Mingliang Zhou 0001, Sam Kwong
Inf. Sci.2
2021 Rate Control Method Based on Deep Reinforcement Learning for Dynamic Video Sequences in HEVC
abstract
Rate control (RC) plays a critical role in the transmission of high-quality video data under certain bandwidth restrictions in High Efficiency Video Coding (HEVC). Most current HEVC RC algorithms based on spatio-temporal information for rate-distortion (R-D) model parameters cannot effectively handle the cases with dynamic video sequences that contain fast moving objects, significant object occlusion or scene changes. In this paper, we propose an RC method based on deep reinforcement learning (DRL) for dynamic video sequences in HEVC to improve the coding efficiency. First, the rate control problem is formulated as a Markov decision process (MDP) problem. Second, with the MDP model, we develop a DRL-based algorithm to find the optimal quantization parameters (QPs) by training a deep neural network. The resulting intelligent agent selects the optimal RC strategy to reduce distortion, buffer and quality fluctuations by observing the current state of the encoder. The asynchronous advantage actor-critic (A3C) method is used to solve the MDP problem. Finally, the proposed DRL-based RC method is implemented in the newest video coding standard. Experimental results show that the proposed method offers substantially enhanced RC accuracy and consistently outperforms HEVC reference software and other state-of-the-art algorithms.
Mingliang Zhou 0001, Xuekai Wei, Sam Kwong, Weijia Jia 0001, Bin Fang 0001
IEEE Trans. Multim.5
2020 Hybrid Active Contour Driven by Double-Weighted Signed Pressure Force for Image Segmentation
abstract
In this paper, we proposed a novel hybrid active contour driven by double-weighted signed pressure force method for image segmentation. First, the Legendre polynomials and global information are integrated into the signed pressure force (SPF) function and a coefficient is applied to weight the effect degrees of the Legendre term and global term. Second, by introducing a weighted factor as the coefficient of inside and outside region fitting center, the curve can be optimally evolved to the interior and branches of the region of interest (ROI). Third, a new edge stopping function is adopted to robustly capture the edge of ROI and speed up the multi-object image segmentation. Experiments show that the proposed method can achieve better accuracy for images with noise, inhomogeneous intensity, blur edge and complex branches, in the meanwhile, it also controls the time-consuming effectively and is insensitive to the initial contour position.
Bin Fang 0001, Mingliang Zhou 0001
ICASSP2
2020 Segmentation Algorithm of the Valid Region in Fisheye Images Using Edge and Region Information
abstract
In this paper, we propose a method to segment the valid region of fisheye images. First, we construct an objective function with three terms, which are the region driving term, the edge driving term and the length regularization term. Second, we minimize this objective function by a modified gradient descent method to find the best segmentation result. Our method can achieve valid region segmentation by making use of both region information and edge information. Experiments show that the proposed method can deal with blurred edges, halation noise and incomplete valid region problems.
Tongxin Du, Bin Fang 0001, Mingliang Zhou 0001, Henjun Zhao, Weizhi Xian, Xuegang Wu
ICIP2
2020 Legendre Based Adaptive Image Segmentation Combining The Gradient Information
abstract
In this paper, we propose an adaptive variable exponent level set method based on Legendre polynomials for object segmentation in complex visual environment. First, we use a set of Legendre basis functions to approximate the region intensity, which enable us to accommodate heterogeneous objects. Second, an improved function is presented to update exponent adaptively and ensures the image gradient information embedding into the model easily. The proposed method is robust to low contrast, blurred boundaries, noise and the 10-cation of initial contour, and sufficient in handling large scale intensity variations. Experimental results demonstrate that the proposed method can achieve relatively high segmentation accuracy and less computational time.
Jiajie Zhu 0001, Bin Fang 0001, Mingliang Zhou 0001, Hengjun Zhao, Futing Luo
ICIP2
2020 MEMS Inertial Sensor Fault Diagnosis Using a CNN-Based Data-Driven Method
abstract
In this paper, we propose a novel fault diagnosis (FD) approach for micro-electromechanical systems (MEMS) inertial sensors that recognize the fault patterns of MEMS inertial sensors in an end-to-end manner. We use a convolutional neural network (CNN)-based data-driven method to classify the temperature-related sensor faults in unmanned aerial vehicles (UAVs). First, we formulate the FD problem for MEMS inertial sensors into a deep learning framework. Second, we design a multi-scale CNN which uses the raw data of MEMS inertial sensors as input and which outputs classification results indicating faults. Then we extract fault features in the temperature domain to solve the non-uniform sampling problem. Finally, we propose an improved adaptive learning rate optimization method which accelerates the loss convergence by using the Kalman filter (KF) to train the network efficiently with a small dataset. Our experimental results show that our method achieved high fault recognition accuracy and that our proposed adaptive learning rate method improved performance in terms of loss convergence and robustness on a small training batch.
Wei Sheng, Mingliang Zhou 0001, Bin Fang 0001, Liping Zheng
Int. J. Pattern Recognit. Artif. Intell.4
2020 DSRPH: Deep semantic-aware ranking preserving hashing for efficient multi-label image retrieval
Yong Feng 0002, Bin Fang 0001, Mingliang Zhou 0001, Sam Kwong, Baohua Qiang
Inf. Sci.3
2020 Just Noticeable Distortion-Based Perceptual Rate Control in HEVC
abstract
In this paper, we propose a just noticeable distortion (JND)-based perceptual rate control method for high efficiency video coding (HEVC). First, the JND factor of a coding unit has been mathematically shown to be an approximation of the average pixel-level JND weight, which means that it can also be used as a weight for bitrate allocation. Second, rate-distortion (R-D) modelling is conducted based on the JND factor. Finally, the proposed R-D model is integrated into an existing rate control framework to improve the coding efficiency, and the proposed algorithm is implemented in the newest video coding standard. As the experimental results reveal, compared with HEVC reference software, our algorithm achieves significantly improved coding performance, subjective coding quality and bitrate accuracy.
Mingliang Zhou 0001, Xuekai Wei, Sam Kwong, Weijia Jia 0001, Bin Fang 0001
IEEE Trans. Image Process.5
2019 Liver Vessels Segmentation Based on 3d Residual U-NET
abstract
Recently, extraction of blood vessels has aroused widespread interests in medical image analysis. In this work, to accelerate convergence speed and enhance the representation for discriminative features, we introduce the residual block structure in the ResNet into the 3D U-Net, and construct a new 3D Residual U-Net architect to segment the hepatic and portal veins from abdominal CT volumes. In addition, we develop a weighted Dice loss function to cope with the challenges of pixel imbalance, vessel boundary segmentation and small vessels segmentation. Furthermore, based on the prediction results, the post-processing methods of 3D morphological closed operation and volume analysis are employed to smooth the surface of vessels and eliminate noise blocks, respectively. Compared with existing 3D DenseNet, FCN and 3D U-Net, the average Dice coefficients of our method in hepatic veins and portal veins segmentation are 71.7% and 76.5% respectively, which are superior to 55.3% and 53.9% of the 3D DenseNet, 60.2% and 75.6% of the FCN, and 66.4% and 73.9% of the 3D U-Net. Meanwhile, the cross validation results prove that our method is accurate and stable for liver vessel extraction.
Bin Fang 0001, Mingqi Gao 0001, Shenhai Zheng, Yi Wang 0074
ICIP2
2019 Human action recognition based on spatio-temporal three-dimensional scattering transform descriptor and an improved VLAD feature encoding algorithm
Bin Fang 0001, Weibin Yang, Jiye Qian
Neurocomputing2
2019 Feature fusion and non-negative matrix factorization based active contours for texture segmentation
Mingqi Gao 0001, Hengxin Chen, Shenhai Zheng, Bin Fang 0001
Signal Process.4
2018 A BasisEvolution framework for network traffic anomaly detection
Bin Fang 0001, Matthew Roughan, Kenjiro Cho, Paul Tune
Comput. Networks2
2018 B-Spline based globally optimal segmentation combining low-level and high-level information
Shenhai Zheng, Bin Fang 0001, Laquan Li, Mingqi Gao 0001, Kaiyi Peng
Pattern Recognit.2
2018 Tackling class overlap and imbalance problems in software defect prediction
Lin Chen 0023, Bin Fang 0001, Zhaowei Shang, Yuan Yan Tang
Softw. Qual. J.2
2018 Fast and Accurate Vanishing Point Detection and Its Application in Inverse Perspective Mapping of Structured Road
abstract
Fast and accurate visual scene understanding in autonomous vehicles is necessary but still very challenging. An autonomous vehicle must be taught to read the road like a human driver for better controlling the vehicle, so it is important to efficiently detect the road area and road markings. In this paper, we mainly focus on the vanishing point detection and its application in inverse perspective mapping (IPM) for road marking understanding. We first propose a fast and accurate vanishing point detection method for various types of roads, by adopting and improving Weber local descriptor to obtain salient representative texture and orientation information of the road area, and then voting for the dominant vanishing point with a simple line-voting scheme. Experimental results demonstrate that the proposed vanishing point detection approach gains a better performance than some state-of-the-art methods in terms of accuracy and computation time. Furthermore, we introduce the detected vanishing point into the IPM algorithm in the structured road environment, since some important calibration parameters can be automatically calculated by the vanishing point, especially on the rough road. Experiments also show that our proposed vanishing point-based IPM method is adaptive and accurate, which is conducive to the subsequent road marking detection and recognition.
Weibin Yang, Bin Fang 0001, Yuan Yan Tang
IEEE Trans. Syst. Man Cybern. Syst.2
2017 Chinese Handwriting Identification Method Based on Keyword Extraction
abstract
Text-independent handwriting identification methods require that features such as texture are extracted from lengthy document image; while text-dependent handwriting identification methods require that the contents of the documents being compared are identical. In order to overcome these confinements, this paper presents a novel Chinese handwriting identification technique. First, Chinese characters are segmented from handwriting document, then keywords are extracted based on matching and voting of local features of character. Then the same-content keywords are used to build training sets, and these training sets of two documents are compared. Because the keywords are similar to signature, the handwriting identification problem is transformed into signature verification problem. Experiments on HIT-MW, HIT-SW and CASIA show this method outperforms many text-independent handwriting identification methods.
Bin Fang 0001, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.2
2017 A Novel Race Classification Method Based on Periocular Features Fusion
abstract
Race identification is an essential ability for human eyes. Race classification by machine based on face image can be used in some practical application fields. Employing holistic face analysis, local feature extraction and 3D model, many race classification methods have been introduced. In this paper, we propose a novel fusion feature based on periocular region features for classifying East Asian from Caucasian. With the periocular region landmarks, we extract five local textures or geometrical features in some interesting regions which contain available discriminating race information. And then, these effective features are fused into a remarkable feature by Adaboost training. On the composed OFD-FERET face database, our method gets perfect performance on average accuracy rate. Meanwhile, we do a plenty of additional experiments to discuss the effect on the performance caused by gender, landmark detection, glasses and image size.
Hengxin Chen, Mingqi Gao 0001, Karl Ricanek, Weiliang Xu 0003, Bin Fang 0001
Int. J. Pattern Recognit. Artif. Intell.5
2017 Off-Line Signature Verification Based on Multi-Scale Local Structural Pattern
abstract
The paper proposes a theoretically very simple, and effective multi-scale local structural patterns (MS-LSP) for off-line signature verification. The proposed strategy models the signature by probabilistically counting the distribution of each local structural pattern on multiple scales. This representation correlates with the writing style of an individual and depict the shape and structural information of signatures. Extensive experiments on two challenging databases (MCYT, GPDS960) show that the proposed approach achieves noticeably verification rates and it is expected to provide a powerful discriminative representation of the signature.
Bin Fang 0001
Int. J. Pattern Recognit. Artif. Intell.2
2016 A factorization based active contour model for texture segmentation
abstract
This paper presents a factorization based active contour model for 2-phase texture segmentation. We utilize the local spectral histogram as the texture features, and then establish a novel energy function based on the theory of the matrix decomposition. Unlike the existing methods, we only choose the combination weights from object region and background region to handle the motion of curve. We compare the proposed method to the recently active contour methods and the experiments are performed on synthetic and the real-world images. The experimental results show that our model is more robust against the complex background than the other strategies.
Mingqi Gao 0001, Hengxin Chen, Shenhai Zheng, Bin Fang 0001
ICIP4
2016 Multi-scale B-spline level set segmentation based on Gaussian kernel equalization
abstract
Images with weak contrast, overlapped noise and texture of the object and background make many PDE based methods disabled. To address these problems, this paper presents a novel combined multi-scale variational framework level set segmentation model. Its level set formulation consists edge-based term, region-based term and shape constraint term. The edge-based term is constructed using a newly defined edge stopping function. The region-based term is derived from parameter-free Gaussian probability density function (pdf) and multiple Gaussian kernel are used to gray equalization. The shape constraint term is used to constrain contour evolution at different scales of image pyramid. For an intrinsic smoothing segmentation contours, the level set function is explicitly represented by B-spline basis functions. Finally, a convolution is used during the energy minimization. Experimental results on synthetic and real images validate the robustness and high accuracy boundaries detection for low contrast, noise and texture images.
Shenhai Zheng, Bin Fang 0001, Patrick Shen-Pei Wang, Laquan Li, Mingqi Gao 0001
ICIP2
2016 Texture image segmentation using fused features and active contour
abstract
This paper introduces an effective active contour model for texture segmentation. To improve the robustness against noise and illumination, a novel descriptor named local statistical variation degree (LSVD) is presented to express textural features, which uses corner point deletion and isolated region detection operations to eliminate image patches unrelated with object regions. And then the fused features combined LSVD with Gabor can be constructed to express image structure in many scene. During the texture segmentation stage, a factorization based fitting energy is proposed to measure the weights of representative features in the features computed from image regions. This fitting energy can be used to localize region boundary more accurately. Moreover, a boundary shrinking method is put forward to improve the reliability of representative features. By comparing our proposed method with the recent texture segmentation models on synthetic images and natural images, we demonstrate that the novel active contour model can obtain accurate segmentation results and is robust to noise, illumination and position of initial contour.
Mingqi Gao 0001, Hengxin Chen, Shenhai Zheng, Bin Fang 0001
ICPR4
2015 NNMap: A method to construct a good embedding for nearest neighbor classification
Jing Chen 0008, Yuan Yan Tang, C. L. Philip Chen, Bin Fang 0001, Zhaowei Shang, Yuewei Lin
Neurocomputing4
2015 Negative samples reduction in cross-company software defects prediction
Lin Chen 0023, Bin Fang 0001, Zhaowei Shang, Yuan Yan Tang
Inf. Softw. Technol.2
2015 A novel item anomaly detection approach against shilling attacks in collaborative recommendation systems using the dynamic time interval segmentation technique
Bin Fang 0001, Yuan Yan Tang
Inf. Sci.2
2015 3D CAD model retrieval based on the combination of features
Bin Fang 0001, Yong-Mei Yu
Multim. Tools Appl.2
2014 Automatic decision support by information energy decision tree algorithm
abstract
The application of information entropy to decision tree algorithms has been shown to produce very accurate classifiers. Information entropy is utilized to ensure that the average distance of paths from the non-leaf node to each descendant leaf node of the decision tree is shortest. Therefore, it works well for data set which covers all the underlying rules. But it is lack of prediction ability when the training data set can not cover all the underlying rules. In this paper, we propose a novel indicator, information energy, to generate decision tree. Information energy describes the distance from the current state of a data set to its balance state. Proper selection of attribute can divide a data set into a state of higher information energy and produce classification rules of prediction ability. A generator of random sample sets and rules is designed to provide synthetic samples for experimental verification. Experimental results show that information energy outperforms information entropy in both speed and accuracy when the training data set can not cover all the underlying rules.
Runzong Liu, Yuan Yan Tang, Bin Fang 0001
SMC3
2014 An enhanced version and an incremental learning version of visual-attention-imitation convex hull algorithm
Runzong Liu, Yuan Yan Tang, Bin Fang 0001, Jingrui Pi
Neurocomputing3
2014 An approximate closed-form solution to correlation similarity discriminant analysis
Taiping Zhang, Yuan Yan Tang, C. L. Philip Chen, Zhaowei Shang, Bin Fang 0001
Neurocomputing5
2014 Extracting sparse error of robust PCA for face recognition in the presence of varying illumination and occlusion
Xiao Luan, Bin Fang 0001, Linghui Liu, Weibin Yang, Jiye Qian
Pattern Recognit.2
2014 Topological Coding and Its Application in the Refinement of SIFT
abstract
Point pattern matching plays a prominent role in the fields of computer vision and pattern recognition. A technique combining the circular onion peeling and the radial decomposition is proposed to analyze the topology structure of a point pattern. The analysis derives a feature which records the topological structure of a point pattern. This novel feature is free from isometric assumption. It can resist various deformations such as adding points, suppressing points, affine transformations, projective transformations and elastic transformations to some degree. A refinement solution of the well known scale invariant feature transform (SIFT) algorithm is also proposed based on the probabilistic analysis of this feature. Experimental results show that the proposed refinement solution for SIFT using this feature is effective and robust.
Runzong Liu, Yuan Yan Tang, Bin Fang 0001
IEEE Trans. Cybern.3
2013 A no-reference image sharpness estimation based on expectation of wavelet transform coefficients
abstract
In this work, the expectation of wavelet transform coefficients is used for estimating an image sharpness. It's based on the observation that the greater the probability of big detail coefficients, the more pixels appear sharply, and consequently, the sharper the image. Specifically, an input image is firstly decomposed into three directional sub-bands by a separable discrete wavelet transform. Then these directional sub-bands are viewed as three random variables, and their expectations are computed. Finally, The proposed sharpness index is the weighted sum of three expectations. The experiments show that, despite its simplicity, the proposed sharpness index is competitive with the current best-performance techniques for no-reference image sharpness estimation.
Hengjun Zhao, Bin Fang 0001, Yuan Yan Tang
ICIP2
2013 Face recognition with contiguous occlusion using linear regression and level set method
Xiao Luan, Bin Fang 0001, Linghui Liu, Lifang Zhou
Neurocomputing2
2013 Visual saliency detection with center shift
Weibin Yang, Yuan Yan Tang, Bin Fang 0001, Zhaowei Shang, Yuewei Lin
Neurocomputing3
2013 Learning orthogonal projections for Isomap
Bin Fang 0001, Yuan Yan Tang, Taiping Zhang, Ruizong Liu
Neurocomputing2
2013 Automatic Multi-Scale Segmentation of Intrahepatic Vessel in CT Images for Liver Surgery Planning
abstract
The processing of blood vessels is an indispensable part in complicated surgeries of livers and hearts as the development of medical image technologies, which requires an automatic segmentation system over CT images of organs. However, the vascular pattern of livers in CT images suffers from low contrast to background so that the existing segmentation technologies are not able to extract the blood vessels completely. In the paper, we propose a new algorithm to extract the blood vessels of livers based on the adaptive multi-scale segmentation. First, we prove that the background histogram of normal scale blood vessels obeys the Gaussian distribution in CT images, and obtain the vascular distribution function from the vascular signal segmented from the background with a local optimal threshold. Second, Hessian matrix is employed to enhance the thin blood vessels before the extraction, and a complete and clear segmentation system for blood vessels is constructed by combining the major and thin blood vessels via filtering. Experimental results show the effectiveness of the proposed method, which is able to extract more complete blood vessels for 3D system, and assist the clinical liver surgeries efficiently.
Yi Wang 0074, Bin Fang 0001, Jingrui Pi, Patrick Shen-Pei Wang, Hongguang Wang
Int. J. Pattern Recognit. Artif. Intell.2
2013 A Fast and Complete Convex-Hull Algorithm Architecture Based on Ellipse and Elastic Ellipse Methods
abstract
The number of inner points excluded in an initial convex hull (ICH) is vital to the efficiency getting the convex hull (CH) in a planar point set. The maximum inscribed circle method proposed recently is effective to remove inner points in ICH. However, limited by density distribution of a planar point set, it does not always work well. Although the affine transformation method can be used, it is still hard to have a better performance. Furthermore, the algorithm mentioned above fails to deal with the exceptional distribution: the gravity centroid (GC) of a planar point set is outside or on the edge formed by the extreme points in ICH. This paper considers how to remove more inner points in ICH when GC is inside of ICH and completely process the case which mentioned above. Further, we presented a complete algorithm architecture: (1) using the ellipse and elasticity ellipse methods (EM and EEM) to remove more inner points in ICH and process the cases: GC is inside or outside of ICH. (2) Using the traditional methods to process the situation: the initial centroid is on the edge in ICH. It is adaptive to more data sets than other algorithms. The experiments under seven distributions show that the proposed method performs better than other traditional algorithms in saving time and space.
Xuegang Wu, Bin Fang 0001, Yuan Yan Tang, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.2
2013 An Adaptive Fuzzy Fusion Framework for Face Recognition under Illumination variation Based on Local Multiple Patterns
abstract
Local binary pattern (LBP) operator offers an efficient way to recognize face under varying illumination, while it has the drawback of abandoning some important texture features. Local multiple patterns (LMP) has alleviated the problem by a hierarchical model. However, the LMP method can bring out the rapid expansion of feature dimension, so a special feature encoding method is adopted by this paper. Meanwhile, we find that the LMP features of different layers can be used to recognize face independently so that it would preserve more abundant recognition information. Most importantly, the contribution of the LMP features from different layers is blurry under varying illumination. We propose a fuzzy framework to fuse the recognition result of different layers and use adaptive weights to calculate contribution rates of different layers under varying illumination. Experimental results demonstrate that the proposed method outperforms other state-of-the-art methods on four databases such as Yale B, Extended Yale B, CMU PIE and Outdoor.
Lifang Zhou, Bin Fang 0001, Weisheng Li 0001, Hengxin Chen, Lidou Wang
Int. J. Pattern Recognit. Artif. Intell.2
2013 A Visual-Attention Model Using Earth Mover's Distance-Based Saliency Measurement and Nonlinear Feature Combination
abstract
This paper introduces a new computational visual-attention model for static and dynamic saliency maps. First, we use the Earth Mover's Distance (EMD) to measure the center-surround difference in the receptive field, instead of using the Difference-of-Gaussian filter that is widely used in many previous visual-attention models. Second, we propose to take two steps of biologically inspired nonlinear operations for combining different features: combining subsets of basic features into a set of super features using the Lm-norm and then combining the super features using the Winner-Take-All mechanism. Third, we extend the proposed model to construct dynamic saliency maps from videos by using EMD for computing the center-surround difference in the spatiotemporal receptive field. We evaluate the performance of the proposed model on both static image data and video data. Comparison results show that the proposed model outperforms several existing models under a unified evaluation setting.
Yuewei Lin, Yuan Yan Tang, Bin Fang 0001, Zhaowei Shang, Song Wang 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2013 Large Margin Subspace Learning for feature selection
Bin Fang 0001, Xinwang Liu 0002, Jie Chen 0001, Zhenghong Huang, Xiping He
Pattern Recognit.2
2013 Multi-focus image fusion based on the neighbor distance
Hengjun Zhao, Zhaowei Shang, Yuan Yan Tang, Bin Fang 0001
Pattern Recognit.4
2012 Visual saliency estimation using support value transform
abstract
This paper proposes a novel method for estimating visual saliency based on a typical agreement that image saliency depends mainly on local and global contrast from various feature channels. We compute the contrast between image patches on different low-level feature maps which are generated by color space conversion and support value transform. To obtain the representative measurement effectively, we calculate the dissimilarity in a reduced dimensional principal component space. In addition, our method may be easily extended for more conspicuous feature channels in an efficient manner. Experimental results on two public available human eye fixation datasets demonstrate that our method outperforms other seven state-of-the-art saliency models.
Weibin Yang, Bin Fang 0001, Yuan Yan Tang, Zhaowei Shang, Hengjun Zhao
ICIP2
2012 Orthogonal Isometric Projection
Yuan Yan Tang, Bin Fang 0001, Taiping Zhang
ICPR3
2012 A fast convex hull algorithm with maximum inscribed circle affine transformation
Runzong Liu, Bin Fang 0001, Yuan Yan Tang, Jiye Qian
Neurocomputing2
2012 Fragmented edge structure coding for Chinese writer identification
Bin Fang 0001, Junlin Chen, Yuan Yan Tang, Hengxin Chen
Neurocomputing2
2012 Multi-Scale Gradient Invariant for Face Recognition under varying Illumination
abstract
In this paper, a novel approach derived from image gradient domain called multi-scale gradient faces (MGF) is proposed to abstract multi-scale illumination-insensitive measure for face recognition. MGF applies multi-scale analysis on image gradient information, which can discover underlying inherent structure in images and keep the details at most while removing varying lighting. The proposed approach provides state-of-the-art performance on Extended YaleB and PIE: Recognition rates of 99.11% achieved on PIE database and 99.38% achieved on YaleB which outperforms most existing approaches. Furthermore, the experimental results on noised Yale-B validate that MGF is more robust to image noise.
Yuan Yan Tang, Bin Fang 0001, Zhaowei Shang
Int. J. Pattern Recognit. Artif. Intell.3
2012 Document Clustering in Correlation Similarity Measure Space
abstract
This paper presents a new spectral clustering method called correlation preserving indexing (CPI), which is performed in the correlation similarity measure space. In this framework, the documents are projected into a low-dimensional semantic space in which the correlations between the documents in the local patches are maximized while the correlations between the documents outside these patches are minimized simultaneously. Since the intrinsic geometrical structure of the document space is often embedded in the similarities between the documents, correlation as a similarity measure is more suitable for detecting the intrinsic geometrical structure of the document space than euclidean distance. Consequently, the proposed CPI method can effectively discover the intrinsic structures embedded in high-dimensional document space. The effectiveness of the new method is demonstrated by extensive experiments conducted on various data sets and by comparison with existing document clustering methods.
Taiping Zhang, Yuan Yan Tang, Bin Fang 0001, Yong Xiang 0001
IEEE Trans. Knowl. Data Eng.3
2011 Kernel-view based discriminant approach for embedded feature extraction in high-dimensional space
Bin Fang 0001, Chi-Man Pun, Yuan Yan Tang
Neurocomputing2
2011 A Multi-Layer Contrast Analysis Method for Texture Classification Based on LBP
abstract
Texture classification is one of the important fields in pattern recognition and machine vision research. LBP method,13–15 proposed by Ojala, can be used to classify texture images effectively. And the LBP method has rotation-invariant, illumination-invariant, multi-resolution characteristics. But, since the contrast is not considered between neighbor pixels, the correct classification rate produced by this method has been remarkably influenced by light source type and light source orientation. The LMLCP (Local Multiple Layer Contrast Pattern) method, proposed by this paper, maps the contrast value between two near pixels to a rank value, which represent a relative contrast value range, and computes the statistic histogram referring to the work in LBP method. The LMLCP method can bring out the rapid expansion of feature dimension, so a special feature encoding method used in 3DLBP6 is adopted by this paper. The experiment, which is built based on Outex_TC_00012,12 demonstrates that the LMLCP can evidently make a more accurate classification rate than LBP method.
Hengxin Chen, Yuan Yan Tang, Bin Fang 0001, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.3
2011 Illumination Invariant Face Recognition Using Fabemd Decomposition with Detail Measure Weight
abstract
With varying illumination conditions, facial features obtained from images are distorted nonlinearly by variant lighting intensity and direction, so face recognition becomes very difficult. According to the "common assumption", illumination varies slowly and the face intrinsic feature (including 3D surface and reflectance) varies rapidly in local area, we can then consider high frequency features that represent the face intrinsic structure. FABEMD8 (Fast and Adaptive Bidimensional Empirical Mode Decomposition) is a fast and adaptive method of BEMD22 (Bidimensional Empirical Mode Decomposition), and not using time-consuming plane interpolation computation, it can decompose the image into multilayer high frequency images representing detail features and low frequency images representing analogy features. But we cannot make a quantitative analysis of how many detail features can be used to eliminate illumination variation. So we propose two measures to quantify the detail features, and with these measure weights, we can activitate FABEMD based multilayer detail images matching for face recognition under varying illumination. With PCA, the experiments based on Yale face database B and MU PIE face database show that the method proposed in this paper can get remarkable performance.
Hengxin Chen, Yuan Yan Tang, Bin Fang 0001
Int. J. Pattern Recognit. Artif. Intell.3
2011 A Computational and Theoretical Analysis of Local Null Space Discriminant Method for Pattern Classification
abstract
Many problems in pattern classification and feature extraction involve dimensionality reduction as a necessary processing. Traditional manifold learning algorithms, such as ISOMAP, LLE, and Laplacian Eigenmap, seek the low-dimensional manifold in an unsupervised way, while the local discriminant analysis methods identify the underlying supervised submanifold structures. In addition, it has been well-known that the intraclass null subspace contains the most discriminative information if the original data exist in a high-dimensional space. In this paper, we seek for the local null space in accordance with the null space LDA (NLDA) approach and reveal that its computational expense mainly depends on the quantity of connected edges in graphs, which may be still unacceptable if a great deal of samples are involved. To address this limitation, an improved local null space algorithm is proposed to employ the penalty subspace to approximate the local discriminant subspace. Compared with the traditional approach, the proposed method can achieve more efficiency so that the overload problem is avoided, while slight discriminant power is lost theoretically. A comparative study on classification shows that the performance of the approximative algorithm is quite close to the genuine one.
Bin Fang 0001, Yuan Yan Tang, Hengxin Chen
Int. J. Pattern Recognit. Artif. Intell.2
2011 Bionic Face Recognition Using Gabor Transformation
abstract
In this paper, we propose a bionic face recognition method based on Gabor feature. First, Gabor features are extracted from face images, followed by dimensionality reduction using 2DPCA algorithm, which serves as the feature vectors of the proposed method. Finally, the bionic classifier is trained for classification. The experiment on AR and PIE face database is reported to show the effectiveness of the proposed method and compare it with Gabor-2DPCA algorithm and Gabor-PCA algorithm.
Yuan Yan Tang, Bin Fang 0001, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.3
2010 A Computational Model for Saliency Maps by Using Local Entropy
abstract
This paper presents a computational framework for saliency maps. It employs the Earth Mover's Distance based on weighted-Histogram (EMD-wH) to measure the center-surround difference, instead of the Difference-of-Gaussian (DoG) filter used by traditional models. In addition, the model employs not only the traditional features such as colors, intensity and orientation but also the local entropy which expresses the local complexity. The major advantage of combining the local entropy map is that it can detect the salient regions which are not complex regions. Also, it uses a general framework to integrate the feature dimensions instead of summing the features directly. This model considers both local and global salient information, in contrast to the existing models that consider only one or the other. Furthermore, the "large scale bias" and "central bias" hypotheses are used in this model to select the fixation locations in the saliency map of different scales. The performance of this model is assessed by comparing their saliency maps and human fixation density. The results from this model are finally compared to those from other bottom-up models for reference.
Yuewei Lin, Bin Fang 0001, Yuan Yan Tang
AAAI2
2010 A Least-Squares Model to Orthogonal Linear Discriminant Analysis
abstract
Orthogonal transformation can delete the correlations among candidate features such that the extracted features do not disturb each other. An orthogonal set of discriminant vectors is more powerful than the classical discriminant vectors. In this paper, we present a new orthogonal linear discriminant analysis (OLDA) model based on least-squares approximation called LS-OLDA for pattern classification, which aims to find an orthogonal transformation W and a diagonal matrix D such that the difference between [Formula: see text] and WDWT is minimized in the least-squares sense, and the trace of D is maximized simultaneously. Theoretical analysis shows that the proposed model coincides with classical OLDA criterion. The experimental results on different standard data sets compared with related methods show that LS-OLDA achieves or approximates closely to the best accuracy, and has lower computational cost.
Taiping Zhang, Bin Fang 0001, Yuan Yan Tang, Zhaowei Shang
Int. J. Pattern Recognit. Artif. Intell.2
2010 Marginal discriminant projections: An adaptable margin discriminant approach to feature reduction and extraction
Bin Fang 0001, Yuan Yan Tang
Pattern Recognit. Lett.2
2010 Incremental Embedding and Learning in the Local Discriminant Subspace With Application to Face Recognition
abstract
Dimensionality reduction and incremental learning have recently received broad attention in many applications of data mining, pattern recognition, and information retrieval. Inspired by the concept of manifold learning, many discriminant embedding techniques have been introduced to seek low-dimensional discriminative manifold structure in the high-dimensional space for feature reduction and classification. However, such graph-embedding framework-based subspace methods usually confront two limitations: (1) since there is no available updating rule for local discriminant analysis with the additive data, it is difficult to design incremental learning algorithm and (2) the small sample size (SSS) problem usually occurs if the original data exist in very high-dimensional space. To overcome these problems, this paper devises a supervised learning method, called local discriminant subspace embedding (LDSE), to extract discriminative features. Then, the incremental-mode algorithm, incremental LDSE (ILDSE), is proposed to learn the local discriminant subspace with the newly inserted data, which applies incremental learning extension to the batch LDSE algorithm by employing the idea of singular value-decomposition (SVD) updating algorithm. Furthermore, the SSS problem is avoided in our method for the high-dimensional data and the benchmark incremental learning experiments on face recognition show that ILDSE bears much less computational cost compared with the batch algorithm.
Bin Fang 0001, Yuan Yan Tang, Taiping Zhang
IEEE Trans. Syst. Man Cybern. Part C2
2010 Generalized Discriminant Analysis: A Matrix Exponential Approach
abstract
Linear discriminant analysis (LDA) is well known as a powerful tool for discriminant analysis. In the case of a small training data set, however, it cannot directly be applied to high-dimensional data. This case is the so-called small-sample-size or undersampled problem. In this paper, we propose an exponential discriminant analysis (EDA) technique to overcome the undersampled problem. The advantages of EDA are that, compared with principal component analysis (PCA) + LDA, the EDA method can extract the most discriminant information that was contained in the null space of a within-class scatter matrix, and compared with another LDA extension, i.e., null-space LDA (NLDA), the discriminant information that was contained in the non-null space of the within-class scatter matrix is not discarded. Furthermore, EDA is equivalent to transforming original data into a new space by distance diffusion mapping, and then, LDA is applied in such a new space. As a result of diffusion mapping, the margin between different classes is enlarged, which is helpful in improving classification accuracy. Comparisons of experimental results on different data sets are given with respect to existing LDA extensions, including PCA + LDA, LDA via generalized singular value decomposition, regularized LDA, NLDA, and LDA via QR decomposition, which demonstrate the effectiveness of the proposed EDA method.
Taiping Zhang, Bin Fang 0001, Yuan Yan Tang, Zhaowei Shang
IEEE Trans. Syst. Man Cybern. Part B2
2009 Weightiness image Partition in 3D Face Recognition
abstract
In this paper we present a novel algorithm suitable to improve the accuracy of 3D face recognition. In the proposed algorithm, we represent the 3D points by point signatures and partition the facial data into fifteen regions according to ¿three courtyards and five eyes¿ theory in pencil sketch on facial image in Chinese traditional art. Then in each partition we use ICA getting eigenvalues of feature and structure character and depth information to represent the 3D facial data. We assign different weightiness to each sub-image according to the result of sub-image variety. In order to match incomplete data under structural constraints, we proposed a reformative robust structural Hausdorff distance to handle these possible cases. Experiments on FRGC v2.0 data set show that the proposed algorithm is robust and effective to 3D face with expression, lighting and expression variance.
Yuan Yan Tang, Bin Fang 0001, Taiping Zhang
SMC3
2009 False positive reduction in urinary particle recognition
Bin Fang 0001, Jiye Qian, Lin Chen 0023, Ying Liu 0004
Expert Syst. Appl.2
2009 Improving the discriminant ability of local margin based learning method by incorporating the global between-class separability criterion
Bin Fang 0001, Yuan Yan Tang
Neurocomputing1
2009 Combining Eodh and Directional Gradient Density for Offline Signature Verification
abstract
The main problem to identify skilled forgeries for offline signature verification lies in the fact that it is difficult to formalize distinguished feature representation of the signature patterns and design appropriate fusion scheme for various types of feature vectors. To tackle these problems, in this paper, we propose an approach to extract robust Edge Orientation Distance Histogram (EODH) descriptor which effectively reflects signature structure variations. In addition, directional gradient density features are employed for skilled forgery verification attempt. To exploit the full capacity of two sets of features, we designed the multilevel weighted fuzzy classifier and fuse match scores by way of selection priority. Experiments were conducted on a subcorpus of open MCYT signature database which is widely used for performance evaluation. It shows that the proposed method was able to improve verification accuracy.
Bin Fang 0001, Yuan Yan Tang, Patrick Shen-Pei Wang, Taiping Zhang
Int. J. Pattern Recognit. Artif. Intell.2
2009 Model-based signature verification with rotation invariant features
Bin Fang 0001, Yuan Yan Tang, Taiping Zhang
Pattern Recognit.2
2009 Multiscale facial structure representation for face recognition under varying illumination
Taiping Zhang, Bin Fang 0001, Yuan Yuan 0001, Yuan Yan Tang, Zhaowei Shang, Fangnian Lang
Pattern Recognit.2
2009 Face Recognition Under Varying Illumination Using Gradientfaces
abstract
In this correspondence, we propose a novel method to extract illumination insensitive features for face recognition under varying lighting called the Gradientfaces. Theoretical analysis shows Gradientfaces is an illumination insensitive measure, and robust to different illumination, including uncontrolled, natural lighting. In addition, Gradientfaces is derived from the image gradient domain such that it can discover underlying inherent structure of face images since the gradient domain explicitly considers the relationships between neighboring pixel points. Therefore, Gradientfaces has more discriminating power than the illumination insensitive measure extracted from the pixel domain. Recognition rates of 99.83% achieved on PIE database of 68 subjects, 98.96% achieved on Yale B of ten subjects, and 95.61% achieved on Outdoor database of 132 subjects under uncontrolled natural lighting conditions show that Gradientfaces is an effective method for face recognition under varying illumination. Furthermore, the experimental results on Yale database validate that Gradientfaces is also insensitive to image noise and object artifacts (such as facial expressions).
Taiping Zhang, Yuan Yan Tang, Bin Fang 0001, Zhaowei Shang
IEEE Trans. Image Process.3
2008 Total variation norm-based nonnegative matrix factorization for identifying discriminant representation of image patterns
Taiping Zhang, Bin Fang 0001, Weining Liu, Yuan Yan Tang
Neurocomputing2
2008 Topology Preserving Non-negative Matrix Factorization for Face Recognition
abstract
In this paper, a novel topology preserving non-negative matrix factorization (TPNMF) method is proposed for face recognition. We derive the TPNMF model from original NMF algorithm by preserving local topology structure. The TPNMF is based on minimizing the constraint gradient distance in the high-dimensional space. Compared with L(2) distance, the gradient distance is able to reveal latent manifold structure of face patterns. By using TPNMF decomposition, the high-dimensional face space is transformed into a local topology preserving subspace for face recognition. In comparison with PCA, LDA, and original NMF, which search only the Euclidean structure of face space, the proposed TPNMF finds an embedding that preserves local topology information, such as edges and texture. Theoretical analysis and derivation given also validate the property of TPNMF. Experimental results on three different databases, containing more than 12,000 face images under varying in lighting, facial expression, and pose, show that the proposed TPNMF approach provides a better representation of face patterns and achieves higher recognition rates than NMF.
Taiping Zhang, Bin Fang 0001, Yuan Yan Tang
IEEE Trans. Image Process.2
2007 Offline signature verification: A new rotation invariant approach
abstract
Rotation problem is one of the major difficulties to distinguish signature patterns in off-line skilled signature verification. This paper presents a new approach utilizing Ring- Peripheral features to tackle this problem. In principle, Ring- Peripheral features are able to describe internal and external structure of signatures with different phase shift. In order to extract stable and consistent presentation of signature patterns for verification purpose, FFT is used to eliminate phase effects. In the training samples stage, we employ a selection function to pick up reasonable samples for better threshold estimation. Experiment results demonstrated that the proposed method was successful to improve verification accuracy.
Bin Fang 0001, Yuan Yan Tang, Taiping Zhang
SMC2
2006 Two-step single parameter regularization fisher discriminant method for face recognition
abstract
In face recognition tasks, Fisher discriminant analysis (FDA) is one of the promising methods for dimensionality reduction and discriminant feature extraction. The objective of FDA is to find an optimal projection matrix, which maximizes the between-class-distance and simultaneously minimizes within-class-distance. The main limitation of traditional FDA is the so-called Small Sample Size (3S) problem. It induces that the within-class scatter matrix is singular and then the traditional FDA fails to perform directly for pattern classification. To overcome 3S problem, this paper proposes a novel two-step single parameter regularization Fisher discriminant (2SRFD) algorithm for face recognition. The first semi-regularized step is based on a rank lifting theorem. This step adjusts both the projection directions and their corresponding weights. Our previous three-to-one parameter regularized technique is exploited in the second stage, which just changes the weights of projection directions. It is shown that the final regularized within-class scatter matrix approaches the original within-class scatter matrix as the single parameter tends to zero. Also, our method has good computational complexity. The proposed method has been tested and evaluated with three public available databases, namely ORL, CMU PIE and FERET face databases. Comparing with existing state-of-the-art FDA-based methods in solving the S3 problem, the proposed 2SRFD approach gives the best performance.
Pong C. Yuen, Jian Huang 0009, Bin Fang 0001
Int. J. Pattern Recognit. Artif. Intell.4
2006 Handwriting-based personal identification
abstract
Handwriting-based personal identification, which is also called handwriting-based writer identification, is an active research topic in pattern recognition. Despite continuous effort, offline handwriting-based writer identification still remains as a challenging problem because writing features can only be extracted from the handwriting image. As a result, plenty of dynamic writing information, which is very valuable for writer identification, is unavailable for offline writer identification. In this paper, we present a novel wavelet-based Generalized Gaussian Density (GGD) method for offline writer identification. Compared with the 2-D Gabor model, which is currently widely acknowledged as a good method for offline handwriting identification, GGD method not only achieves a better identification accuracy but also greatly reduces the elapsed time on calculation in our experiments.
Zhenyu He 0001, Xinge You, Yuan Yan Tang, Bin Fang 0001, Jianwei Du
Int. J. Pattern Recognit. Artif. Intell.4
2006 Thinning Character Using Modulus Minima of Wavelet Transform
abstract
An essential step in character recognition is to extract the skeleton characteristics of the character. In this paper, an efficient algorithm is proposed to extract visually satisfactory skeleton from printed and handwritten characters, which overcomes fundamental shortcomings of our previous skeletonization technique based on the maximum modulus symmetry of wavelet transform (WT). The proposed method is motivated from some desirable properties of the WT with constructed wavelet functions: namely, the local modulus minima of the WT are scale-independent at different level scales and are located at the medial axis of the symmetrical contours of character stroke. Thus the modulus minima of the WT are computed as the intrinsic skeletons of character strokes. To achieve faster implementation, a multiscale processing technique is employed. Thus major structures of the skeleton are extracted using the coarse scale, while fine structures are extracted using the fine scale. We have tested the algorithm on handwritten and printed character images. Experimental results show that the proposed algorithm is applicable to not only binary image but also gray-level image where it can be impractical to use other skeletonization techniques, such as thinning and distance transforms. Further, it can effectively remove unwanted artifacts and branches from the extracted skeletons at the intersections and junctions of character strokes and is robust against noises while most existing methods perform poorly.
Xinge You, Qiuhui Chen, Bin Fang 0001, Yuan Yan Tang
Int. J. Pattern Recognit. Artif. Intell.3
2005 A Novel Method for Off-line Handwriting-based Writer Identification
abstract
Handwriting-based writer identification is a hot research topic in the pattern recognition field. Nowadays, online handwriting-based writer identification is steadily growing toward its maturity. On the contrary, offline handwriting-based writer identification still remains as a challenging problem because writing features only can be extracted from the handwriting image in this situation. As a result, plenty of dynamic writing information, which is very valuable for writer identification, is lost. At present, 2D Gabor filter method is widely acknowledged as a good method for offline handwriting identification, however it still suffers from some inherent disadvantages, such as the high computational cost. In this paper, we present a novel wavelet-based GGD method to replace the traditional 2D Gabor filters. Shown in our experiments, this novel method not only achieves better experiment results but also greatly reduces the elapsed time on calculation.
Zhenyu He 0001, Yuan Yan Tang, Bin Fang 0001, Jianwei Du, Xinge You
ICDAR3
2005 Improved DTW Algorithm for Online Signature Verification Based on Writing Forces
Ping Fang, ZhongCheng Wu, YunJian Ge, Bin Fang 0001
ICIC (1)5
2005 Similarity Measurement for Off-Line Signature Verification
Xinge You, Bin Fang 0001, Zhenyu He 0001, Yuan Yan Tang
ICIC (1)2
2005 Locating Vessel Centerlines in Retinal Images Using Wavelet Transform: A Multilevel Approach
Xinge You, Bin Fang 0001, Yuan Yan Tang, Zhenyu He 0001, Jian Huang 0009
ICIC (1)2
2005 The closed-loop human eye-brain-hand to computer (EBH-C) interface for hand sensory-motor coordination based on force tablet
abstract
The mechanism of the sensory-to-motor transformation as well as motor-to-sensory transformation of human beings has attracted much attention in recent years. As no efficient device can record hand intrinsic behavior, it is difficult to get its neuro-physiological models of sensory-motor coordination. In this paper, we offer a system to acquire the kinematics and kinetics information of human hand movement through handwriting. Joined with human being, an eye-brain-hand to computer (EBH-C) interaction system is presented. In this human-in-the-loop-testing system, human beings acquire image, voice or text from computer by eyes or ears then write them down. The core part of the system, named as F-Tablet/spl trade/ is able to acquire the trajectory and three-axis forces of pen-tip directly and simultaneously. With the help of this system, we designed an experiment to evaluate the handwriting movements and forces controlling ability of different ages. Some experiment results were present. Aided with conventional analysis of electroencephalography (EEG) and magnetoencephalography (MEG), the whole procedures of information transmitting, acquired by eyes and ears, processed by brain, outputted and actuated by hand, can be recorded.
ZhongCheng Wu, Mingxu Wei, Yong Yu 0003, Bin Fang 0001
IROS6
2005 Morphological structure reconstruction of retinal vessels in fundus images
abstract
Vessels in retinal fundus images are useful in revealing the severity of eye-related diseases. In addition, they can act as landmarks for localizing lesions or the central vision area, and guide laser treatment of neovascularization. In this paper, we propose a two-stage scheme to extract vessels and reconstruct the morphological structure of vessels in retinal images. First, we employ mathematical morphology techniques to highlight large and small vessels with respect to their spatial properties. Different curvature response between vessel and noise patterns allows the use of curvature evaluation to remove enhanced vessel-like noise. A set of linear filters finalize the vessel map. However, the resulting vascular structure is incomplete of some important features in bifurcation points and central reflex. In order to rectify the pitfall, a reconstruction process is performed using dynamic local region growth to recover the morphological structure of vessels. Average performance of our method to extract vessels is 83.7% of TPR(True positive rate) and 3.8% of FPR(False positive rate) for 35 retinal images which include 21 abnormal images.
Bin Fang 0001, Xinge You, Yuan Yan Tang
Int. J. Pattern Recognit. Artif. Intell.1
2005 A wavelet-based approach to ridge thinning in fingerprint images
abstract
As a global feature of fingerprints, the thinning of ridges, extraction of minutiae and computation of orientation field are very important for automatic fingerprint recognition. Many algorithms have been proposed for their computation and estimation, but their results are unsatisfactory, especially for poor quality fingerprint images. In this paper, a robust wavelet-based method to create thinned ridge map of fingerprint for automatic recognition is proposed. Properties of modulus minima based on the spline wavelet function are substantially investigated. Desirable characteristics show that this method is suitable to describe the skeleton of the ridge of the fingerprint image. A multi-scale thinning algorithm based on the modulus minima of wavelet transform is presented. The proposed algorithm is able to improve the skeleton representation of the ridge of the fingerprint without side-effects and limitations of the existing methods. The thinned ridge map can facilitate the extraction of the minutiae for matching in fingerprint recognition. Experiments have been conducted to validate the effectiveness and efficiency of the proposed method.
Xinge You, Bin Fang 0001, Yuan Yan Tang, Zhenyu He 0001
Int. J. Pattern Recognit. Artif. Intell.2
2005 Improved class statistics estimation for sparse data problems in offline signature verification
abstract
Sparse data problems are prominent in applications of offline signature verification. By using a small number of training samples, the class statistics estimation errors may be significant, resulting in worsened verification performance. In this paper, we propose two methods to improve the statistics estimation. The first approach employs an elastic distortion model to artificially generate additional training samples for pairs of genuine signatures. These additional samples, together with original genuine samples, are used to compute statistic parameters for a Mahalanobis distance threshold classifier. The other approach is to adopt regularization techniques to overcome the problem of inverting an ill-conditioned sample covariance matrix due to insufficient training samples. A ridge-like estimator is modeled to add some constant values for diagonal elements of the sample covariance matrix. Experimental results showed that both methods were able to improve verification accuracy when they were incorporated with a set of peripheral features. Effectiveness of the methods was validated by quantity analysis.
Bin Fang 0001, Yuan Yan Tang
IEEE Trans. Syst. Man Cybern. Part C1
2003 Off-line signature verification by the tracking of feature and stroke positions
Bin Fang 0001, Cheung Hoi Leung, Yuan Yan Tang, K. W. Tse 0001, Paul C. K. Kwok
Pattern Recognit.1
2001 Offline Signature Verification by the Analysis of Cursive Strokes
abstract
In this paper, a method is proposed for offline signature verification. It is based on a smoothness criterion. It is observed that the cursive segments of forgery signatures are generally less smooth and less natural than the genuine ones, especially for those signatures that consist of cursive graphic patterns. Two approaches are proposed to extract a smoothness feature: a crossing method and a fractal dimension method. When the proposed smoothness feature is combined with other global shape features for signature verification, satisfactory results are obtained.
Bin Fang 0001, Y. Y. Wang, Cheung Hoi Leung, K. W. Tse 0001, Yuan Yan Tang, Paul C. K. Kwok
Int. J. Pattern Recognit. Artif. Intell.1
1999 A Smoothness Index based Approach for Off-line Signature Verification
abstract
Proposes a method to tackle the problem of detecting skilled forgeries in off-line signature verification. Inspired by the approach adopted by expert examiners, it is based on a smoothness criterion. From a collection of genuine and forged signatures, it is observed that, although skilled forgery signatures are very similar to genuine ones on a global scale, they are generally less smooth and natural on a detailed scale than the genuine ones, especially for those skilled forgery signatures which consist of cursive graphic patterns. A smoothness index is derived from such signatures. This is combined with other global shape features and used for verification. Satisfactory results are obtained.
Bin Fang 0001, Y. Y. Wang, Cheung Hoi Leung, Yuan Yan Tang, Paul C. K. Kwok, K. W. Tse 0001
ICDAR1