Patrick Shen-Pei Wang

dblp:05/4537 · also Patrick S. P. Wang, Patrick Wang 0004 · DBLP profile ↗
← Back
117ranked-venue papers
24as first author
24since 2021 · last 2026
0000-0002-9336-3155ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 98 · 13 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-authorDatabases, data management, data science and information retrieval · 9 · 7 first-authorApplied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorTheory of computation · 4 · 4 first-authorSystems, architecture and hardware · 2 · 2 first-authorSecurity and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2026 Unlocking the Potential of Auxiliary Captions via Dual-Branch Multi-Scale Network for Composed Image Retrieval
abstract
Composed image retrieval (CIR) aims to retrieve target images by combining a reference image with a modification text. Traditional CIR methods often struggle with feature-level multimodal fusion, leading to deviations from the original embedding space. To address this, we propose a Dual-Branch Multi-Scale Network (DMN) that integrates a combining branch and a complete text branch. To enhance the use of captions generated by advanced image captioning models for CIR, the DMN leverages an attribute-driven disentanglement layer to separate features into distinct latent factors and employs a dual-path multimodal fusion module for effective feature integration. Additionally, a multi-scale matching module incorporating both global and local matching strategies is introduced to enhance fine-grained feature discrimination. Experimental results on the FashionIQ, Shoes, and CIRR datasets demonstrate that our DMN model consistently outperforms state-of-the-art methods, achieving improvements of up to 1.43% in mean recall metrics.
Jinhong Xu, Xichun Li, Thomas Wu 0001, Yuan Yan Tang, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.6
2026 Similarity-Guided Denoising Reconstruction for Unsupervised Image Captioning
abstract
Image captioning aims to generate natural and accurate textual descriptions of given images. Although significant progress has been made in image captioning models in recent years, most existing approaches heavily rely on high-quality image-text paired datasets that require expensive human annotation, thus limiting model scalability. Current unsupervised image captioning methods primarily focus on leveraging zero-shot learning capabilities of large pre-trained models (e.g. CLIP, GPT-2), yet still face persistent challenges including modality gaps, inefficient inference, and excessive noise incorporation, which constrain model accuracy and generalization capabilities. To address these limitations, we propose SGDR-Cap ( Similarity- Guided Denoising Reconstruction for Captioning), a novel unsupervised image captioning method that bridges the vision-language modality gap through a similarity-guided denoising reconstruction module. Our method leverages similarity information to guide the reconstruction of authentic text features during caption generation while simultaneously forcing the model to learn how to extract crucial image-relevant features and filter out unnecessary noise information. This enhances both coarse- and fine-grained cross-modal alignment. Furthermore, our approach jointly optimizes denoising reconstruction loss and language modeling loss, ensuring accuracy and fluency, and promoting greater diversity. Extensive evaluations on the MSCOCO and Flickr30K benchmarks demonstrate that our method achieves state-of-the-art results across all major metrics, with the most notable gain on the CIDEr score, improving from 101.1 to 104.4.
Dongnan Yang, Thomas Wu 0001, Xichun Li, Yuan Yan Tang, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.6
2025 STSODNet: Scale Transformer Small Object Detection Network
abstract
Dense small object detection in complex scenes is a valuable and challenging research field. While deep learning has driven significant advancements in computer vision, traditional object detection models still struggle to achieve high accuracy in detecting small objects, particularly in large-scale aerial images. Challenges such as scale variations, occlusions, and complex backgrounds continue to hinder the effective detection of dense small objects. In this paper, we present the Scale Transformer Small Object Detection Network (STSODNet), a novel architecture designed to address these challenges. First, we conceptualize the pronounced scale variation in drone images as an anomalous disturbance and propose a multiscale feature enhancement module (MSFEM), built upon the Spatial Transformer Network, to mitigate this effect. The multiscale feature enhancement module performs learnable, multi-point magnification on regions surrounding objects based on spatial saliency, enhancing the model’s scale invariance. Second, to generate a more accurate global saliency map and heighten the model’s focus on small target regions, we introduce a refined spatial attention mechanism, termed Spatial Region Attention. This mechanism combines coarse region attention with fine spatial attention to produce a more detailed saliency map and improve long-range dependency capture. Third, to achieve more accurate spatial regression of small objects, the traditional three-layer detection head is improved by expanding its output layer, resulting in a finer and larger output while maintaining the same number of parameters. Extensive experiments on the VisDrone and SeaPerson benchmark datasets validate that STSODNet achieves superior precision and robustness, outperforming current state-of-the-art object detection methods for small object detection.
Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.3
2025 Analysis of Job Offers to Measure Gender Barriers through Natural Language Processing and Soft Computing Techniques
abstract
Gender-biased language is still traced in job advertisements. Legal requirements to avoid direct gender-biased adjectives, and the usage of special software to detect and substitute gender-based words, scale up the issue more than solve it. The veil of discrimination on gender in job advertisements becomes more sophisticated with each succeeding level of its official and technical (including AI) prevention. This paper is mainly focused on the application of natural language processing (NLP) to detect gender-biased and discrimination of candidates by analyzing job offers posted online. NLP is an Artificial Intelligence tool that was applied in combination with Term Frequency-Inverse Document Frequency (TF-IDF) and Latent Dirichlet Allocation (LDA) to analyze the type of language used in job advertisements, detect the most relevant words used in the ads, and ultimately detect gender-bias. The main objective of this work is to provide equal access to employment opportunities from the very initial stage of the recruitment process. In addition, clustering techniques were applied to create groups based on the target public and the type of language used, providing evidence of gender-biased practices. The system was tested using a database of 2000 job ads in four different sectors: nursery, secretarial, managerial, and engineering.
Cristina Puente, Ivan Sanchez-Perez, Evhenia Kolomiyets-Ludwig, Clara Palacios-Castrillo, Patrick Shen-Pei Wang, Rafael Palacios
Int. J. Pattern Recognit. Artif. Intell.5
2025 Multi-Scale Adaptive Diffusion Feature Fusion with Causal Invariance Learning for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification is a significant research area in remote sensing with a wide range of application scenarios. Recently, numerous HSI classification methods based on convolutional neural networks (CNNs) and Transformers have demonstrated promising classification performance. However, these methods demonstrate insufficient capability in mining spectral–spatial relationships from limited HSI samples and fail to extract features pertinent to the target category. To address these challenges, we propose a multi-scale adaptive diffusion feature fusion and causal invariance learning (MSDF-CIL) framework based on diffusion models. Specifically, the framework establishes spectral–spatial distribution relationships through forward and backward diffusion processes. The forward process gradually introduces noise to the HSI input. In the backward diffusion process, our pre-trained hyperspectral denoising network extracts semantically rich multi-scale diffusion features from complex spectral–spatial relationships. A multi-scale adaptive diffusion feature fusion (MSADFF) module is designed to learn key information about each scale and fuse it to enhance the representation. In addition, a causal invariance learning (CIL) module is designed to focus on features causally related to the target class, enabling the model to eliminate spurious correlations among diffusion features. Experimental results on three public HSI datasets show that the proposed MSDF-CIL outperforms other state-of-the-art HSI classification methods, even with minimal samples. When the number of training samples in each class reaches 30, the MSDF-CIL achieves overall accuracy improvements of 4.0% on the Pavia University dataset, 2.3% on the Indian Pines dataset, and 2.3% on the Houston13 dataset, respectively.
Wanxing Zha, Huafu Xu, Bingzhen Wang, Yuwen Lin, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.8
2024 Configurable Customized Information Extraction and Processing Pipeline
abstract
Extracting information from scanned business documents, while a necessary commercial task, continues to be mostly done manually, requiring significant human effort. Current solutions for automated document information extraction still have limited capabilities in regards to user-required customizability and extraction of dataset-specific information, leaving the area as a very active field of research. In this paper, we propose modifications and improvements to our previously developed custom pipeline for extracting and tabulating key-value pairs from commercial invoice documents. Our design changes and additions adapt the pipeline to a wider variety of document types and use cases, primarily through the implementation of dataset-specific configuration files that promote customizability along with new technical modules that address both general and dataset-specific complexities. We compare our pipeline’s performance against current machine learning and commercial solutions on a real-world dataset, and demonstrate that it is able to extract a wider variety of fields while maintaining competitive or greater accuracies compared to the alternate solutions.
Pierce Lai, Dariyan Khan, Kevin Zhao, Brian Le, Alex Luchianov, Margaret Yu, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.8
2024 Abnormal Detection of Commutator Surface Defects Based on YOLOv8
abstract
The YOLOv8 model has high detection efficiency and classification accuracy in detecting commutator surface defects, aimed at the problem of low working efficiency of a commutator, caused by commutator surface defects. First, the theoretical framework of Region-based Convolutional Neural Networks (R-CNN), spatial pyramid pooling (SPP)-net, Fast R-CNN, and Faster R-CNN is introduced, and the detection principle and process are described in detail. Secondly, the principle of the YOLOv8 network structure, head structure, neck structure, and C2f module are explained, and the loss function is described. The average precision of the proposed algorithm for detecting cracks and small points is more than 98%, and the frames per second (FPS) is 27. The detection results are mapped to the original image, and the visualization of the commutator surface defect detection is obtained, which has a higher robustness, accuracy, and real-time performance than the R-CNN, SPP-net, Fast R-CNN, and Faster R-CNN algorithms.
Ban-Hoe Kwan, Mau-Luen Tham, Oon-Ee Ng, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.5
2024 LCSTR: Scene Text Recognition with Large Convolutional Kernels
abstract
The task of scene text recognition involves processing information from two modalities: images and text, thereby requiring models to have the ability to extract features from images and model sequences simultaneously. Although linguistic knowledge greatly aids scene text recognition tasks, the extensive use of language models in sequence modeling and model prediction stages in recent years has made model architectures increasingly complex and inefficient. In this paper, we propose LCSTR, a pure convolutional visual model that can complete text recognition without the need for attention mechanisms or language models. This approach applies large kernels to text recognition tasks for the first time, extracting word-level text information through large text-aware blocks, capturing long-range dependencies between characters, and using small text-aware blocks to obtain local features within characters. Experiments show that this model strikes a good trade-off between accuracy and speed, achieving notable results on seven public benchmarks, validating the generalizability and effectiveness of this method. Furthermore, owing to the absence of a language module, this model demonstrates remarkable accuracy even in limited sample scenarios, and the lightweight and low computational overhead features make it suitable for engineering applications.
Jing Wang 0233, Patrick Shen-Pei Wang, Xichun Li, Huiwu Luo, Huafu Xu
Int. J. Pattern Recognit. Artif. Intell.6
2023 Constraint-Based Adversarial Networks for Unsupervised Abstract Text Summarization
abstract
Abstract text summarization is a classic sequence-to-sequence natural language generation task. In order to improve the quality of unsupervised abstract text summarization in unsupervised mode, we propose two constraints for training text summarization model, embedding space constraint and information ratio constraint. We construct a generative adversarial network with two discriminators based on these two constraints (TC-SUM-GAN). We use unsupervised and supervised methods to train the model in the experiment. Experimental results show that the ROUGE-1 value of the unsupervised TC-SUM-GAN increases by [Formula: see text] points compared with the basic model and at least 1.96 points compared with other comparative models. The ROUGE scores of the supervised TC-SUM-GAN are also improved. TC-SUM-GAN achieves very competitive results for the metrics of ROUGE-1 and ROUGE-2. In addition, the abstracts generated by our model are closer to those generated manually.
Liwei Jing, Yujian Yuan, Zuqiang Meng, Yifeng Tan, Patrick Shen-Pei Wang, Xichun Li
Int. J. Pattern Recognit. Artif. Intell.6
2023 Customized Information Extraction and Processing Pipeline for Commercial Invoices
abstract
Extracting information from scanned invoices and other commercial documents, a critical component of corporate function, typically requires significant manual processing. Much research has been conducted in the field of automated information extraction and document processing to alleviate the manual resources used for document analysis, but resultant literature and commercially available products have demonstrated limitations in customizability for identifying specific information. In this paper, we propose a customized machine learning-based pipeline for extracting and tabulating relevant key–value pairs from commercial invoice documents. Specifically, the pipeline combines general document understanding, OCR extraction, and key–value matching with custom rules pertaining to a provided invoice dataset. Then, we demonstrate that the pipeline greatly outperforms a commercially available product and can significantly reduce the amount of manual labor required to process invoice documents. Future work will focus on generalizing the pipeline, so as to apply it on more varied datasets.
Pierce Lai, Abhishek Mohan, Jung Soo Victor Chu, Samuel Lee, Prabhakar Kafle, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.7
2023 3D PET/CT Tumor Co-Segmentation Based on Background Subtraction Hybrid Active Contour Model
abstract
Accurate tumor segmentation in medical images plays an important role in clinical diagnosis and disease analysis. However, medical images usually have great complexity, such as low contrast of computed tomography (CT) or low spatial resolution of positron emission tomography (PET). In the actual radiotherapy plan, multimodal imaging technology, such as PET/CT, is often used. PET images provide basic metabolic information and CT images provide anatomical details. In this paper, we propose a 3D PET/CT tumor co-segmentation framework based on active contour model. First, a new edge stop function (ESF) based on PET image and CT image is defined, which combines the grayscale standard deviation information of the image and is more effective for blurry medical image edges. Second, we propose a background subtraction model to solve the problem of uneven grayscale level in medical images. Apart from that, the calculation format adopts the level set algorithm based on the additive operator splitting (AOS) format. The solution is unconditionally stable and eliminates the dependence on time step size. Experimental results on a dataset of 50 pairs of PET/CT images of non-small cell lung cancer patients show that the proposed method has a good performance for tumor segmentation.
Laquan Li, Chuangbo Jiang, Patrick Shen-Pei Wang, Shenhai Zheng
Int. J. Pattern Recognit. Artif. Intell.3
2023 Edge-Labeled and Node-Aggregated Graph Neural Networks for Few-Shot Relation Classification
abstract
Relation classification as a core technique for building knowledge graphs becomes a critical task in natural language processing. The fact that humans can learn by summarizing and generalizing limited knowledge motivates scholars to explore few-shot learning. Graph neural networks provide a method to measure the distance between nodes, which improves the model effect in the problem of few-shot relation classification. However, graph neural network methods focus only on node information and ignore edge information which implies inter-class and intra-class relations. This paper proposes edge-labeled and node-aggregated graph neural networks (ENGNNs) for few-shot relation classification: edge labels are encoded and used for node information aggregation. In addition, a process of semi-supervised learning is designed to discover a better solution for one-shot learning. Compared with previous methods, experimental results show that the proposed ENGNN model improves the performance of the graph neural network on the FewRel dataset.
Xichun Li, Patrick Shen-Pei Wang, Zuqiang Meng
Int. J. Pattern Recognit. Artif. Intell.4
2022 Phylogenetic Analysis: A Novel Method of Protein Sequence Similarity Analysis
abstract
Protein sequence similarity analysis (PSSA) is a significant task in bioinformatics, which can obtain information about unknown sequences such as protein structures and homology relationships. Protein sequence refers to the series of amino acids with rich physical and chemical properties, namely the basic structure of proteins. However, sequence similarity analysis and phylogenetic analysis between different species which have complex amino acid sequences is a challenging problem. In this paper, nine properties of amino acids were considered and the sequence was converted into numerical values by principal component analysis (PCA); with Haar Wavelet Transform, and Higuchi fractal dimension (HFD), a new feature vector is constructed to represent the sequence; Spearman distance was selected to calculate the distance matrix and the phylogenetic tree was constructed. In this paper, two representative protein sequences (9 ND5 (NADH dehydrogenase 5) and 8 ND6 (NADH dehydrogenase 6)) were selected for similarity analysis and phylogenetic analysis, and compared with MEGA software and other existing methods. The extensive results show that our method is outperforming and results consistent with the known facts.
Wei Li 0330, Zuqiang Meng, Patrick Shen-Pei Wang, Xichun Li
Int. J. Pattern Recognit. Artif. Intell.5
2022 Improving Utterance Rewriter Based on MMI and Text Data Augmentation
abstract
In multi-round dialogue tasks, how to maintain the consistency of model answers is a major research challenge. Every answer to the model should be time dependent, causal, and logical. In order to maintain the consistency of the personality, dialogue style, and context of the model, it is necessary to retain the key information in the historical dialogue as much as possible so that the model can generate more accurate answers. Utterance rewriting is a technique that replenishes the information of the current sentence by analyzing the historical dialogue, so as to retain the key information. This paper mainly uses text augmentation, Maximum Mutual Information (MMI) method and character correction method based on Knuth–Morria–Pratt (KMP) algorithm to improve the effect of utterance rewriting generation. The number of original statement rewriting datasets is limited, and the cost of manual manufacturing is too high. By using the method of text data augmentation based on coreference resolution, the positive dataset that is missing from the statement rewriting dataset is repaired. At the same time, the existing datasets are expanded to increase the number of data. The generated results are optimized by using the MMI method, and the KMP character correction method is used to modify the wrong characters to improve the overall accuracy.
Wei Li 0330, Zuqiang Meng, Patrick Shen-Pei Wang, Xichun Li, Huiwu Luo
Int. J. Pattern Recognit. Artif. Intell.5
2022 Automatic Detection of Bridge Surface Crack Using Improved YOLOv5s
abstract
Bridge crack detection is a key task in the structural health monitoring of Civil Engineering. In the traditional bridge crack detection methods, there exist some problems such as high cost, low speed, and complex structure. This paper developed a bridge surface crack detection system based on improved YOLOv5s. The GhostBottleneck module was employed to replace the classic C3 module of the YOLOv5s backbone network, meanwhile the channel attention module namely ECA-Net was also added to the network, which not only reduced the amount of calculation, but also enhanced the ability of the network in extracting cross-channel information features. The adaptive spatial feature fusion (ASFF) was introduced to address the conflict problem caused by the inconsistency of feature scale in the network feature fusion stage, and the transfer learning was utilized to train the network. The experimental results showed that the improved YOLOv5s performed better than Faster R-CNN, SSD, YOLOv3, and YOLOv5s, with the Precision of 93.6%, Recall of 95.4%, and mAP of 98.4%. Further, the improved YOLOv5s was deployed in PyQt5 to realize the real-time detection of bridge cracks. This research showed that the proposed model not only provides a novel solution for bridge surface crack detection, but also has certain industrial application value.
Thomas Wu 0001, Zuqiang Meng, Youju Huang, Patrick Shen-Pei Wang, Xichun Li
Int. J. Pattern Recognit. Artif. Intell.6
2022 Multi-Content Merging Network Based on Focal Loss and Convolutional Block Attention in Hyperspectral Image Classification
abstract
Simultaneous extraction of spectral and spatial features and their fusion is currently a popular solution in hyperspectral image (HSI) classification. It has achieved satisfactory results in some research. Because the scales of objects are often different in HSI, it is necessary to extract multi-scale features. However, this aspect was not taken into account in many spectral-spatial feature fusion methods. This causes the model to be unable to get sufficient features on scales with a large difference range. The model (MCMN: Multi-Content Merging Network) proposed in this paper designs a multi-branch fusion structure to extract multi-scale spatial features by using multiple dilated convolution kernels. Considering the interference of the surrounding heterogeneous objects, the useful information from different directions is also fused together to realize the merging of multiple regional features. MCMN introduces a convolution block attention mechanism, which fully extracts attention features in both spatial and spectral directions, so that the network can focus on more useful parts, which can effectively improve the performance of the model. In addition, since the number of objects in each class is often discrepant, it will have some impact on the training process. We apply the focal loss function to eliminate the negative factor. The experimental results of MCMN on three data sets have a breakthrough compared with the other comparison models, which highlights the role of MCMN structure.
Fengqi Zhang, Patrick Shen-Pei Wang, Xichun Li, Huiwu Luo
Int. J. Pattern Recognit. Artif. Intell.3
2022 Multi-scale spatial-spectral fusion based on multi-input fusion calculation and coordinate attention for hyperspectral image classification
Fengqi Zhang, Patrick Shen-Pei Wang, Xichun Li, Zuqiang Meng
Pattern Recognit.3
2021 Real-Time Stair Detection Using Multi-stage Ground Estimation Based on KMeans and RANSAC
Patrick Shen-Pei Wang
WorldCIST (1)3
2021 W-core Transformer Model for Chinese Word Segmentation
Patrick Shen-Pei Wang
WorldCIST (1)3
2021 Improved Multi-scale Fusion of Attention Network for Hyperspectral Image Classification
Fengqi Zhang, Patrick Shen-Pei Wang
WorldCIST (1)4
2021 Guest Editorial
Yue Lu 0001, Nicole Vincent, Ching Y. Suen, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.4
2021 35th Anniversary of IJPRAI
Patrick Shen-Pei Wang, Xiaoyi Jiang 0001, Frank Y. Shih, Terence Sim
Int. J. Pattern Recognit. Artif. Intell.1
2021 Smart Home Privacy Protection Based on the Improved LSB Information Hiding
abstract
Smart home is an emerging form of the Internet of Things (IoT), enabling people to enjoy a convenient and intelligent life. The data generated by smart home devices are transmitted through the public channel, which is not secure enough, so the secret data in smart home are easily intercepted by malicious adversaries. In order to solve this problem, this paper proposes a smart home privacy protection method combining DES encryption and the improved Least Significant Bit (LSB) information hiding algorithm, changing the practice of directly exposing smart home secret information to the Internet, first, using Data Encryption Standard (DES) encryption to encrypt the smart home information and second, the improved LSB information hiding algorithm is used to hide the ciphertext, so that the adversary cannot detect the smart home secret information. The goal of the scheme is to provide a double protection for the secure transmission of the smart home secret information. If an attacker wants to carry out an attack, it has to break through at least two defense lines, which seems impossible to do. Experiment results show that the improved LSB algorithm is more robust than the existing algorithms, and it is very safe. Therefore, the scheme proposed in this paper is very practical for protecting the smart home secret information.
Haiyu Deng, Ren Ping Liu 0001, Patrick Shen-Pei Wang, Xiaocui Dang, Yuan Yan Tang, Xichun Li
Int. J. Pattern Recognit. Artif. Intell.4
2021 PseKNC and Adaboost-Based Method for DNA-Binding Proteins Recognition
abstract
DNA-binding proteins are an essential part of the DNA. It also an integral component during life processes of various organisms, for instance, DNA recombination, replication, and so on. Recognition of such proteins helps medical researchers pinpoint the cause of disease. Traditional techniques of identifying DNA-binding proteins are expensive and time-consuming. Machine learning methods can identify these proteins quickly and efficiently. However, the accuracies of the existing related methods were not high enough. In this paper, we propose a framework to identify DNA-binding proteins. The proposed framework first uses PseKNC (ps), MomoKGap (mo), and MomoDiKGap (md) methods to combine three algorithms to extract features. Further, we apply Adaboost weight ranking to select optimal feature subsets from the above three types of features. Based on the selected features, three algorithms (k-nearest neighbor (knn), Support Vector Machine (SVM), and Random Forest (RF)) are applied to classify it. Finally, three predictors for identifying DNA-binding proteins are established, including [Formula: see text], [Formula: see text], [Formula: see text]. We utilize benchmark and independent datasets to train and evaluate the proposed framework. Three tests are performed, including Jackknife test, 10-fold cross-validation and independent test. Among them, the accuracy of ps+md is the highest. We named the model with the best result as psmdDBPs and applied it to identify DNA-binding proteins.
Patrick Shen-Pei Wang, Xichun Li
Int. J. Pattern Recognit. Artif. Intell.4
2019 Improving Text-Independent Chinese Writer Identification with the Aid of Character Pairs
abstract
Text-independent Chinese writer identification does not depend on the text content of the query and reference handwritings. In order to deal with the uncertainty of the text content, text-independent approaches usually give special attention to the global writing style of handwriting, rather than the properties of each individual character or word. Thanks to the existence of high-frequency characters, some characters probably appear in both the query and reference handwritings in most cases. If character images in the query handwriting are similar to those in the reference handwriting, this query handwriting and the corresponding reference handwriting are very likely to be written by the identical writer. In this paper, we exploit the above characteristic to improve the performance of Chinese writer identification. We first present an identification scheme using edge co-occurrence feature (ECF). Then, we detect the character pairs in the query and reference handwritings using a two-step framework and propose the displacement field-based similarity (DFS) to determine whether a character pair is written by the identical writer. The character pairs help to re-rank the candidate list obtained by text-independent ECF-based similarity and finally decide the writer of the query handwriting. The proposed method is evaluated on the HIT-MW and CASIA-2.1 datasets. Experimental results demonstrate that our proposed method outperforms the existing ones, and its Top-1 accuracy on the two datasets reaches 97.1% and 98.3%, respectively.
Yujie Xiong, Li Liu 0010, Shujing Lyu, Patrick Shen-Pei Wang, Yue Lu 0001
Int. J. Pattern Recognit. Artif. Intell.4
2019 A Fractal Dimension and Empirical Mode Decomposition-Based Method for Protein Sequence Analysis
abstract
In bioinformatics, the biological functions of proteins and their interactions can often be analyzed by the similarity of their sequences. In this paper, the authors combine the fractal dimension, empirical mode decomposition (EMD), and sliding window for protein sequence comparison. First, the protein sequence is characterized and digitized into a signal, and then the signal characteristics are obtained by using EMD and fractal dimension. Each protein sequence can be decomposed into Intrinsic Mode Functions (IMFs). The fixed window’s fractal dimension is applied to each IMF and the original signal to extract the protein sequence characteristics. Experiments have shown that the feature extracted by this hybrid method is superior to the EMD method alone.
Pu Wei, Zuqiang Meng, Patrick Shen-Pei Wang, Yuan Yan Tang
Int. J. Pattern Recognit. Artif. Intell.5
2019 Multi-Level Downsampling of Graph Signals via Improved Maximum Spanning Trees
abstract
Graph signal processing (GSP) is an emerging field in the signal processing community. Novel GSP-based transforms, such as graph Fourier transform and graph wavelet filter banks, have been successfully utilized in image processing and pattern recognition. As a rapidly developing research area, graph signal processing aims to extend classical signal processing techniques to signals with irregular underlying structures. One of the hot topics in GSP is to develop multi-scale transforms such that novel GSP-based techniques can be applied in image processing or other related areas. For designing graph signal multi-scale frameworks, downsampling operations that ensuring multi-level downsampling should be specifically constructed. Among the existing downsampling methods in graph signal processing, the state-of-the-art method was constructed based on the maximum spanning tree (MST). However, when using this method for multi-level downsampling of graph signals defined on unweighted densely connected graphs, such as social network data, the sampling rates are not close to [Formula: see text]. This phenomenon is summarized as a new problem and called downsampling unbalance problem in this paper. Due to the unbalance, MST-based downsampling method cannot be applied to construct graph signal multi-scale transforms. In this paper, we propose a novel and efficient method to detect and reduce the downsampling unbalance generated by the MST-based method. For any given graph signal, we apply the graph density to construct a measurement of the downsampling unbalance generated by the MST-based method. If a graph signal has large unbalance possibility, the multi-level downsampling is conducted after the MST is improved. The experimental results on synthetic and real-world social network data show that downsampling unbalance can be efficiently detected and then reduced by our method.
Xianwei Zheng, Yuan Yan Tang, Jiantao Zhou 0001, Jianjia Pan, Shouzhi Yang, Youfa Li, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.7
2018 A Fusion Strategy for the Single Shot Text Detector
abstract
In this paper, we propose a new fusion strategy for scene text detection. The system is based on a single fully convolution network, which outputs the coordinates of text bounding boxes at multiple scales. We improve the performance of text detection by combining a fusion strategy. This strategy obtains precise text bounding boxes according to the confidence of candidate text boxes. It exhibits promising robustness and discriminative power by fusing text boxes. Experimental results on ICDAR2011 and ICDAR2013 datasets indicate the effectiveness and robustness of the proposed fusion strategy with an F-measure of 87%, which outperforms the base network 2%.
Shujing Lyu, Yue Lu 0001, Patrick Shen-Pei Wang
ICPR4
2018 Granularity Approach for Multi-Criteria Decision Making About Hybrid Evaluation Information
abstract
In this study, a new version of TOPSIS method is reconstructed to deal with the problem of multi-criteria decision making. Here, the data representation of all alternatives is varied according to different criteria, such as real number, interval-valued number, set-valued number and intuitionistic fuzzy-valued number, etc. Because the distinguishing ability of each criterion can be reflected by its knowledge granularity, naturally, a knowledge granularity method is constructed to measure the criteria weights. Besides, the approach of how to select the ideal solution is redefined, especially for the case that the content of criterion according to all alternatives is not a totally ordered set anymore. What is more, the decision maker’s personal preference is considered, and the concrete indicator value can be calculated by the convex combination of the distance from possible alternatives to ideal solutions. Finally, the validity of the proposed decision-making algorithm is illustrated by a synthetic example.
Shihu Liu, Fusheng Yu, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.3
2018 A Hazy Image Database with Analysis of the Frequency Magnitude
abstract
Image haze removal has been extensively studied, but there has been no such an image database regarding the haze level. It is not convenient for readers to verify the assumptions or priors that are supposed to be useful for haze removal, and meanwhile, it is not fair to compare the performance of haze removal methods, which are effective for images with different haze levels. To solve this problem, we built a database consisting of more than 3464 images of different kinds of outdoor scenes. The images of the database are grouped into four classes regarding the haze level. Along with the database, we also observe a frequency magnitude prior, i.e. the frequency magnitude decreases with the increasing haze level, which can be used as a prior to develop haze removal methods. Our purpose is to help develop image haze removal methods, as well as verify existing statistical priors and discover new ones that can be used for image processing.
Shuhang Wang, Patrick Shen-Pei Wang, Petra Perner
Int. J. Pattern Recognit. Artif. Intell.4
2017 Chinese Handwriting Identification Method Based on Keyword Extraction
abstract
Text-independent handwriting identification methods require that features such as texture are extracted from lengthy document image; while text-dependent handwriting identification methods require that the contents of the documents being compared are identical. In order to overcome these confinements, this paper presents a novel Chinese handwriting identification technique. First, Chinese characters are segmented from handwriting document, then keywords are extracted based on matching and voting of local features of character. Then the same-content keywords are used to build training sets, and these training sets of two documents are compared. Because the keywords are similar to signature, the handwriting identification problem is transformed into signature verification problem. Experiments on HIT-MW, HIT-SW and CASIA show this method outperforms many text-independent handwriting identification methods.
Bin Fang 0001, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.3
2017 Empirical Mode Decomposition - Window Fractal (EMDWF) Algorithm in Classification of Fingerprint of Medicinal Herbs
abstract
This paper presents a new approach called the empirical mode decomposition — window fractal (EMDWF) algorithm in classification of fingerprint of medicinal herbs. In this way, we consider a glycyrrhiza fingerprint of medicinal herb as a signal sequence, and apply empirical mode decomposition (EMD) and Hiaguchis fractal dimension to construct a feature vector. By using EMD, the glycyrrhiza fingerprint of medicinal herb can be decomposed into some intrinsic mode functions (IMFs). As window fractal dimension (WFD) is applied to each IMF and original signal, the features of the glycyrrhiza fingerprint of medicinal herb can be obtained. Thereafter, SVM is applied as a classifier. The results of the experiments state clearly that the feature extracted by EMDWF is better than that of the existing methods including the pure EMD. With the increase of the number of training samples and the increase of the number of layers in EMD, the classification result achieves more stability.
Jianwei Du, Zhengguang Xu, Zhichun Mu, Patrick Shen-Pei Wang, Yuan Yan Tang, Huiwu Luo
Int. J. Pattern Recognit. Artif. Intell.4
2017 Rank Factor Granules with Fuzzy Collaborative Clustering and Factor Space Theory
abstract
This paper makes a discussion on the ranking problem of factor granules where each granule is composed by three parts: the patterns, the factors and the factor-induced information. Hereinto, the factor-induced information refers to the pattern’s attributes and the relationship between any two patterns. The overall ranking process is based on the ideology of fuzzy collaborative clustering, by considering a referential factor granule. The collaborative information, i.e. the partition matrices of factor granules, are used to collaborate the clustering for the referential factor granule. These collaborative information are obtained from different sources by different methods. Specially, one kind is obtained from the qualitative data by factor theory-based method. By comparing the difference of the referential factor granule before and after collaboration in aspect of clustering results, we can sort these factor granules: the little the difference, the closer to the top of the sequence.
Shihu Liu, Fusheng Yu, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.3
2017 Prediction of Television Audience Rating Based on Fuzzy Cognitive Maps with Forward Stepwise Regression
abstract
The television audience rating is an important indicator of the quality of television programs and important reference for decision-television operator. As many factors that affect the ratings and the trends are complex, the article proposes a television rating mining predictive model based on fuzzy cognitive maps (FCMs) with forward stepwise regression. The FCMs use the causal relationship among various concept nodes to simulate the fuzzy reasoning, and enhance the dynamic behavior of the simulation system with its feedback mechanism, which is suitable for system to predict the trend of television audience rating. A FCM-based model for predicting television audience rating is proposed in this paper. The forward stepwise regression algorithm is used to obtain concept nodes of coarse weight matrix for FCMs, and then a training weight algorithm is used to refine the coarse weight matrix model. The FCM model is applied to mine the television audience rating, realizing to predict the television playback volume. The experimental result shows that the modeling method is effective.
Nan Ma 0002, Patrick Shen-Pei Wang, Wenjia Li, Zhang Huan
Int. J. Pattern Recognit. Artif. Intell.2
2017 Off-line Text-Independent Writer Recognition: A Survey
abstract
Writer recognition is to identify a person on the basis of handwriting, and great progress has been achieved in the past decades. In this paper, we concentrate ourselves on the issue of off-line text-independent writer recognition by summarizing the state of the art methods from the perspectives of feature extraction and classification. We also exhibit some public datasets and compare the performance of the existing prominent methods. The comparison demonstrates that the performance of the methods based on frequency domain features decreases seriously when the number of writers becomes larger, and that spatial distribution features are superior to both frequency domain features and shape features in capturing the individual traits.
Yujie Xiong, Yue Lu 0001, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.3
2016 Multi-scale B-spline level set segmentation based on Gaussian kernel equalization
abstract
Images with weak contrast, overlapped noise and texture of the object and background make many PDE based methods disabled. To address these problems, this paper presents a novel combined multi-scale variational framework level set segmentation model. Its level set formulation consists edge-based term, region-based term and shape constraint term. The edge-based term is constructed using a newly defined edge stopping function. The region-based term is derived from parameter-free Gaussian probability density function (pdf) and multiple Gaussian kernel are used to gray equalization. The shape constraint term is used to constrain contour evolution at different scales of image pyramid. For an intrinsic smoothing segmentation contours, the level set function is explicitly represented by B-spline basis functions. Finally, a convolution is used during the energy minimization. Experimental results on synthetic and real images validate the robustness and high accuracy boundaries detection for low contrast, noise and texture images.
Shenhai Zheng, Bin Fang 0001, Patrick Shen-Pei Wang, Laquan Li, Mingqi Gao 0001
ICIP3
2016 Information-theoretic atomic representation for robust pattern classification
abstract
Representation-based classifiers (RCs) including sparse RC (SRC) have attracted intensive interest in pattern recognition in recent years. In our previous work, we have proposed a general framework called atomic representation-based classifier (ARC) including many popular RCs as special cases. Despite the empirical success, ARC and conventional RCs utilize the mean square error (MSE) criterion and assign the same weights to all entries of the test data, including both severely corrupted and clean ones. This makes ARC sensitive to the entries with large noise and outliers. In this work, we propose an information-theoretic ARC (ITARC) framework to alleviate such limitation of ARC. Using ITARC as a general platform, we develop three novel representation-based classifiers. The experiments on public real-world datasets demonstrate the efficacy of ITARC for robust pattern recognition.
Yulong Wang 0002, Yuan Yan Tang, Luoqing Li, Patrick Shen-Pei Wang
ICPR4
2016 Maximal level estimation and unbalance reduction for graph signal downsampling
abstract
The emerging field of graph signal processing requires a solid design of downsampling operation for graph signals to extend pattern recognition, machine learning and signal processing techniques into the graph setting. The state-of-the-art downsampling method is constructed upon the maximum spanning trees of the graphs. However, under the framework of this method, unbalanced downsampling often occurs for signals defined on densely connected unweighted graphs, such as social network data. The unbalance also significantly reduces the maximal downsampling level, making it smaller than the level we expect. In applications, the maximal level must be estimated to ensure that it is larger than the expected level; meanwhile, the unbalance has to be reduced, if it occurs. In this paper, we propose a novel method to jointly estimate the maximal level and reduce the downsampling unbalance. This method also offers an estimation of the possibility of unbalanced downsampling. If a graph signal is classified to be with high unbalance possibility, the maximum spanning tree will be updated to generate a balanced downsampling. The simulation results on synthesis and real world data support the theoretical analysis.
Xianwei Zheng, Yuan Yan Tang, Jiantao Zhou 0001, Patrick Shen-Pei Wang
ICPR4
2016 A hybrid swarm optimization for neural network training with application in stock price forecasting
abstract
A improved swarm optimization method based on particle swarm optimization (PSO) and simplified swarm optimization (SSO) is proposed to adjust the weight in artificial neural network. This method is a modification of traditional PSO and SSO, and combines them to a new optimization method (PSOSSO for short). The proposed method overcomes some of the drawbacks of SSO and improves its ability to train the weight of ANN. In the experiments, the PSOSSO is employed to train fuzzy wavelet neural network (FWNN) forecasting model to predict the prices of Hong Kong Hang Seng Index. The experimental results present that the PSOSSO is more efficient than traditional PSO and SSO methods.
Jianjia Pan, Yuan Yan Tang, Yulong Wang 0002, Xianwei Zheng, Huiwu Luo, Patrick Shen-Pei Wang
SMC7
2016 Road curve fitting by multi-resolution analysis
abstract
In this paper, we propose a new method for road curve fitting in urban environment based on multi-resolution analysis. The main technical contributions of the proposed method are the reconstructed approximation on the basis of the wavelet decomposition structure for curve fitting and the de-noising via the wavelet coefficients thresholding. The carried out experimental tests show promising results in a series of continuous driving images, validating our suggested method can fit the road curve effectively and computational efficiently.
Yuan Yan Tang, Patrick Shen-Pei Wang
SMC4
2016 Improving unbalanced downsampling via maximum spanning trees for graph signals
abstract
The state-of-the-art downsampling method for graph signals has been constructed by using maximum spanning trees (MSTs) of the graphs. For the graph signals defined on unweighted densely connected graphs, such as social network data, the sampling rates via MST-based downsampling are not close to 1/2, leading to a unbalanced downsampling phenomenon on multi-level downsampling. The unbalance hinders the applications of MST-based downsampling on constructing graph signal multiscale transforms, such as graph wavelet decomposition and multiscale pyramid transform. In this paper, we propose a simple but efficient method to improve the performance of the MST-based method on downsampling balance. For every graph signal, we first propose an unbalance possibility to measure the unbalance of the MST-based downsampling. If the unbalance possibility is high, the downsampling will be conducted on an improved MST, which is constructed by rearranging the structure of the MST to reduce the downsampling unbalance. The experiment results on synthesis graph signal show that the proposed improved MST leads to balanced downsampling. That is, the sampling rates produced by the improved MST are closer to 1/2 in multi-level downsampling than the original MST-based method.
Xianwei Zheng, Yuan Yan Tang, Jiantao Zhou 0001, Patrick Shen-Pei Wang
SMC4
2015 Gender Classification from Face Images Based on Gradient Directional Pattern (GDP)
Faisal Ahmed 0004, Padma Polash Paul, Patrick Shen-Pei Wang, Marina L. Gavrilova
ICCSA (2)3
2015 Text-independent writer identification using SIFT descriptor and contour-directional feature
abstract
This paper presents a method for text-independent writer identification using SIFT descriptor and contour-directional feature (CDF). The proposed method contains two stages. In the first stage, a codebook of local texture patterns is constructed by clustering a set of SIFT descriptors extracted from images. Using this codebook, the occurrence histograms are calculated to determine the similarities between different images. For each image, we obtain a candidate list of reference images. The next stage is to refine the candidate list using the contour-directional feature and SIFT descriptor. The proposed method is evaluated with two datasets: the ICFHR2012-Latin dataset and the ICDAR2013 dataset. Experimental results show that the proposed method outperforms the state-of-the-art algorithms and archives the best performance.
Yujie Xiong, Ying Wen 0003, Patrick Shen-Pei Wang, Yue Lu 0001
ICDAR3
2014 A General Framework for Manifold Reconstruction from Dimensionality Reduction
abstract
Recently, many dimensionality reduction (DR) algorithms have been developed, which are successfully applied to feature extraction and representation in pattern classification. However, many applications need to re-project the features to the original space. Unfortunately, most DR algorithms cannot perform reconstruction. Based on the manifold assumption, this paper proposes a General Manifold Reconstruction Framework (GMRF) to perform the reconstruction of the original data from the low dimensional DR results. Comparing with the existing reconstruction algorithms, the framework has two significant advantages. First, the proposed framework is independent of DR algorithm. That is to say, no matter what DR algorithm is used, the framework can recover the structure of the original data from the DR results. Second, the framework is space saving, which means it does not need to store any training sample after training. The storage space GMRF needed for reconstruction is far less than that of the training samples. Experiments on different dataset demonstrate that the framework performs well in the reconstruction.
Xian'en Qiu, Guo-Can Feng, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.4
2013 Filtering Terms from the Web for Image Annotations
abstract
In this paper, we propose a novel automatic image annotation model by mining the web. In our approach, the terms or words appearing in the associated text are extracted and filtered as labels or annotations for the corresponding web images. Sure, much noise exists in those selected labels. In order to reduce the influence caused by the noisy labels, for each label or potential word, we improve web image-word relationships using Mixture Gaussian Distribution Model. By doing so, the relationships between words and images are re-weighted both in terms of sematic relevance and in terms of visual feature similarity. In fact, all the words associated to an image are not semantically independent. We use co-occurrences between two words to describe their semantic relevance. Thus, we further use a method, called Word Promotion, to co-enhance the weights of all the words associated to a given image based on their co-occurrences. Our experiments are conducted in several ways and the results show that our annotation method can achieve a satisfactory performance in respects of system scalability and sematic evolution.
Zhiguo Gong, Jingzhi Guo, Yuan Yan Tang, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.4
2013 Kernel-Based 2D Fisher Discriminant Analysis with parameter Optimization for Face Recognition
abstract
As is known, kernel methods are developed to handle nonlinear classification. However, we have found that based on vector representation of face images, it is not easy to improve the performance of LDA-based methods by incorporating the kernel trick. Actually, in the case that the dimensionality of face images in vector form is far greater than the number of samples, linear classifiers are adequate and it is illogical to increase dimensions of vectors using kernel methods. In order to give full play to the strong point of the kernel technique for manipulating the nonlinearity of pattern distribution, we propose a new image feature extraction method for face recognition, called kernel-based two-dimensional Fisher discriminant analysis (K2DFDA), which deals with a face image directly as a matrix, instead of a stacked vector from rows or columns of the image. Moreover, we present a kernel parameter optimization scheme for K2DFDA, based on the maximum margin criterion and the damped Newton's method. Experimental results show the effectiveness of K2DFDA and its parameter optimization scheme.
Xiaozhang Liu, Patrick Shen-Pei Wang, Guo-Can Feng
Int. J. Pattern Recognit. Artif. Intell.2
2013 Automatic Multi-Scale Segmentation of Intrahepatic Vessel in CT Images for Liver Surgery Planning
abstract
The processing of blood vessels is an indispensable part in complicated surgeries of livers and hearts as the development of medical image technologies, which requires an automatic segmentation system over CT images of organs. However, the vascular pattern of livers in CT images suffers from low contrast to background so that the existing segmentation technologies are not able to extract the blood vessels completely. In the paper, we propose a new algorithm to extract the blood vessels of livers based on the adaptive multi-scale segmentation. First, we prove that the background histogram of normal scale blood vessels obeys the Gaussian distribution in CT images, and obtain the vascular distribution function from the vascular signal segmented from the background with a local optimal threshold. Second, Hessian matrix is employed to enhance the thin blood vessels before the extraction, and a complete and clear segmentation system for blood vessels is constructed by combining the major and thin blood vessels via filtering. Experimental results show the effectiveness of the proposed method, which is able to extract more complete blood vessels for 3D system, and assist the clinical liver surgeries efficiently.
Yi Wang 0074, Bin Fang 0001, Jingrui Pi, Patrick Shen-Pei Wang, Hongguang Wang
Int. J. Pattern Recognit. Artif. Intell.5
2013 A Fast and Complete Convex-Hull Algorithm Architecture Based on Ellipse and Elastic Ellipse Methods
abstract
The number of inner points excluded in an initial convex hull (ICH) is vital to the efficiency getting the convex hull (CH) in a planar point set. The maximum inscribed circle method proposed recently is effective to remove inner points in ICH. However, limited by density distribution of a planar point set, it does not always work well. Although the affine transformation method can be used, it is still hard to have a better performance. Furthermore, the algorithm mentioned above fails to deal with the exceptional distribution: the gravity centroid (GC) of a planar point set is outside or on the edge formed by the extreme points in ICH. This paper considers how to remove more inner points in ICH when GC is inside of ICH and completely process the case which mentioned above. Further, we presented a complete algorithm architecture: (1) using the ellipse and elasticity ellipse methods (EM and EEM) to remove more inner points in ICH and process the cases: GC is inside or outside of ICH. (2) Using the traditional methods to process the situation: the initial centroid is on the edge in ICH. It is adaptive to more data sets than other algorithms. The experiments under seven distributions show that the proposed method performs better than other traditional algorithms in saving time and space.
Xuegang Wu, Bin Fang 0001, Yuan Yan Tang, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.4
2012 Preface
Xiaoyi Jiang 0001, Matthew Y. Ma, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.3
2012 Fingerprint Enhancement Based on Wavelet and Anisotropic Filtering
abstract
The importance of high-fidelity enhancement in low quality fingerprint image cannot be overemphasized. Most of the existing fingerprint enhancement methods are contextual filter-based methods and they often suffer from two shortcomings: (1) there is block effect on the enhanced images; and (2) they blur or destroy ridge structures around singular points. In order to well preserve the ridge structures in singular regions and avoid block effect, we develop a new method for fingerprint enhancement combining nontensor product wavelet filter banks and anisotropic filter. We first decompose the fingerprint image using the nontensor product wavelet filter banks. Then we modify the approximation subimage using anisotropic filtering and adjust the high frequency coefficients of the three other subimages by applying the adaptive approach to reduce the noises according to the geometry feature of images. Finally, the inverse transform is applied to map the result and a final contrast enhancement is done subsequently. Experiments have been conducted on the fingerprint database FVC2004 in our study. The results demonstrate that the proposed approach is capable of overcoming block effect and enhancing low quality fingerprint while preserving the ridge structures around singular points.
Jiajia Lei, Qinmu Peng, Xinge You, Hiyam Hatem Jabbar, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.5
2012 Dynamic Management of Multiple Classifiers in Complex Recognition System
abstract
There are different kinds of multiple classifiers in complex recognition systems in pursuit of better recognition capabilities. To exploit the classifiers' potential as individual ones sufficiently and enable them to work cooperatively for the best classification results, they need to be considered as a whole and be dynamically managed according to the changing recognition occasions. In this paper, we present the conception of distributed Multiple Classifiers Management (MCM) and a self-adaptive recursive MCM model based on Mixture-of-Experts (ME). A control subsystem is consisted in the model, which allows the classification progress to be controlled by the systems' priori information when necessary. The model adjusts its parameters dynamically according to the current recognition state and gives the recognition results by combining the current individual classifiers' results with the previous combination result under priori information's control. An algorithm based on one step error correction is presented to acquire the model's parameters dynamically. It takes the previous times' ensemble classification results as true and corrects the current weights of the classifiers. At last, an experiment on the recognition of space objects is simulated. The experiment results show that the MCM model in this paper is effective for complex recognition system containing heterogeneous classifiers on improving the recognition rate and robustness.
Hui-Min Liu, Patrick Shen-Pei Wang, Hongqiang Wang 0001, Xiang Li 0014
Int. J. Pattern Recognit. Artif. Intell.2
2012 Cost-Sensitive Neural Network Classifiers for Postcode Recognition
abstract
Most traditional postcode recognition systems implicitly assumed that the distribution of the 10 numerals (0–9) is balanced. However it is far from a reasonable setting because the distribution of 0–9 in postcodes of a country or a city is generally imbalanced. Some numerals appear in more postcodes, while some others do not. In this paper, we study cost-sensitive neural network classifiers to address the class imbalance problem in postcode recognition. Four methods, namely: cost-sampling, cost-convergence, rate-adapting and threshold-moving are considered in training neural networks. Cost-sampling adjusts the distribution of the training data such that the costs of classes are conveyed explicitly by the appearances of their instances. Cost-convergence and rate-adapting are carried out in training phase by modifying the architecture of training algorithms of the neural network. Threshold-moving tries to increase the probability estimations of expensive classes to avoid the samples with higher costs to be misclassified. 10,702 postcode images are experimented using five cost matrices based on the distribution of numerals in postcodes. The results suggest that cost-sensitive learning is indeed effective on class imbalanced postcode analysis and recognition. It also reveals that cost-sampling on a proper cost matrix outperforms others in this application.
Shujing Lu, Li Liu 0010, Yue Lu 0001, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.4
2012 A Novel Supervised Structure Dictionary Learning for Classification Based on Sparse Representation
abstract
Sparse representation based classification has led to interesting image recognition results, while the dictionary used for sparse coding plays a key role in it. This paper presents a novel supervised structure dictionary learning (SSDL) algorithm to learn a discriminative and block structure dictionary. We associate label information with each dictionary item and make each class-specific sub-dictionary in the whole structured dictionary have good representation ability to the training samples from the associated class. More specifically, we learn a structured dictionary and a multiclass classifier simultaneously. Adding an inhomogeneous representation term to the objective function and considering the independence of the class-specific sub-dictionaries improve the discrimination capabilities of the sparse coordinates. An iteratively optimization method be proposed to solving the new formulation. Experimental results on four face databases demonstrate that our algorithm outperforms recently proposed competing sparse coding methods.
Patrick Shen-Pei Wang, Guo-Can Feng
Int. J. Pattern Recognit. Artif. Intell.2
2012 A Review of Wavelet-Based Edge Detection Methods
abstract
Edges are prominent features in images. The detection and analysis of edges are key issues in image processing, computer vision and pattern recognition. Wavelet provides a powerful tool to analyze the local regularity of signals. Wavelet transform has been successfully applied to the analysis and detection of edges. A great number of wavelet-based edge detection methods have been proposed over the past years. The objective of this paper is to give a brief review of these methods, and encourage the research of this topic. In practice, an image is usually of multistructure edge, the identification of different edges, such as steps, curves and junctions play an important role in pattern recognition. In this paper, more attention is paid on the identification of different types of edges. We present the main idea and the properties of these methods.
Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.1
2012 Similarity Learning Based on Semi-Supervised Graph for Classification
abstract
Similarity measurement is crucial for classification. Based on the manifold assumption, many graph-based algorithms were developed. Almost all methods follow the k-rule or ε-rule to construct a graph, and then focus on the algorithms based on the graph. However, the graph may not represent the local structure well, and it does not fully utilize the label information yet. The local structure can be presented by the local density and the distance between the samples and their neighbors. And the graph constructed by the guidance of label information will be better approximate of the relationship of the input data. In this paper, we propose an adaptive semi-supervised graph constructing method. The similarity is learned when constructing the graph. The advantages of the similarity learned by our method include: (1) The similarity is measured along the manifold by constructing a graph; (2) nearby points and points in the same cluster share high similarity; (3) samples from the same class have higher similarity than samples from different classes. Experimental results show that using the proposed similarity for classification task could get better recognition accuracy.
Qianying Wang 0001, Pong C. Yuen, Guo-Can Feng, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.4
2011 A Multi-Layer Contrast Analysis Method for Texture Classification Based on LBP
abstract
Texture classification is one of the important fields in pattern recognition and machine vision research. LBP method,13–15 proposed by Ojala, can be used to classify texture images effectively. And the LBP method has rotation-invariant, illumination-invariant, multi-resolution characteristics. But, since the contrast is not considered between neighbor pixels, the correct classification rate produced by this method has been remarkably influenced by light source type and light source orientation. The LMLCP (Local Multiple Layer Contrast Pattern) method, proposed by this paper, maps the contrast value between two near pixels to a rank value, which represent a relative contrast value range, and computes the statistic histogram referring to the work in LBP method. The LMLCP method can bring out the rapid expansion of feature dimension, so a special feature encoding method used in 3DLBP6 is adopted by this paper. The experiment, which is built based on Outex_TC_00012,12 demonstrates that the LMLCP can evidently make a more accurate classification rate than LBP method.
Hengxin Chen, Yuan Yan Tang, Bin Fang 0001, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.4
2011 Bionic Face Recognition Using Gabor Transformation
abstract
In this paper, we propose a bionic face recognition method based on Gabor feature. First, Gabor features are extracted from face images, followed by dimensionality reduction using 2DPCA algorithm, which serves as the feature vectors of the proposed method. Finally, the bionic classifier is trained for classification. The experiment on AR and PIE face database is reported to show the effectiveness of the proposed method and compare it with Gabor-2DPCA algorithm and Gabor-PCA algorithm.
Yuan Yan Tang, Bin Fang 0001, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.4
2010 Rotation Invariant Multiview Face Detection Using Skin Color Regressive Model and Support Vector Regression
abstract
In this paper, an automatic rotation invariant multiview face detection method, which utilizes modified Skin Color Model (SCM), is presented. First, Gaussian Mixture Model (GMM) and Support Vector Machine (SVM) based hybrid models are used to classify human skin regions from color images. The novelty of the adaptive hybrid model is its ability to predict the chromatic skin color band for individual images based on calibration differences of camera and luminance condition of environment. Classified skin regions are then converted to gray scale image with a threshold based on the predicted chromatic skin color bands, which further enhances detection performance. Next, Principle Component Analysis (PCA) is applied to gray segmented regions. Face detection is carried out based on the PCA-based extracted features, along with selected features, using support vector regression. The output of this procedure is used to report the final result of face detection. The proposed method is also beneficial for the rotation invariant face recognition problem.
Padma Polash Paul, Md. Maruf Monwar, Marina L. Gavrilova, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.4
2009 Editorial
Yiu-Ming Cheung, Yuping Wang 0003, Xinge You, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.4
2009 Combining Eodh and Directional Gradient Density for Offline Signature Verification
abstract
The main problem to identify skilled forgeries for offline signature verification lies in the fact that it is difficult to formalize distinguished feature representation of the signature patterns and design appropriate fusion scheme for various types of feature vectors. To tackle these problems, in this paper, we propose an approach to extract robust Edge Orientation Distance Histogram (EODH) descriptor which effectively reflects signature structure variations. In addition, directional gradient density features are employed for skilled forgery verification attempt. To exploit the full capacity of two sets of features, we designed the multilevel weighted fuzzy classifier and fuse match scores by way of selection priority. Experiments were conducted on a subcorpus of open MCYT signature database which is widely used for performance evaluation. It shows that the proposed method was able to improve verification accuracy.
Bin Fang 0001, Yuan Yan Tang, Patrick Shen-Pei Wang, Taiping Zhang
Int. J. Pattern Recognit. Artif. Intell.4
2009 Facial Biometrics Using Nontensor Product Wavelet and 2D Discriminant Techniques
abstract
A new facial biometric scheme is proposed in this paper. Three steps are included. First, a new nontensor product bivariate wavelet is utilized to get different facial frequency components. Then a modified 2D linear discriminant technique (M2DLD) is applied on these frequency components to enhance the discrimination of the facial features. Finally, support vector machine (SVM) is adopted for classification. Compared with the traditional tensor product wavelet, the new nontensor product wavelet can detect more singular facial features in the high-frequency components. Earlier studies show that the high-frequency components are sensitive to facial expression variations and minor occlusions, while the low-frequency component is sensitive to illumination changes. Therefore, there are two advantages of using the new nontensor product wavelet compared with the traditional tensor product one. First, the low-frequency component is more robust to the expression variations and minor occlusions, which indicates that it is more efficient in facial feature representation. Second, the corresponding high-frequency components are more robust to the illumination changes, subsequently it is more powerful for classification as well. The application of the M2DLD on these wavelet frequency components enhances the discrimination of the facial features while reducing the feature vectors dimension a lot. The experimental results on the AR database and the PIE database verified the efficiency of the proposed method.
Dan Zhang 0008, Xinge You, Patrick Shen-Pei Wang, Svetlana N. Yanushkevich, Yuan Yan Tang
Int. J. Pattern Recognit. Artif. Intell.3
2008 Metric Learning: A general dimension reduction framework for classification and visualization
abstract
A new general dimension reduction framework based on similar and dissimilar metric learning is proposed in this paper which allows us to exploit the geometry of data to reduce the data dimension for classification and visualization. The general formulation can unify the existing dimension reduction algorithms within a common framework. Furthermore, this metric learning framework can be used as a general platform for developing new dimension reduction algorithms. By utilizing this framework as a tool, we propose a novel supervised dimension reduction algorithm named sub-manifold preserving analysis (SMPA) in which the intrinsic sub-manifold structure will be preserved while the margin of interclass will be separated. Experimental evidences show that performance of our proposed SMPA algorithm is better than other algorithms.
Chunyuan Lu, Guo-Can Feng, Jianmin Jiang, Patrick Shen-Pei Wang
ICPR4
2008 A scene-based video watermarking technique using SVMs
abstract
In this paper we present a scene-based video watermarking scheme using support vector machines (SVMs). In a given scene, the algorithm uses the first h′ frames to train an embedding SVM, and uses this SVM to watermark the rest of frames. In the extracting phrase, the detector uses center h frames of the first h′ frames to train an extracting SVM. The final extracted watermark in a given scene is the average of watermarks extracted from the remaining frames. Watermarks are embedded in l longest scenes of a video such that it is time efficient and capable to resist possible frames swapping/deleting /duplicating attacks. Collusion attacks on a watermarked video are examined. The proposed algorithm is shown to be robust to compression and collusion attacks, and it has novelty and practicability on SVM-applications.
Shwu-Huey Yen, Hsiao-Wei Chang, Chia-Jen Wang, Patrick Shen-Pei Wang, Mei-Chueh Chang
ICPR4
2008 Intelligent pattern recognition and biometrics
abstract
This talk deals with advanced concepts of artificial intelligence (AI) and pattern recognition (PR), and their applications to solving real life problems including biometrics applications. It basically covers the following topics: (1) Overview of pattern recognition (PR),(2) Overview of artificial intelligence (AI),(3) the relation between PR and AI, (4) Analysis and learning: pattern recognition concept : foundation and theories, (5) Importance of ambiguity, and its applications: theory and applications, (6) An overall Interactive intelligent pattern recognition (IPR) system, (7) Concepts of syntax, semantics, and pragmatics: theories and applications, (8) Importance of ambiguity, and semantics: graphics and its applications, (9) How it works: IPR and applications to solving real life problems including biometrics and face recognition, (10)some more Illustrations, discussions and future directions.
Patrick Shen-Pei Wang
ISI1
2008 Classifier Combination and its Application in Iris Recognition
abstract
Classifier combination is an effective method to improve the recognition accuracy of a biometric system. It has been applied to many practical biometric systems and achieved excellent performance. However, there is little literature involving theoretical analysis on the effectiveness of classifier combination. In this paper, we investigate classifiers combined with the max and min rules. In particular, we compute the recognition performance of each combined classifier, and illustrate the condition in which the combined classifier outperforms the original unimodal classifier. We focus our study on personal verification, where the input pattern is classified into one of two categories, the genuine or the impostor. For simplicity, we further assume that the matching score produced by the original classifier follows a normal distribution and the outputs of different classifiers are independent and identically distributed. Randomly-generated data are employed to test our conclusion. The influence of finite samples is explored at the same time. Moreover, an iris recognition system, which adopts multiple snapshots to identify a subject, is introduced as a practical application of the above discussions.
Xinhua Feng, Xiaoqing Ding, Youshou Wu, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.4
2008 Noniterative 3D Face Reconstruction Based on Photometric Stereo
abstract
3D face reconstruction is a popular area within the computer vision domain. 3D face reconstruction should ideally be achieved easily and cost-effectively, without requiring specialized equipment to estimate 3D shapes. As a result of this, many techniques for retrieving 3D shapes from 2D images have been proposed. In this paper, a novel method for 3D face reconstruction based on photometric stereo, which estimates the surface normal from shading information in multiple images, hence recovering the 3D shape of a face, is proposed. In order to overcome the problems of previous approaches related to prior-knowledge regarding lighting conditions and iterative algorithms, the exemplar is synthesized with known lighting conditions from at least three images, under arbitrary lighting conditions and using an illumination reference. Experiments in 3D face reconstruction were made by verifying the proposed approach using the illumination subset of the Max-Planck Institute face database and Yale face database B. Experimental results demonstrate that the proposed method is effective for 3D shape reconstruction of faces from 2D images.
Patrick Shen-Pei Wang, Svetlana N. Yanushkevich, Seong-Whan Lee
Int. J. Pattern Recognit. Artif. Intell.2
2008 Facial Metamorphosis Using Geometrical Methods for Biometric Applications
abstract
Facial expression modeling has been a popular topic in biometrics for many years. One of the emerging recent trends is capturing subtle details such as wrinkles, creases and minor imperfections that are highly important for biometric modeling as well as matching. In this paper, we suggest a novel approach to the problem of expression modeling and morphing based on a geometry-based paradigm. In 2D image space, a distance-based morphing system is utilized to create a line drawing style facial animation from two input images representing frontal and profile views of the face. Aging wrinkles and expression lines are extracted and mapped back to the synthesized facial NPR (nonphotorealistic) sketches. In 3D object space, we present a metamorphosis system that combines the traditional free-form deformation (FFD) model with data interpolation techniques based on the proximity preserving Voronoi diagram. With feature points selected from two images of the target face, the proposed system generates the 3D target facial model by transforming a generic model. Experimental results demonstrate that morphing sequences generated by our systems are of convincing quality.
Marina L. Gavrilova, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.3
2008 Extracting Faces and Facial Features from Color Images
abstract
In this paper, we present image processing and pattern recognition techniques to extract human faces and facial features from color images. First, we segment a color image into skin and non-skin regions by a Gaussian skin-color model. Then, we apply mathematical morphology and region filling techniques for noise removal and hole filling. We determine whether a skin region is a face candidate by its size and shape. Principle component analysis (PCA) is used to verify face candidates. We create an ellipse model to locate eyes and mouths areas roughly, and apply the support vector machine (SVM) to classify them. Finally, we develop knowledge rules to verify eyes. Experimental results show that our algorithm achieves the accuracy rate of 96.7% in face detection and 90.0% in facial feature extraction.
Frank Y. Shih, Shouxian Cheng, Chao-Fa Chuang, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.4
2008 Performance Comparisons of Facial Expression Recognition in Jaffe Database
abstract
Facial expression provides an important behavioral measure for studies of emotion, cognitive processes, and social interaction. Facial expression recognition has recently become a promising research area. Its applications include human-computer interfaces, human emotion analysis, and medical care and cure. In this paper, we investigate various feature representation and expression classification schemes to recognize seven different facial expressions, such as happy, neutral, angry, disgust, sad, fear and surprise, in the JAFFE database. Experimental results show that the method of combining 2D-LDA (Linear Discriminant Analysis) and SVM (Support Vector Machine) outperforms others. The recognition rate of this method is 95.71% by using leave-one-out strategy and 94.13% by using cross-validation strategy. It takes only 0.0357 second to process one image of size 256 × 256.
Frank Y. Shih, Chao-Fa Chuang, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.3
2008 A Multimedia Watermarking Technique Based on SVMS
abstract
In this paper we present an improved support vector machines (SVMs) watermarking system for still images and video sequences. By a thorough study on feature selection for training SVM, the proposed system shows significant improvements on computation efficiency and robustness to various attacks. The improved algorithm is extended to be a scene-based video watermarking technique. In a given scene, the algorithm uses the first h' frames to train an embedding SVM, and uses the trained SVM to watermark the rest of the frames. In the extracting phrase, the detector uses only the center h frames of the first h' frames to train an extracting SVM. The final extracted watermark in a given scene is the average of watermarks extracted from the remaining frames. Watermarks are embedded in l longest scenes of a video such that it is computationally efficient and capable to resist possible frames swapping/deleting/duplicating attacks. Two collusion attacks, namely temporal frame averaging and watermark estimation remodulation, on video watermarking are discussed and examined. The proposed video watermarking algorithm is shown to be robust to compression and collusion attacks, and it is novel and practical for SVM-applications.
Chia-Jen Wang, Shwu-Huey Yen, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.3
2008 Editorial
Svetlana N. Yanushkevich, David Hurley, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.3
2007 Biometrics Intelligence Information Systems and Applications
abstract
This cutting-edge research lecture deals with some fundamental aspects of biometrics and its applications. It basically includes the following subtopics: (1) overview of biometric technology and applications (2) importance of security: a scenario of terrorists attack, (3) what are biometric technologies?, (4) biometrics: analysis vs. synthesis. (5)analysis: pattern recognition concept, (6) how it works: fingerprint extraction and matching, iris, and facial analysis, (7) authentication applications, (8)thermal Imaging: (9) emotion recognition. (10) synthesis in biometrics, (11) modeling and simulation. (12) examples and applications. The cutting-edge research lecture will also show broad methodology developments and applications in bioinformatics.
Patrick Shen-Pei Wang
BIBE1
2006 Content-based Image Retrieval Trained by Adaboost for Mobile Application
abstract
This paper proposes a Content-Based Image Retrieval (CBIR) system applicable in mobile devices. Due to the fact that different queries to a content-based image retrieval (CBIR) system emphasize different subsets of a large collection of features, most CBIR systems using only a few features are therefore only suitable for retrieving certain types of images. In this research we combine a wide range of features, including edge information, texture energy, and the HSV color distributions, forming a feature space of up to 1053 dimensions, in which the system can search for features most desired by the user. Through a training process using the AdaBoost algorithm9 our system can efficiently search for important features in a large set of features, as indicated by the user, and effectively retrieve the images according to these features. The characteristics of the system meet the requirements of mobile devices for performing image retrieval. The experimental results show that the performance of the proposed system is sufficiently applicable for mobile devices to retrieve images from a huge database.
Hwei-Jen Lin, Yang-Ta Kao, Fu-Wen Yang, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.4
2006 Editorial
Matthew Y. Ma, Jinhong Katherine Guo, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.3
2004 From Pixels To True XML Structures In Digital Document Images
abstract
XML has been widely used as metadata for image retrieval. As a standard, it makes it easier to index and retrieve information across different platforms. However, how to automatically convert an image into XML format remains a challenge. In this paper, a system for generating structured document in XML from digitally captured document images is presented. The system is aimed at providing an easy to use tool for average users without requiring depth of knowledge in the document processing areas. Further, a XML/XSL generator is developed to accurately represent a document in a XML structure, yet in a representation that reflects its original layout.
Matthew Y. Ma, Jinhong Katherine Guo, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.3
2001 3D Object Recognition and Visualization on the Web
Patrick Shen-Pei Wang
Web Intelligence1
2001 Intelligent Agent Technology - Introduction
Jiming Liu 0001, Ning Zhong 0001, Yuan Yan Tang, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.4
2000 3D Artificial Objects Recognition under Virtual Environment
abstract
Presents a method for visualization, understanding and recognition of artificial objects. The method using linear combination is simple and needs only a few learning samples. Furthermore, it can strengthen the advantages of conventional methods while overcoming their drawbacks. Also, it is able to distinguish objects with very similar patterns and is more accurate than other conventional methods in the literature. Four experiments are included to demonstrate the method's simplicity and accuracy.
Min Yi, Patrick Shen-Pei Wang
ICPR2
2000 A New Method of Color Image Segmentation Based on Intensity and Hue Clustering
abstract
A new method of color image segmentation is proposed. It is based on the K-means algorithm in HSI color space and has the advantage over those based on the RGB space. Both the hue and the intensity components are fully utilized. In the process of hue clustering, the special cyclic property of the hue component is taken into consideration. The paper gives the definition of the distance and the center in the hue space, based on which the hue-clustering algorithm is implemented. Utilised in medical image processing, the new method gives a good performance.
Patrick Shen-Pei Wang
ICPR2
2000 3D Articulated Object Understanding, Learning, and Recognition from 2D Images
abstract
This paper is aimed at 3D object understanding from 2D images, including articulated objects in active vision environment, using interactive, and internet virtual reality techniques. Generally speaking, an articulated object can be divided into two portions: main rigid portion and articulated portion. It is more complicated that "rigid" object in that the relative positions, shapes or angles between the main portion and the articulated portion have essentially infinite variations, in addition to the infinite variations of each individual rigid portions due to orientations, rotations and topological transformations. A new method generalized from linear combination is employed to investigate such problems. It uses very few learning samples, and can describe, understand, and recognize 3D articulated objects while the objects status is being changed in an active vision environment.
Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.1
1999 A Simple and Robust Thinning Algorithm
abstract
In this paper a new thinning algorithm is presented. It is simple yet robust, and can maintain the same advantages of other key skeletonization algorithms. In this new algorithm, only a very small set of rules of criteria for deleting pixels is used. It is faster and easier to implement. Its advantages and limitations are discussed and compared with others. Several examples are illustrated.
C. L. Lee, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.2
1999 Parallel Matching of 3D Articulated Object Recognition
abstract
In dealing with large volume image data, sequential methods usually are too slow and unsatisfactory. This paper introduces a new system employing parallel matching in high-level recognition of 3D articulated objects. A new structural strategy using linear combination and parallel graphic matching techniques is presented for 3D polyhedral objects representable by 2D line-drawings. It solves one of the basic concerns in diffusion tomography complexities, i.e. patterns can be reconstructed through fewer projections, and 3D objects can be recognized by a few learning sample views. It also improves some of the current methods while overcoming their drawbacks. Furthermore, it can distinguish very similar objects and is more accurate than other methods in the literature. An online webpage system for understanding and recognizing 3D objects is also illustrated.
Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.1
1999 A knowledge-based segmentation algorithm for enhanced recognition of handwritten courtesy amounts
Karim Hussein, Arun Agarwal, Amar Gupta, Patrick Shen-Pei Wang
Pattern Recognit.4
1998 Ink Matching of Cursive Chinese Handwritten Annotations
abstract
In this paper, we discuss the notion of treating electronic ink as first class data without attempting to recognize it by presenting two different variations of approximate ink matching (AIM) for searching ink data. We also illustrate a pen-based electronic document annotating and browsing system and methods for searching handdrawn personal notes employing the described matching schemes. Adapting from the Learning by Knowledge paradigm, we propose a semantic matching network that applies semantics of Chinese language early in the process of ink matching. Finally we evaluate several key components in our entire ink matching network via experiments. Preliminary experimental results show the approximate ink matching algorithms perform well, despite the informal and highly variable nature of Chinese handwriting. Our experiments also show some promising results on semantic matching and the feasibility of our semantic matching architecture.
Daniel P. Lopresti, Matthew Y. Ma, Patrick Shen-Pei Wang, Jill D. Crisman
Int. J. Pattern Recognit. Artif. Intell.3
1998 Intriguing Aspects of Oriental Languages
abstract
This paper includes a description of 3 affiliated oriental languages: Chinese, Japanese, and Korean. It includes a description of the origins of these 3 languages and the inter-relationship among them. Drawn from the viewpoints of several experienced researchers in the field of OCR (Optical Character Recognition) and computational linguistics, it attempts to bring out the intriguing aspects of these 3 ideographic languages, including the formation and composition of pictograms, special features, learning, understanding, contextual information, and recognition of characters and words, and their relations to poetic expressions and pattern recognition techniques. Numerous references are given and comments on future trends are also presented.
Ching Y. Suen, Shunji Mori, Hae-Chang Rim, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.4
1996 A parallel thinning algorithm with two-subiteration that generates one-pixel-wide skeletons
abstract
Many algorithms for vectorization by thinning have been devised and applied to a great variety of pictures and drawings for data compression, pattern recognition and raster-to-vector conversion. But parallel thinning algorithms which generate one-pixel-wide skeletons can have difficulty preserving the connectivity of an image. In this paper, we propose a 2-subiteration parallel thinning algorithm with template matching (PTA2T) which preserves image connectivity, produces thinner results, maintains very fast speed and generates one-pixel-wide skeletons.
Edward Y. Y. Zhang, Patrick Shen-Pei Wang
ICPR2
1995 Detection of courtesy amount block on bank checks
abstract
This paper discusses a technique for locating the courtesy amount block on bank checks. In the analysis and recognition process, connected components in the image are identified first. Then, strings are constructed on the basis of proximity and horizontal alignment of characters. Next, a set of rules and heuristics are applied to these strings to choose the correct one. The chosen string is only accepted if it passes a verification test, which includes an attempt to recognize the currency sign. A deterministic finite automaton system is then used for segmenting the handprinted courtesy amount. Finally, the separated components are passed on to a neural network based recognition system.
Arun Agarwal, Len Granowetter, Karim Hussein, Amar Gupta, Patrick Shen-Pei Wang
ICDAR5
1995 Perception and visualization of line images
abstract
This paper deals with state-of-the-art novel ideas of visualization, understanding and interpretation of two-dimensional (2D) line images. A new strategy using fast two-pass parallel pattern matching techniques is presented. It can learn, represent, visualize, and interpret 2D polyhedral line images with only very few learning samples. It can also distinguish similar patterns more accurately than other methods. Such a technique is applied to articulated object recognition and algorithms for articulated feature extraction are presented. Several illustrations are given and future research topics are discussed.
Patrick Shen-Pei Wang
ICIP (3)1
1995 Analysis and Design of Parallel Thinning Algorithms - A Generic Approach
abstract
In this paper, a generic approach for the analysis and design of parallel thinning algorithms is presented. The procedures for designing 4-subcycle/iteration, 2-subcycle/iteration, and 1-subcycle/iteration thinning algorithms are developed. The experimental results show that these thinning algorithms designed by proposed approach preserve image connectivity, produce thinner results, and obtain faster speed than many existing key thinning algorithms. As an example, a 1-subcycle/iteration algorithm and its experimental results were illustrated, which is in general about two times faster than others in the literature.
Edward Y. Y. Zhang, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.2
1994 A new thinning algorithm
abstract
This paper introduces a new thinning algorithm which is simple yet can maintain the same advantages of other key algorithms. In this new algorithm, only a very small set of rules of criteria for deleting pixels is used. It is faster and easier to implement. Several examples are illustrated.
C. L. Lee, Patrick Shen-Pei Wang
ICPR (1)2
1994 An Adaptive Modular Neural Network with Application to Unconstrained Character Recognition
abstract
The topology and the capacity of a traditional multilayer neural system, as measured by the number of connections in the network, has surprisingly little impact on its generalization ability. This paper presents a new adaptive modular network that offers superior generalization capability. The new network provides significant fault tolerance, quick adaption to novel inputs, and high recognition accuracy. We demonstrate this paradigm on recognition of unconstrained handwritten characters.
Lik Mui, Arun Agarwal, Amar Gupta, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.4
1994 Three-Dimensional Sequential/Parallel Universal Array Grammars for Polyhedral Object Pattern Analysis
abstract
We introduce a sequential/parallel parsing algorithm for analyzing 3-dimensional polyhedral objects represented by 3-d universal array grammars. The mechanism serves as a compromise between purely sequential methods, which normally take too much time, and purely parallel methods, which take too much hardware for large digital arrays. Several examples of 3-d polyhedral objects of various shapes and their corresponding parsing sequences are illustrated. Future research topics are discussed.
Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.1
1994 A New Parallel Thinning Methodology
abstract
A perfectly parallel thinning algorithm (PPTA) is proposed. It can generate perfect skeletons, which consist of end points, break points, and hole points only. Experimental results show that the proposed PPTA can also preserve image connectivity, produce thinner skeletons, and is faster than many existing thinning algorithms. For example, it is twice as fast as one of the fastest parallel thinning algorithm by Holt, Stewart, Clint and Perrorte.
Edward Y. Y. Zhang, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.2
1994 3D Line Image Analysis - A Heuristic Parallel Approach with Learning and Recognition
Patrick Shen-Pei Wang
Inf. Sci.1
1993 Machine visualization, understanding and interpretation of polyhedral line-drawings in document analysis
abstract
The author deals with high level visualization, understanding, and interpretation of three-dimensional (3-D) polyhedral objects from two-dimensional (2-D) line-drawing images for document analysis. A new scheme is presented, which is aimed at learning, representing and recognizing 3-D objects with only very few learning samples. It can strengthen advantages of some current key methods while overcome their drawbacks. Further, it will be able to distinguish objects with very similar patterns and is more accurate than other existing methods. Several illustrations are given. Future research topics are discussed.>
Patrick Shen-Pei Wang
ICDAR1
1993 A Thinning Algorithm Based on the Force Between Charged Particles
abstract
A new thinning algorithm based on the well known concept of the force of attraction or repulsion between charged particles is presented. This algorithm generates connected skeletons which preserve the shape and end-points of the original patterns. Its performance is experimentally compared with four other known algorithms published in the literature. For the sake of comparison, a reasoned set of test data is introduced. The results of our comparison reveal that the proposed CPM (Charge Particle Method) algorithm is almost as fast as the fastest of those compared. For thin images obtainable from low resolution scanners or coarse scanned images, the CPM algorithm is the fastest.
Akila Arumugam, Thiruvengadam Radhakrishnan, Ching Y. Suen, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.4
1993 An Integrated Architecture for Recognition of Totally Unconstrained Handwritten Numerals
abstract
A multi-staged system for off-line handwritten numeral recognition is presented here. After scanning, the digitized binary bitmap image of the source document is passed through a preprocessing stage which performs segmentation, thinning and rethickening, normalization, and slant correction. The recognizer is a three-layered neural net trained with back-propagation algorithm. While a few systems that use three-layered nets for recognition have been presented in the literature, the contribution of our system is based on two aspects: elaborate preprocessing based on structural pattern recognition methods combined with a neural net based recognizer; and integration of neural net based and structural pattern recognition methods to produce high accuracies.
Amar Gupta, M. V. Nagendraprasad, Patrick Shen-Pei Wang, S. Ayyadurai
Int. J. Pattern Recognit. Artif. Intell.4
1993 Analytical Comparison of Thinning Algorithms
abstract
This paper deals with analyzing and comparing several key thinning algorithms in terms of various different methodologies. From these analyses and comparisons, a new sequential model thinning algorithm using heuristic, hybrid methods is presented. It intends to produce, based on past experience, thinned skeletons of thickness one from input patterns, keep connectivity, and eliminate unnecessary pixels. Several illustrative examples on various patterns were tested and the results compared. The parallel model thinning algorithm is still open.
Edward Y. Y. Zhang, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.2
1992 Character segmentation techniques for handwritten text-a survey
abstract
The paper is a survey of techniques for segmenting images of handwritten text into individual characters. The topic is broken into two categories: segmentation and segmentation-recognition techniques. Several approaches to each are outlined, and each is analyzed for its relevance to printed, cursive, on-line and off-line input data.>
Christopher E. Dunn, Patrick Shen-Pei Wang
ICPR (2)2
1992 An improved algorithm for thinning binary digital patterns
abstract
A number of image processing and pattern recognition applications demand that a raw digitized binary pattern array be normalized, so that the constituent components of that array are of uniform thickness. The thinning process reduces such components to a thickness of one pixel, or sometimes a few pixels. This paper describes a strategy for farther enhancing the operational speed of one of the fastest parallel algorithms proposed in the literature.>
M. V. Nagendraprasad, Patrick Shen-Pei Wang, Amar Gupta
ICPR (3)2
1992 Analysis of thinning algorithms
abstract
Uses a combinatorial approach to analyse the differences of thinning algorithms among iteration and sub-iteration, thinning by coordinate and by edges of patterns. The authors also propose a serial model thinning algorithm in which the skeletons only consist of three types of points: connected points, end-points, and hole-points.>
Edward Y. Y. Zhang, Patrick Shen-Pei Wang
ICPR (3)2
1992 Three-Dimensional Object Pattern Representation by Array Grammars
abstract
A formal model for three-dimensional object representation is introduced. It uses parallel techniques and significantly reduces the time required for dealing with three-dimensional image analysis problems. Its fundamental properties are investigated and several interesting examples are illustrated.
Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.1
1991 An Improved Structural Approach for Automated Recognition of handprinted Characters
abstract
This paper examines several line-drawing pattern recognition methods for handwritten character recognition. They are the picture descriptive language (PDL), Berthod and Maroy (BM), extended Freeman's chain code (EFC), error transformation (ET), tree grammar (TG), and array grammar (AG) methods. A new character recognition scheme that uses improved extended octal codes as primitives is introduced. This scheme offers the advantages of handling flexible sizes, orientations, and variations, the need for fewer learning samples, and lower degree of ambiguity. Finally, the simulation of off-line character recognition by the real-time on-line counterpart is investigated.
Patrick Shen-Pei Wang, Amar Gupta
Int. J. Pattern Recognit. Artif. Intell.1
1989 Pushdown recognizers for Array Pattern
abstract
We investigate the factors that make it difficult to generalize pushdown automata for one-dimensional strings to two-dimensional arrays. Then we resolve the problems and construct two-dimensional pushdown array automata (PDAA). The relationship between isometric context-free array languages and pushdown array automata is established. Several examples of array automata are presented, and a pushdown array automaton is tested on VAX8650/VMS using PASCAL.
Hwei-Jen Lin, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.2
1989 A Fast and Flexible Thinning Algorithm
abstract
A fast serial and parallel algorithm for thinning digital patterns is presented. The processing speed is faster than previous algorithms in that it reads pixels along the edge of the input pattern rather than all pixels in each iteration. Using this algorithm, an experiment is conducted and the patterns such as 'X', 'H', 'A', 'moving body', and 'leaf' are tested. The results show that this algorithm is faster, structure-preserving, and more flexible in that it can be done either sequentially or in parallel.>
Patrick Shen-Pei Wang, Edward Y. Y. Zhang
IEEE Trans. Computers1
1988 A maximum algorithm for thinning digital patterns
abstract
An approach to thinning is presented that deletes pixels without using iterative transformations. It consists of four steps: computing an addition matrix, which assigns maximal values to the pixels on the medial axis; computing a comparative matrix; deleting the nonmaximum pixels; and deleting the endpoints. Experimental results show that this algorithm makes the skeleton closer to the medial axis and makes it more convenient to reconstruct the original pattern.>
Edward Y. Y. Zhang, Patrick Shen-Pei Wang
ICPR2
1988 A modified parallel thinning algorithm
abstract
A parallel thinning algorithm of C.M. Holt et al. (1987) is compared with an algorithm of D. Rutovitz (1966) and one by T.Y. Zhang and C.Y. Suen (1984). Analyses and experiments show that the Holt algorithm is similar to the Rutovitz algorithm. A heuristic modification to Rutovitz' algorithm is also proposed and the modified algorithm is faster than Holt's algorithm.
Edward Y. Y. Zhang, Patrick Shen-Pei Wang
ICPR2
1988 Knowledge Pattern Representation of Chinese Characters
abstract
This article discusses some intelligence aspects of Chinese characters. Some basic concepts of two-dimensional pattern representation and artificial intelligence such as semantic networks, forward chaining, deduction and the resolution principle are used to analyze and interpret the syntactic structure, representation, semantics and evolution of Chinese characters. The concept of degrees of ambiguity and the principle of new characters are investigated. It is found that Chinese characters are actually not only artistically elegant and culturally rich but also semantically meaningful and intelligently sound. Finally some topics for future research such as intelligent pattern recognition for Chinese characters, automatic learning and translation, and knowledge-based Chinese language understanding are discussed.
Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.1
1985 A new character recognition scheme with lower ambiguity and higher recognizability
Patrick Shen-Pei Wang
Pattern Recognit. Lett.1
1984 An application of array grammars to clustering analysis for syntactic patterns
Patrick Shen-Pei Wang
Pattern Recognit.1
1983 Hierarchical Structures and Complexities of Parallel Isometric Languages
abstract
The relationship between parallel isometric array languages and sequential isometric array languages is examined. Their hierarchical structures are investigated and a hierarchy is established by introducing parallel context-free array languages (PCFAL), derivation bounded array languages (DBAL), linear array languages (LAL), and extended regular afray languages (ERAL). It is interesting to find that some fundamental aspects that hold in one-dimensional string languages do not hold in their two-dimensional counterparts. Some parsing techniques are also explored. It is shown that while parallel parsing grammars may be simpler to write and parallel processing usually takes less time than sequential ones, the nature of parallel parsing is very complicated. Finally, several future research topics concerned with parallel isometric array languages including their complexities, hierarchical structures, and application to pattern recognition are discussed.
Patrick Shen-Pei Wang
IEEE Trans. Pattern Anal. Mach. Intell.1
1982 A New Hierarchy of Two-Dimensional Array Languages
Patrick Shen-Pei Wang
Inf. Process. Lett.1
1981 Finite-Turn Repetitive Checking Automata and Sequential/Parallel Matrix Languages
abstract
A class of machines capable of recognizing two-dimensional sequential/parallel matrix languages is introduced. It is called "finite-turn repetitive checking automata" (FTRCA). This checking automaton is provided with three READ-WRITE heads, one READ only head, and a control register in its memory to keep track of the internal transition states and the stack symbols. The value of the register ranges from 0 to a finite positive integer k and can repetitively appear as many times as necessary. A one-to-one correspondence is established between the class of FIRCA and (R:R)ML–the smallest class of the sequential/parallel matrix languages. Several closure properties are also investigated.
Patrick Shen-Pei Wang
IEEE Trans. Computers1
1980 Some New Results on Isotonic Array Grammars
Patrick Shen-Pei Wang
Inf. Process. Lett.1
1976 Sequential/parallel matrix picture languages
abstract
A new system, that of matrix grammars, for two-dimensional picture processing, is introduced. A hierarchy, induced on Chomsky's is found. Language operations such as union, catenation (row and column), Kleene's closure (row and column), and homomorphisms are investigated. It is found that the smallest class of these languages may serve as the class of arrays, which is defined as the smallest class of arrays closed under union, catenation (row and column) and Kleene's closure (row and column). Eight possible ways of defining a matrix language are discussed and it is suggested that one of them may lead to a normal form of matrix grammars. The method is advantageous over others on several points. Perhaps the most interesting of all is that it provides a compromise between purely sequential methods, which take too much time for large arrays and purely parallel methods, which usually take too much hardware for large arrays.
Patrick Shen-Pei Wang, William I. Grosky
SIGGRAPH1
1976 Recursiveness of Monotonic Array Grammars and a Hierarchy of Array Languages
Patrick Shen-Pei Wang, William I. Grosky
Inf. Process. Lett.1
1976 Erratum: Recursiveness of Monotonic Array Grammars and a Hierarchy of Array Languages
Patrick Shen-Pei Wang, William I. Grosky
Inf. Process. Lett.1