Xiaoqing Ding

dblp:53/82 · DBLP profile ↗
← Back
116ranked-venue papers
1as first author
0since 2021 · last 2018
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 80 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 49Databases, data management, data science and information retrieval · 31Human-computer interaction and ubiquitous computing · 4Applied, interdisciplinary, general and emerging computing · 3Security and privacy · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
13 papers
Representation and self-supervised learning · 51% Video understanding and tracking · 14% Face, body and person analysis · 13%
Computer graphics and multimedia
4 papers
Image and video processing · 100%
Databases, data mining, and information retrieval
2 papers
Data mining · 72% Information retrieval · 22% Recommender systems · 7%

Topics — the 28 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction
1.452018
Heteroscedastic Max-Min Distance Analysis for Dimensionality Reduction · IEEE Trans. Image Process. 2018
Discriminative Dimensionality Reduction for Multi-Dimensional Sequences · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Discriminative Transformation for Multi-Dimensional Temporal Sequences · IEEE Trans. Image Process. 2017
Machine learning › Representation and self-supervised learning › representation learning
metric learning
0.622018
Heteroscedastic Max-Min Distance Analysis for Dimensionality Reduction · IEEE Trans. Image Process. 2018
Discriminative Transformation for Multi-Dimensional Temporal Sequences · IEEE Trans. Image Process. 2017
Computer vision › Face, body and person analysis
face recognition
0.522018
Heteroscedastic Max-Min Distance Analysis for Dimensionality Reduction · IEEE Trans. Image Process. 2018
Advanced Joint Bayesian Method for Face Verification · IEEE Trans. Inf. Forensics Secur. 2015
Computer vision › Video understanding and tracking
action recognition
0.522017
Unsupervised Hierarchical Dynamic Parsing and Encoding for Action Recognition · IEEE Trans. Image Process. 2017
Hierarchical Dynamic Parsing and Encoding for Action Recognition · ECCV (4) 2016
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
supervised dimensionality reduction
0.312018
Discriminative Dimensionality Reduction for Multi-Dimensional Sequences · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Computer vision › Video understanding and tracking › temporal modeling
temporal dynamics modeling
0.312017
Unsupervised Hierarchical Dynamic Parsing and Encoding for Action Recognition · IEEE Trans. Image Process. 2017
Machine learning › Representation and self-supervised learning › multi-view learning
canonical correlation analysis
0.212016
Sufficient Canonical Correlation Analysis · IEEE Trans. Image Process. 2016
Computer vision › Segmentation and scene understanding › scene parsing
hierarchical parsing
0.212016
Hierarchical Dynamic Parsing and Encoding for Action Recognition · ECCV (4) 2016
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian classification
0.212015
Advanced Joint Bayesian Method for Face Verification · IEEE Trans. Inf. Forensics Secur. 2015
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
discriminant analysis
0.212015
Heteroscedastic max-min distance analysis · CVPR 2015
Computer vision › Face, body and person analysis › face recognition
face verification
0.212015
Advanced Joint Bayesian Method for Face Verification · IEEE Trans. Inf. Forensics Secur. 2015
Machine learning › Kernel, tree and ensemble methods › ensemble learning
boosting
0.212014
Learning Cascaded Shared-Boost Classifiers for Part-Based Object Detection · IEEE Trans. Image Process. 2014
Computer vision › Image recognition and object detection
object detection
0.212014
Learning Cascaded Shared-Boost Classifiers for Part-Based Object Detection · IEEE Trans. Image Process. 2014
Computer vision › Image recognition and object detection › object detection
part-based object detection
0.212014
Learning Cascaded Shared-Boost Classifiers for Part-Based Object Detection · IEEE Trans. Image Process. 2014
Image and video processing › image restoration
image deblurring
0.212014
Robust Image Restoration via Adaptive Low-Rank Approximation and Joint Kernel Regression · IEEE Trans. Image Process. 2014
Image and video processing
image restoration
0.212014
Robust Image Restoration via Adaptive Low-Rank Approximation and Joint Kernel Regression · IEEE Trans. Image Process. 2014
Image and video processing › super-resolution
image super-resolution
0.212014
Robust Image Restoration via Adaptive Low-Rank Approximation and Joint Kernel Regression · IEEE Trans. Image Process. 2014
Machine learning › Representation and self-supervised learning › representation learning
sequence representation
0.212013
Linear Sequence Discriminant Analysis: A Model-Based Dimensionality Reduction Method for Vector Sequences · ICCV 2013
Data mining
dimensionality reduction
0.212013
Linear Sequence Discriminant Analysis: A Model-Based Dimensionality Reduction Method for Vector Sequences · ICCV 2013
Data mining › clustering › image clustering
face clustering
0.112009
Face based image navigation and search · ACM Multimedia 2009
Information retrieval
multimedia analysis and retrieval
0.112009
Face based image navigation and search · ACM Multimedia 2009
Natural language and speech › Information extraction and text analysis
document analysis
0.112007
Character Independent Font Recognition on a Single Chinese Character · IEEE Trans. Pattern Anal. Mach. Intell. 2007
Natural language and speech › Information extraction and text analysis › document understanding › document image analysis
font recognition
0.112007
Character Independent Font Recognition on a Single Chinese Character · IEEE Trans. Pattern Anal. Mach. Intell. 2007
Data mining › predictive modeling › classification › structured classification
sequence classification
0.012013
Linear Sequence Discriminant Analysis: A Model-Based Dimensionality Reduction Method for Vector Sequences · ICCV 2013
Image and video processing › image restoration › document image restoration
document image rectification
0.012003
A Cylindrical Surface Model to Rectify the Bound Document Image · ICCV 2003
Image and video processing › document image analysis
document layout analysis
0.012003
Unified HMM-based layout analysis framework and algorithm · Sci. China Ser. F Inf. Sci. 2003
Recommender systems › content recommendation
image recommendation
0.012009
Face based image navigation and search · ACM Multimedia 2009
Natural language and speech › Information extraction and text analysis › document analysis
document structure extraction
0.012003
Unified HMM-based layout analysis framework and algorithm · Sci. China Ser. F Inf. Sci. 2003

Methods — techniques the papers use, named apart from their topics

bisection search · 0.5linear discriminant analysis · 0.4trace quotient · 0.3chernoff distance · 0.3temporal clustering · 0.3semidefinite programming · 0.3rank pooling · 0.3convex relaxation · 0.3sufficient dimension reduction · 0.2encoding · 0.2dynamic parsing · 0.2hidden markov model · 0.2nonlocal self-similarity · 0.2low-rank approximation · 0.2joint kernel regression · 0.2dictionary learning · 0.2fisher discriminant analysis · 0.2eigendecomposition · 0.2
YearPublicationVenuePosition
2018 Discriminative Dimensionality Reduction for Multi-Dimensional Sequences
abstract
Since the observables at particular time instants in a temporal sequence exhibit dependencies, they are not independent samples. Thus, it is not plausible to apply i.i.d. assumption-based dimensionality reduction methods to sequence data. This paper presents a novel supervised dimensionality reduction approach for sequence data, called Linear Sequence Discriminant Analysis (LSDA). It learns a linear discriminative projection of the feature vectors in sequences to a lower-dimensional subspace by maximizing the separability of the sequence classes such that the entire sequences are holistically discriminated. The sequence class separability is constructed based on the sequence statistics, and the use of different statistics produces different LSDA methods. This paper presents and compares two novel LSDA methods, namely M-LSDA and D-LSDA. M-LSDA extracts model-based statistics by exploiting the dynamical structure of the sequence classes, and D-LSDA extracts the distance-based statistics by computing the pairwise similarity of samples from the same sequence class. Extensive experiments on several different tasks have demonstrated the effectiveness and the general applicability of the proposed methods.
Bing Su 0001, Xiaoqing Ding, Hao Wang 0005, Ying Wu 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2018 Heteroscedastic Max-Min Distance Analysis for Dimensionality Reduction
abstract
Max-min distance analysis (MMDA) performs dimensionality reduction by maximizing the minimum pairwise distance between classes in the latent subspace under the homoscedastic assumption, which can address the class separation problem caused by the Fisher criterion, but is incapable of tackling heteroscedastic data properly. In this paper, we propose two heteroscedastic MMDA (HMMDA) methods to employ the differences of class covariances. Whitened HMMDA (WHMMDA) extends MMDA by utilizing the Chernoff distance as the separability measure between classes in the whitened space. Orthogonal HMMDA (OHMMDA) incorporates the maximization of the minimal pairwise Chernoff distance and the minimization of class compactness into a trace quotient formulation with an orthogonal constraint of the transformation, which can be solved by bisection search. Two variants of OHMMDA further encode the margin information by using only neighboring samples to construct the intra-class and inter-class scatters. Experiments on several UCI datasets and two face databases demonstrate the effectiveness of the HMMDA methods.
Bing Su 0001, Xiaoqing Ding, Changsong Liu, Ying Wu 0001
IEEE Trans. Image Process.2
2017 Attention Based RNN Model for Document Image Quality Assessment
abstract
Document Image Quality Assessment (DIQA) is an essential step preceding Optical Character Recognition (OCR). In this paper we propose an attention based Recurrent Neural Network (RNN) model for camera based DIQA. Convolutional Neural Network (CNN) and RNN are integrated into our model to capture spatial features for several glimpse regions step by step within an image patch. Reinforcement learning is adopted to train a locator to generate the optimal location of a glimpse region for the next time step so that attention can be payed to the salient part. Given an input document image, patches are generated with a sliding window, and the pure background ones are sifted out. Quality scores are obtained for all the sifted patches by applying the proposed attention based RNN method, and the patch scores are averaged over each input image as the result of DIQA. We conduct experiments on two public datasets and make comparisons with several other reported methods. Experimental results show that our model achieves the state of the art performance.
Pengchao Li, Liangrui Peng, Junyang Cai, Xiaoqing Ding, Shuangkui Ge
ICDAR4
2017 Local Discriminant Training and Global Optimization for Convolutional Neural Network Based Handwritten Chinese Character Recognition
abstract
This paper investigates local discriminant training and global optimization methods for Convolutional Neural Network (CNN) to improve its discriminant ability and recognition accuracy. For local discriminant training, we propose to combine triplet loss and softmax with cross-entropy loss as the loss function. The triplet loss is incorporated into an additional fully-connected layer before the final fully-connected layer of a CNN model. For global optimization, we use Conditional Random Field (CRF) to further utilize the pairwise distance of the CNN feature vectors trained with triplet loss. Experiments with different CNN models on handwritten Chinese character samples show that the combined local discriminant training and global optimization scheme achieves better character recognition accuracy and confidence analysis performance.
Xiangsheng Zeng, Donglai Xiang, Liangrui Peng, Changsong Liu, Xiaoqing Ding
ICDAR5
2017 Discriminative Transformation for Multi-Dimensional Temporal Sequences
abstract
Feature space transformation techniques have been widely studied for dimensionality reduction in vector-based feature space. However, these techniques are inapplicable to sequence data because the features in the same sequence are not independent. In this paper, we propose a method called max-min inter-sequence distance analysis (MMSDA) to transform features in sequences into a low-dimensional subspace such that different sequence classes are holistically separated. To utilize the temporal dependencies, MMSDA first aligns features in sequences from the same class to an adapted number of temporal states, and then, constructs the sequence class separability based on the statistics of these ordered states. To learn the transformation, MMSDA formulates the objective of maximizing the minimal pairwise separability in the latent subspace as a semi-definite programming problem and provides a new tractable and effective solution with theoretical proofs by constraints unfolding and pruning, convex relaxation, and within-class scatter compression. Extensive experiments on different tasks have demonstrated the effectiveness of MMSDA.
Bing Su 0001, Xiaoqing Ding, Changsong Liu, Hao Wang 0005, Ying Wu 0001
IEEE Trans. Image Process.2
2017 Unsupervised Hierarchical Dynamic Parsing and Encoding for Action Recognition
abstract
Generally, the evolution of an action is not uniform across the video, but exhibits quite complex rhythms and non-stationary dynamics. To model such non-uniform temporal dynamics, in this paper, we describe a novel hierarchical dynamic parsing and encoding method to capture both the locally smooth dynamics and globally drastic dynamic changes. It parses the dynamics of an action into different layers and encodes such multi-layer temporal information into a joint representation for action recognition. At the first layer, the action sequence is parsed in an unsupervised manner into several smooth-changing stages corresponding to different key poses or temporal structures by temporal clustering. The dynamics within each stage are encoded by mean-pooling or rank-pooling. At the second layer, the temporal information of the ordered dynamics extracted from the previous layer is encoded again by rank-pooling to form the overall representation. Extensive experiments on a gesture action data set (Chalearn Gesture) and three generic action data sets (Olympic Sports, Hollywood2, and UCF101) have demonstrated the effectiveness of the proposed method.
Bing Su 0001, Jiahuan Zhou, Xiaoqing Ding, Ying Wu 0001
IEEE Trans. Image Process.3
2016 Hierarchical Dynamic Parsing and Encoding for Action Recognition
Bing Su 0001, Jiahuan Zhou, Xiaoqing Ding, Hao Wang 0005, Ying Wu 0001
ECCV (4)3
2016 Combining multiple biometric traits with an order-preserving score fusion algorithm
Yicong Liang, Xiaoqing Ding, Changsong Liu, Jing-Hao Xue
Neurocomputing2
2016 Sufficient Canonical Correlation Analysis
abstract
Canonical correlation analysis (CCA) is an effective way to find two appropriate subspaces in which Pearson's correlation coefficients are maximized between projected random vectors. Due to its well-established theoretical support and relatively efficient computation, CCA is widely used as a joint dimension reduction tool and has been successfully applied to many image processing and computer vision tasks. However, as reported, the traditional CCA suffers from overfitting in many practical cases. In this paper, we propose sufficient CCA (S-CCA) to relieve CCA's overfitting problem, which is inspired by the theory of sufficient dimension reduction. The effectiveness of S-CCA is verified both theoretically and experimentally. Experimental results also demonstrate that our S-CCA outperforms some of CCA's popular extensions during the prediction phase, especially when severe overfitting occurs.
Yiwen Guo, Xiaoqing Ding, Changsong Liu, Jing-Hao Xue
IEEE Trans. Image Process.2
2016 A Boosting Approach to Exploit Instance Correlations for Multi-Instance Classification
abstract
We propose a Boosting approach for multi-instance (MI) classification. Lp-norm is integrated to localize the witness instances and formulate the bag scores from classifier outputs. The contributions are twofold. First, a flexible and concise model for Boosting is proposed by the Lp-norm localization and exponential loss optimization. The scores for bag-level classification are directly fused from the instance feature space without probabilistic assumptions. Second, gradient and Newton descent optimizations are applied to derive the weak learners for Boosting. In particular, the instance correlations are exploited by fitting the weights and Newton updates for the weak learner construction. The final Boosted classifiers are the sums of iteratively chosen weak learners. Experiments demonstrate that the proposed Lp-norm-localized Boosting approach significantly improves the MI classification performance. Compared with the state of the art, the approach achieves the highest MI classification accuracy on 7/10 benchmark data sets.
Yali Li 0001, Shengjin Wang, Qi Tian 0001, Xiaoqing Ding
IEEE Trans. Neural Networks Learn. Syst.4
2015 Heteroscedastic max-min distance analysis
abstract
Many discriminant analysis methods such as LDA and HLDA actually maximize the average pairwise distances between classes, which often causes the class separation problem. Max-min distance analysis (MMDA) addresses this problem by maximizing the minimum pairwise distance in the latent subspace, but it is developed under the homoscedastic assumption. This paper proposes Heteroscedastic MMDA (HMMDA) methods that explore the discriminative information in the difference of intra-class scatters for dimensionality reduction. WHMMDA maximizes the minimal pairwise Chenoff distance in the whitened space. OHMMDA incorporates this objective and the minimization of class compactness into a trace quotient formulation and imposes an orthogonal constraint to the final transformation, which can be solved by a bisection search algorithm. Two variants of OHMMDA are further proposed to encode the margin information. Experiments on several UCI Machine Learning datasets and the Yale Face database demonstrate the effectiveness of the proposed HMMDA methods.
Bing Su 0001, Xiaoqing Ding, Changsong Liu, Ying Wu 0001
CVPR2
2015 Application of Face Verification in Automated Passenger Clearance System
Wentao Shen, Yicong Liang, Xiaoqing Ding, Changsong Liu
ICPRAM (2)3
2015 Restoring camera-captured distorted document images
Changsong Liu, Baokang Wang, Xiaoqing Ding
Int. J. Document Anal. Recognit.4
2015 MiLDA: A graph embedding approach to multi-view face recognition
Yiwen Guo, Xiaoqing Ding, Jing-Hao Xue
Neurocomputing2
2015 A survey of recent advances in visual feature detection
Yali Li 0001, Shengjin Wang, Qi Tian 0001, Xiaoqing Ding
Neurocomputing4
2015 Exploring More Representative States of Hidden Markov Model in Optical Character Recognition: A Clustering-Based Model Pre-Training Approach
abstract
Hidden Markov Model (HMM) is an effective method to describe sequential signals in many applications. As to model estimation issue, common training algorithm only focuses on the optimization of model parameters. However, model structure influences system performance as well. Although some structure optimization methods are proposed, they are usually implemented as an independent module before parameter optimization. In this paper, the clustering feature of states in HMM is discussed through comparing the mechanism of Quadratic Discriminant Function (QDF) classifier and HMM. Then, through the clustering effect of Viterbi training and Baum–Welch training, a novel clustering-based model pre-training approach is proposed. It can optimize model parameters and model structure by turns, until the representative states of all models are explored. Finally, the proposed approach is evaluated on two typical OCR applications, printed and handwritten Arabic text line recognition. And it is compared with some other optimization methods. The improvement of character recognition performance proves the proposed approach can make more precise state allocation. And the representative states are benefit to HMM decoding.
Xiaoqing Ding, Liangrui Peng, Changsong Liu
Int. J. Pattern Recognit. Artif. Intell.2
2015 Feature representation for statistical-learning-based object detection: A review
Yali Li 0001, Shengjin Wang, Qi Tian 0001, Xiaoqing Ding
Pattern Recognit.4
2015 Importance sampling based discriminative learning for large scale offline handwritten Chinese character recognition
Xiaoqing Ding, Changsong Liu
Pattern Recognit.3
2015 Advanced Joint Bayesian Method for Face Verification
abstract
Generative Bayesian models have recently become the most promising framework in classifier design for face verification. However, we report in this paper that the joint Bayesian method, a successful classifier in this framework, suffers performance degradation due to its underuse of the expectation-maximization algorithm in its training phase. To rectify the underuse, we propose a new method termed advanced joint Bayesian (AJB). AJB has a good convergence property and achieves a higher verification rate than both the Joint Bayesian method and other state-of-the-art classifiers on the labeled faces in the wild face database.
Yicong Liang, Xiaoqing Ding, Jing-Hao Xue
IEEE Trans. Inf. Forensics Secur.2
2014 An MQDF-CNN Hybrid Model for Offline Handwritten Chinese Character Recognition
abstract
An MQDF-CNN hybrid model is presented for offline handwritten Chinese character recognition. The main idea behind MQDF-CNN hybrid model is that the significant difference on features and classification mechanisms between MQDF and CNN can complement each other. Linear confidence accumulation and multiplication confidence criteria are used for fusion outputs of MQDF and CNN. Experiments have been conducted on CASIA-HWDB1.1 and ICDAR2013 offline handwritten Chinese character recognition competition dataset. On both datasets, CNN beats MQDF by more than 1% of the accuracy, and the MQDF-CNN hybrid model has achieved the test accuracies of 92.03% and 94.44% respectively. The result on competition dataset is comparable to the state-of-the-art result though less training samples and only one CNN is used.
Xin Li 0144, Changsong Liu, Xiaoqing Ding, Youxin Chen
ICFHR4
2014 Fisher's linear discriminant embedded metric learning
Yiwen Guo, Xiaoqing Ding, Chi Fang, Jing-Hao Xue
Neurocomputing2
2014 Generalized joint kernel regression and adaptive dictionary learning for single-image super-resolution
Chen Huang 0001, Yicong Liang, Xiaoqing Ding, Chi Fang
Signal Process.3
2014 Topic Language Model Adaption for Recognition of Homologous Offline Handwritten Chinese Text Image
abstract
As the content of a full text page usually focuses on a specific topic, a topic language model adaption method is proposed to improve the recognition performance of homologous offline handwritten Chinese text image. Firstly, the text images are recognized with a character based bi-gram language model. Secondly, the topic of the text image is matched adaptively. Finally, the text image is recognized again with the best matched topic language model. To obtain a tradeoff between the recognition performance and computational complexity, a restricted topic language model adaption method is further presented. The methods have been evaluated on 100 offline Chinese text images. Compared to the general language model, the topic language model adaption has reduced the relative error rate by 11.94%. The restricted topic language model has lessened the running time by 19.22% at the cost of losing 0.35% of the accuracy.
Xiaoqing Ding, Changsong Liu
IEEE Signal Process. Lett.2
2014 Detecting Human Action as the Spatio-Temporal Tube of Maximum Mutual Information
abstract
Human action detection in complex scenes is a challenging problem due to its high-dimensional search space and dynamic backgrounds. To achieve efficient and accurate action detection, we represent a video sequence as a collection of feature trajectories and model human action as the spatio-temporal tube (ST-tube) of maximum mutual information. First, a random forest is built to evaluate the mutual information of feature trajectories toward the action class, and then a one-order Markov model is introduced to recursively infer the action regions at consecutive frames. By exploring the time-continuity property of feature trajectories, the action region is efficiently inferred at large temporal intervals. Finally, we obtain an ST-tube by concatenating the consecutive action regions bounding the human bodies. Compared with the popular spatio-temporal cuboid action model, the proposed ST-tube model is not only more efficient, but also more accurate in action localization. Experimental results on the KTH, CMU and UCF sports datasets validate the superiority of our approach over the state-of-the-art methods in both localization accuracy and time efficiency.
Taiqing Wang, Shengjin Wang, Xiaoqing Ding
IEEE Trans. Circuits Syst. Video Technol.3
2014 Robust Image Restoration via Adaptive Low-Rank Approximation and Joint Kernel Regression
abstract
In recent years, image priors based on nonlocal self-similarity and low-rank approximation have been proven as powerful tools for image restoration. Many restoration methods group similar patches as a matrix and recover the underlying low-rank structure from the corrupted matrix via rank minimization. However, both the nonlocally redundant and low-rank properties are highly content dependent, and whether they can faithfully characterize a wide range of natural images still remains unclear. In this paper, we analyze these two properties and provide quantifications of them in a data-driven and parametric way, respectively, obtaining the new measures of regional redundancy and nonlocal patch rank. Leveraging these prior leads to an adaptive image restoration method with content-awareness. In particular, our method iteratively removes outliers and recovers latent fine details. To handle outliers, we propose an adaptive low-rank and sparse matrix approximation algorithm to encourage the estimated nonlocal rank in the patch matrix. The guidance of regional redundancy further gives rise to the “denoise” quality. In the detail recovery step, we propose an adaptive joint kernel regression algorithm using the redundancy measure to determine the confidence of each regression group. It also bridges the gap between our online and offline dictionary learning schemes. Experiments on synthetic and real-world images show the efficacy of our method in image deblurring and super-resolution tasks, especially when subject to practical outliers such as rain drops.
Chen Huang 0001, Xiaoqing Ding, Chi Fang, Di Wen 0001
IEEE Trans. Image Process.2
2014 Learning Cascaded Shared-Boost Classifiers for Part-Based Object Detection
abstract
This paper focuses on the problem of detecting a number of different class objects in images. We present a novel part-based model for object detection with cascaded classifiers. The coarse root and fine part classifiers are combined into the model. Different from the existing methods which learn root and part classifiers independently, we propose a shared-Boost algorithm to jointly train multiple classifiers. This paper is distinguished by two key contributions. The first is to introduce a new definition of shared features for similar pattern representation among multiple classifiers. Based on this, a shared-Boost algorithm which jointly learns multiple classifiers by reusing the shared feature information is proposed. The second contribution is a method for constructing a discriminatively trained part-based model, which fuses the outputs of cascaded shared-Boost classifiers as high-level features. The proposed shared-Boost-based part model is applied for both rigid and deformable object detection experiments. Compared with the state-of-the-art method, the proposed model can achieve higher or comparable performance. In particular, it can lift up the detection rates in low-resolution images. Also the proposed procedure provides a systematic framework for information reusing among multiple classifiers for part-based object detection.
Yali Li 0001, Shengjin Wang, Qi Tian 0001, Xiaoqing Ding
IEEE Trans. Image Process.4
2014 3D face sparse reconstruction based on local linear fitting
Liu Ding, Xiaoqing Ding, Chi Fang
Vis. Comput.2
2013 Single-Image Super-Resolution via Adaptive Joint Kernel Regression
abstract
Single image super-resolution (SR) methods can be broadly categorized into three classes: interpolation-based methods, reconstruction-based methods [7], and example-based methods [2, 3, 6]. The reconstruction-based methods often incorporate prior knowledge to regularize the ill-posed problem. For example, Zhang et al. [7] assembled the Steering Kernel Regression [5] (SKR)-based local prior and Nonlocal Means [1] (NLM)based nonlocal prior. The example-based methods strongly rely on the chosen dictionary for satisfactory results. This paper focuses on learning good image priors and robust dictionaries for SR reconstruction. Among the extensively studied natural image priors, we choose to exploit the local structural regularity prior and nonlocal self-similarity prior in a coherent framework. We propose in this paper an Adaptive Joint Kernel Regression (AJKR)based prior to simultaneously exploit both image statistics. Our approach differs from others in several ways: 1) we combine a set of NLM-generalized local kernel regressors, which are more consistent with our nonlocal collaborative framework; 2) the proposed regional redundancy measure introduces higher-order statistics at the region level for each regression group, making the overall framework more adaptive (see Fig. 1(a)); and 3) an adaptive PCA-based dictionary learning scheme is adopted to bridge the gap of dictionaries learned online and offline by mixing them, and more importantly such a scheme together with its induced sparsity prior can adapt to the AJKR process in response to the regional redundancy measure (see the block diagram in Fig. 1(b)). The imaging model for SR (assume Y ∈ Rm is low resolution (LR) image, X ∈ Rn is high resolution (HR) image) is usually expressed as
Chen Huang 0001, Xiaoqing Ding, Chi Fang
BMVC2
2013 Linear Sequence Discriminant Analysis: A Model-Based Dimensionality Reduction Method for Vector Sequences
abstract
Dimensionality reduction for vectors in sequences is challenging since labels are attached to sequences as a whole. This paper presents a model-based dimensionality reduction method for vector sequences, namely linear sequence discriminant analysis (LSDA), which attempts to find a subspace in which sequences of the same class are projected together while those of different classes are projected as far as possible. For each sequence class, an HMM is built from states of which statistics are extracted. Means of these states are linked in order to form a mean sequence, and the variance of the sequence class is defined as the sum of all variances of component states. LSDA then learns a transformation by maximizing the separability between sequence classes and at the same time minimizing the within-sequence class scatter. DTW distance between mean sequences is used to measure the separability between sequence classes. We show that the optimization problem can be approximately transformed into an eigen decomposition problem. LDA can be seen as a special case of LSDA by considering non-sequential vectors as sequences of length one. The effectiveness of the proposed LSDA is demonstrated on two individual sequence datasets from UCI machine learning repository as well as two concatenate sequence datasets: APTI Arabic printed text database and IFN/ENIT Arabic handwriting database.
Bing Su 0001, Xiaoqing Ding
ICCV2
2013 Cross-Language Sensitive Words Distribution Map: A Novel Recognition-Based Document Understanding Method for Uighur and Tibetan
abstract
Cross-language document recognition and understanding have urgent realistic needs and extensive application prospects. In this paper, we propose a novel recognition-based Uighur and Tibetan document understanding method, termed "cross-language sensitive words distribution map" (CSWDM). In our unified recognition-understanding framework, digital Uighur/Tibetan document images are first recognized using OCR technology, and then CSWDM labels the Chinese information of sensitive words on the recognized transcriptions or directly on the original digital images, thus the space location and occurrence frequency of these sensitive words can be intuitively represented. With such information, readers can roughly understand the theme and meaning of the cross-language documents.
Bing Su 0001, Xiaoqing Ding, Liangrui Peng, Changsong Liu
ICDAR2
2013 A Novel Baseline-independent Feature Set for Arabic Handwriting Recognition
abstract
HMM-based analytical methods have been widely used for Arabic handwriting recognition. A key factor influencing the performance of HMM-based systems is the features extracted from a sliding window. In this paper, we propose a novel baseline-independent feature set extracted from a wider sliding window to directly capture the contextual information. This feature set is a combination of center of mass based log-space distribution features and inverse percentile features. Center of mass based log-space distribution features use a normalized histogram to describe the distribution of foreground pixels in different direction and distances with respect to the center of mass. Experiments on the IFN/ENIT database demonstrate the effectiveness of the proposed feature set. Further, this feature set can be combined with some popular baseline-independent features to form a large feature set, which achieves comparable results with several state-of-the-art systems using a simple HMM-based architecture.
Bing Su 0001, Xiaoqing Ding, Liangrui Peng, Changsong Liu
ICDAR2
2013 Similar Pattern Discriminant Analysis for Improving Chinese Character Recognition Accuracy
abstract
In this paper, a similar pattern discriminant analysis method is proposed. It optimizes the feature projection matrix based on similar pattern pairs and aims to extract targeted features for similar pattern discrimination. For improving Chinese character recognition accuracy, we introduce a cascade modified quadratic discriminant function (MQDF) model to combine linear discriminant analysis (LDA) and similar pattern discriminant analysis. The proposed method is investigated and compared with compound Mahalanobis function (CMF) on two data sets. The results indicate that the cascade MQDF achieves a better improvement and higher recognition accuracies than CMF. The relative recognition errors have been decreased up to 19.73% and 15.59% respectively on HCL2000 and THU-HCD datasets with respect to single MQDF.
Changsong Liu, Xiaoqing Ding
ICDAR3
2012 Analyzing the information entropy of states to optimize the number of states in an HMM-based off-line handwritten Arabic word recognizer
Xiaoqing Ding, Liangrui Peng, Changsong Liu
ICPR2
2012 A hierarchical algorithm with multi-feature fusion for facial expression recognition
Chi Fang, Xiaoqing Ding
ICPR3
2012 Pose robust face tracking by combining view-based AAMs and temporal filters
Chen Huang 0001, Xiaoqing Ding, Chi Fang
Comput. Vis. Image Underst.2
2012 Continuous Pose Normalization for Pose-Robust Face Recognition
abstract
Pose variation is a great challenge for robust face recognition. In this paper, we present a fully automatic pose normalization algorithm that can handle continuous pose variations and achieve high face recognition accuracy. First, an automatic method is proposed to find pose-dependent correspondences between 2-D facial feature points and 3-D face model. This method is based on a multi-view random forest embedded active shape model. Then we densely map each pixel in the face image onto the 3-D face model and rotate it to the frontal view. The filling of occluded face regions is guided by facial symmetry. Recognition experiments were conducted on the two western databases CMU-PIE, FERET and one eastern database CAS-PEAL. Currently the algorithm has been trained with pose variation up to ±50° in yaw. Our algorithm not only achieves high recognition accuracy for learnt poses but also shows good generalizability for extreme poses. Furthermore, it suggests the promising application to people of different races.
Liu Ding, Xiaoqing Ding, Chi Fang
IEEE Signal Process. Lett.2
2011 A Multi-scale Text Line Segmentation Method in Freestyle Handwritten Documents
abstract
Text lines in free-style handwritten documents are often curved, touch or overlap with each other, which presents a challenge for text line segmentation. In this paper, we proposed a novel text line segmentation method that utilizes the advantages of algorithms in both the small scale and large scale. A path is dynamically detected between each pair of neighboring text lines to separate them. During the process, the line-separating path's coordinate in each step is determined by a three-stage multi-scale method that combines (1) a simple local minima search algorithm, (2) the technique based on following the contour of the foreground component and (3) the piecewise projection profile. Without training, our method has achieved a high segmentation accuracy on plenty of samples, which proves its strong adaptability to various line conditions. Experimental results show that the proposed method outperforms traditional methods.
Yangdong Gao, Xiaoqing Ding, Changsong Liu
ICDAR2
2011 A Novel Short Merged Off-line Handwritten Chinese Character String Segmentation Algorithm Using Hidden Markov Model
abstract
Hidden Markov model (called "HMM" for short) has been a widespread method to segment sequential data in speech recognition and DNA sequence analysis. According to the same principle, it can be also used in segmenting short merged off-line handwritten Chinese character strings, which is a tough issue but often met in practice. Because HMM is still not a common method in this field nowadays, in this paper, we will introduce a novel algorithm using HMM for the segmentation issue above. Eventually, this segmentation algorithm can achieve an applicable performance even when 3755 character classes are compressed into similar characters classes with only 1% amount of original ones, and it also shows an enormous potential of segmenting long text lines.
Xiaoqing Ding, Changsong Liu
ICDAR2
2011 MQDF Discriminative Learning Based Offline Handwritten Chinese Character Recognition
abstract
This paper has proposed a discriminative learning method of modified quadratic discriminant function (MQDF) based on sample importance weights. Firstly, sample importance function is derived from distance based recognition results under bayes decision rule. It weights samples according to extended recognition confidence. On these weighted samples, parameters of MQDF are modulated indirectly by re-estimating the mean vector and covariance matrix. The proposed method is investigated and compared with other discriminative learning methods about MQDF on THU-HCD offline Chinese handwriting sets. The results show that the proposed method has improved the basic MQDF drastically and outperforms other methods compared.
Xiaoqing Ding, Changsong Liu
ICDAR2
2011 An Improved Method Based on Weighted Grid Micro-structure Feature for Text-Independent Writer Recognition
abstract
Writer recognition is a very important branch of biometrics. In our previous research, a Grid Micro-structure Feature (GMSF) based text-independent and script-independent method was adopted and high performance was obtained. However, this method is sensitive to pen-width variation in practical situation. To solve this problem, an inner and inter class variances weighted high-dimensional feature matching method is proposed. The inner and inter class variances are estimated on handwriting samples with different pen-width written by different writers. Experimental results show that our method is effective.
Xiaoqing Ding, Liangrui Peng, Xin Li 0144
ICDAR2
2011 An Improved Scene Text Extraction Method Using Conditional Random Field and Optical Character Recognition
abstract
Over the past few years, research on scene text extraction has developed rapidly. Recently, condition random field (CRF) has been used to give connected components (CCs) 'text' or 'non-text' labels. However, a burning issue in CRF model comes from multiple text lines extraction. In this paper, we propose a two-step iterative CRF algorithm with a Belief Propagation inference and an OCR filtering stage. Two kinds of neighborhood relationship graph are used in the respective iterations for extracting multiple text lines. Furthermore, OCR confidence is used as an indicator for identifying the text regions, while a traditional OCR filter module only considered the recognition results. The first CRF iteration aims at finding certain text CCs, especially in multiple text lines, and sending uncertain CCs to the second iteration. The second iteration gives second chance for the uncertain CCs and filter false alarm CCs with the help of OCR. Experiments based on the public dataset of ICDAR 2005 prove that the proposed method is comparative with the existing algorithms.
Changsong Liu, Xiaoqing Ding, Kongqiao Wang
ICDAR4
2011 2D face fitting-assisted 3D face reconstruction for pose-robust face recognition
Liu Ding, Xiaoqing Ding, Chi Fang
Soft Comput.3
2010 Person-independent head pose estimation based on random forest regression
abstract
In this paper, a novel approach for person-independent head pose estimation in gray-level images is presented. There are two steps of the proposed method. In order to preserve similar patterns of faces under various poses, a novel multi-view face detector using tree-structured cascaded-Adaboost classifiers is applied. Furthermore, based on the cropped face images, randomized regression trees are learned and applied to estimate head pose precisely. Experiments show that our method achieves better pose estimation results in both horizontal and vertical orientations in comparison with the reported result with skin color information.
Yali Li 0001, Shengjin Wang, Xiaoqing Ding
ICIP3
2010 Discriminative Prototype Learning in Open Set Face Recognition
abstract
We address the problem of prototype design for open set face recognition (OSFR) using single sample image. Normalized Correlation (NC), also known as Cosine Distance, offers many benefits in accuracy and robustness compared to other distance measurement in OSFR problem. Inspired by classical Learning Vector Quantization (LVQ), a novel discriminative learning method is proposed to design a discriminative prototype used by NC classifier. Specifically, we develop an objective function that fixes the NC score between the prototype and within-class sample at a high level and minimizes the similarity between the prototype and between-class samples. Several experiments conducted on benchmark databases demonstrate the superior performance of the prototype designed compared to the original one.
Zhongkai Han, Chi Fang, Xiaoqing Ding
ICPR3
2010 Head Pose Estimation Based on Random Forests for Multiclass Classification
abstract
Head pose estimation remains a unique challenge for computer vision system due to identity variation, illumination changes, noise, etc. Previous statistical approaches like PCA, linear discriminative analysis (LDA) and machine learning methods, including SVM and Adaboost, cannot achieve both accuracy and robustness that well. In this paper, we propose to use Gabor feature based random forests as the classification technique since they naturally handle such multi-class classification problem and are accurate and fast. The two sources of randomness, random inputs and random features, make random forests robust and able to deal with large feature spaces. Besides, we implement LDA as the node test to improve the discriminative power of individual trees in the forest, with each node generating both constant and variant number of children nodes. Experiments are carried out on two public databases to show the proposed algorithm outperforms other approaches in both accuracy and computational efficiency.
Chen Huang 0001, Xiaoqing Ding, Chi Fang
ICPR2
2010 Part Detection, Description and Selection Based on Hidden Conditional Random Fields
abstract
In this paper, the problem of part detection, description and selection is discussed. This problem is crucial in the learning algorithms of part-based models, but can't be solved well when some candidate parts are extracted from background. This paper studies this problem and introduces a new algorithm, HCRF-PS (Hidden Conditional Random Fields for Part Selection), for part detection, description, especially selection. Our algorithm is distinguished for its power to optimize multiple kinds of information at the same time, including texture, color, location and part label. Finally, we did some experiments with HCRF-PS algorithm which give good results on both virtual and real data.
Wenhao Lu, Shengjin Wang, Xiaoqing Ding
ICPR3
2010 Development of Recognition Engine for Baby Faces
abstract
Existing face recognition approaches are mostly developed based on adult faces which may not work well in distinguishing faces of kids. Especially, baby faces tend to have common features such as round cheeks and chins, so that current face recognition engines often fail to differentiate them. In this paper, we present methods for discriminating baby faces from adult faces, and for training a special engine to recognize faces of different babies. To achieve these, we collected a huge number of baby face images and developed a software system to annotate the image database. Experimental results prove that the trained baby face recognizer achieves dramatic improvement on differentiating baby faces and the fusion of it with the conventional adult face recognition engine also works well on the overall data set containing both baby and adult faces.
Di Wen 0001, Chi Fang, Xiaoqing Ding, Tong Zhang 0007
ICPR3
2010 Multi-font printed Mongolian document recognition system
Liangrui Peng, Changsong Liu, Xiaoqing Ding, Jianming Jin, Youshou Wu, Yanhua Bao
Int. J. Document Anal. Recognit.3
2010 Eye/eyes tracking based on a unified deformable template and particle filtering
Yali Li 0001, Shengjin Wang, Xiaoqing Ding
Pattern Recognit. Lett.3
2010 Action and Gait Recognition From Recovered 3-D Human Joints
abstract
A common viewpoint-free framework that fuses pose recovery and classification for action and gait recognition is presented in this paper. First, a markerless pose recovery method is adopted to automatically capture the 3-D human joint and pose parameter sequences from volume data. Second, multiple configuration features (combination of joints) and movement features (position, orientation, and height of the body) are extracted from the recovered 3-D human joint and pose parameter sequences. A hidden Markov model (HMM) and an exemplar-based HMM are then used to model the movement features and configuration features, respectively. Finally, actions are classified by a hierarchical classifier that fuses the movement features and the configuration features, and persons are recognized from their gait sequences with the configuration features. The effectiveness of the proposed approach is demonstrated with experiments on the Institut National de Recherche en Informatique et Automatique Xmas Motion Acquisition Sequences data set.
Junxia Gu, Xiaoqing Ding, Shengjin Wang, Youshou Wu
IEEE Trans. Syst. Man Cybern. Part B2
2009 Improved 3D assisted pose-invariant face recognition
abstract
Recent face recognition algorithm can achieve high accuracy when testing face samples are frontal. However, when face pose changes largely, the performance of existing methods drop drastically. In this paper, we propose an improved algorithm aiming at recognizing faces of different poses when each face class has only one frontal training sample. For each sample, a 3D face is constructed by using 3D morphable model (3DMM). The shape and texture parameters of 3DMM are recovering by fitting the model to the 2D face sample which is a non-linear optimization problem. The virtual faces of different views are generated from the 3DMM to assist face recognition. Different from the conventional optimization energy function, proposed energy function takes not only image intensity but also shape constraint into account. In this paper, we locate 88 sparse points from the 2D face sample by automatic face fitting and use their correspondence in the 3D face as shape constraint. We experiment proposed method on the publicly available CMUPIE database which includes faces viewed from 11 different poses and the results show that proposed method is effective and the face recognition results towards pose-variant are promising.
Liu Ding, Xiaoqing Ding, Chi Fang
ICASSP3
2009 Face based image navigation and search
abstract
People are often the most important subjects in photos, and the ability of finding photos of a particular person easily and quickly in an image collection is highly desired. In this paper, we present a face clustering system which automatically groups photos into clusters, with each cluster containing photos of the same person. This is done based on an advanced face recognition engine and a semi-supervised clustering approach. The system achieved good clustering accuracy when tested on different image sets and by different users. Moreover, features such as adding new images, face cluster navigation and face based image retrieval are added that greatly improve the usability of the system. It also facilitates efficient manual manipulations of clustering results. On top of this technology, image navigation systems have been built, including the "face bubble" visualization which provides one-glance view of a photo collection, and shows the relations among people.
Tong Zhang 0007, Jun Xiao 0003, Di Wen 0001, Xiaoqing Ding
ACM Multimedia4
2009 Vehicle Detection and Tracking in Relatively Crowded Conditions
abstract
Aiming at vehicle detection and tracking problems in video monitoring and controlling system, this paper mainly studies vehicle detection and tracking problems in conditions of high traffic density in daytime. This paper is distinguished by two key contributions. First, we develop an improvement — SEAP(Simple but Efficient After Process) which checks the detection results in an accurate way and is an after process of Adaboost [1] detector which used to detect car in every frame. Second, we propose a tracking algorithm named 4-states tracking algorithm based on Kalman[5] linear filter. Tracking results turn unsteady as traffic density grows higher because of much more false positives and false negatives appear. However, 4-states tracking algorithm can solve this problem in an easy way by introducing FSM (Finite State Machine) into tracking algorithm. Finally, we implement a real-time vehicle detection and tracking system with the upper methods. Experiments give good results in relative crowded Conditions.
Wenhao Lu, Shengjin Wang, Xiaoqing Ding
SMC3
2008 Adaptive particle filter with body part segmentation for full body tracking
abstract
This paper presents a novel approach for marker-less 3D full body pose tracking using adaptive particle filter. Firstly, the search space decomposition strategy and body part segmentation method are used to reduce the calculation complexity due to the large degrees of freedom. Then an adaptive particle filter is adopted to track each body part. This new technique is a significant improvement over the standard particle filter with the advantage of adaptive particle number for each body part. Experimental results on tracking several challenging action sequences have shown that the proposed 3D full body tracker is able to effectively handle rapid no-linear movements, large changes of viewpoint, and different actors. The average errors of joint position are from 0.56 to 1.13 voxel in these action sequences.
Junxia Gu, Xiaoqing Ding, Shengjin Wang, Youshou Wu
FG2
2008 Full body tracking-based human action recognition
abstract
In this paper, we present a novel method for human action recognition with the combined global movement feature and local configuration feature. The human action is represented as a sequence of joints in the 4D spatio-temporal space, and modeled by two HMMs, a conventional HMM for global movement feature and an exemplar-based HMM for configuration feature. Firstly, an adaptive particle filter is adopted to track the marker-less actor's 3D joints. Then, the combined features are extracted from the full body tracking results. Finally, the actions are classified by fusing two HMMs. The effectiveness of the proposed algorithm is demonstrated with experiments on 7 actions by 12 actors. The results prove robustness of the proposed method with respect to viewpoints and actors.
Junxia Gu, Xiaoqing Ding, Shengjin Wang, Youshou Wu
ICPR2
2008 Bhattacharyya boosting
abstract
In this paper, we discuss a new feature selection criterion in boosting. Our method directly optimizes Bhattacharyya distance between the weighted positive and negative samples to find the feature vector instead of brute-force selection from a feature pool. Unlike some similar work including FisherBoost and MRCBoost, the new criterion connects the feature selection process with cost function minimization. Thus the coefficients of the member classifiers and sample distribution update are both theoretically accordant with feature selection. Our criterion can be regarded as a generalized version of FisherBoost and MRCBoost. The experiments on several data sets validate our algorithm’s effectiveness on both training and testing sets. 1.
Xiaoqing Ding
ICPR2
2008 Pose adaptive LDA based face recognition
abstract
In this paper, a novel method based on pose adaptive linear discriminant analysis (PALDA) is proposed to deal with pose variation problems in face recognition when each person has only one frontal training sample. The basic idea of the PALDA method is described as follows: first, the pose style of the test sample is estimated by one pose classifier; then, the corresponding LDA feature, which is robust to the variation between the estimated pose style and the frontal pose style, is extracted. Since human faces are very similar objects with similar geometrical shape and configuration, the facial images from the same pose style are similar to each other. So an effective pose classifier is designed. The variations of each specific person’s facial images, due to changes of pose, are also rather similar to each other, so varieties of LDA projection matrixes, each of which is robust to one variation between one specific pose style and the frontal pose style, are trained with offline face samples containing various poses. The comparison with other pose robust face recognition methods on CMUPie database has confirmed the effectiveness and the robustness of the proposed method.
Zhenger Wang, Xiaoqing Ding, Chi Fang
ICPR2
2008 Asymmetric Real Adaboost
abstract
A cost-sensitive extension of Real Adaboost denoted as asymmetric Real Adaboost(RAB) is proposed. The two main differences between Asymmetric RAB and the naïve RAB are (1) a Chernoff measurement is used to evaluate the best weak classifier during training, rather than a Bhattacharyya measurement used in naïve RAB, and (2) the weights are updated separately for positives and negatives at each boosting step. The upper bound on training error is also provided. Experiment results are shown to demonstrate its cost-sensitivity when selecting weak classifiers, and also show that it outperforms previously proposed cost-sensitive extensions of Discrete Adaboost(DAB) and several extensions of Real Adaboost. Besides, it also consumes much less time than previously proposed DAB extensions.
Zhanjun Wang, Chi Fang, Xiaoqing Ding
ICPR3
2008 Arbitrary warped document image restoration based on segmentation and Thin-Plate Splines
abstract
Warping is a common appearance in camera captured document images. It is the primary factor that makes such kind of document images hard to be recognized. Therefore it is necessary to restore warped document image before recognition. In this paper, a novel restore method is presented. The method takes a rough line segmentation and character segmentation firstly in order to estimate the warping direction. Then several pairs of key points mapping between the original image and the restored image are determined and thin-plate splines (TPS), which is an interpolation algorithm, is introduced to restore the image. Such process can effectively describe the warping direction of the document and successfully restore the image. Some experimental results show the effect of the image restoration and compare the recognition rate before and after the restoration based on a same OCR application.
Changsong Liu, Xiaoqing Ding, Yanming Zou
ICPR3
2008 Classifier Combination and its Application in Iris Recognition
abstract
Classifier combination is an effective method to improve the recognition accuracy of a biometric system. It has been applied to many practical biometric systems and achieved excellent performance. However, there is little literature involving theoretical analysis on the effectiveness of classifier combination. In this paper, we investigate classifiers combined with the max and min rules. In particular, we compute the recognition performance of each combined classifier, and illustrate the condition in which the combined classifier outperforms the original unimodal classifier. We focus our study on personal verification, where the input pattern is classified into one of two categories, the genuine or the impostor. For simplicity, we further assume that the matching score produced by the original classifier follows a normal distribution and the outputs of different classifiers are independent and identically distributed. Randomly-generated data are employed to test our conclusion. The influence of finite samples is explored at the same time. Moreover, an iris recognition system, which adopts multiple snapshots to identify a subject, is introduced as a practical application of the above discussions.
Xinhua Feng, Xiaoqing Ding, Youshou Wu, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.2
2008 Visual Tracker Using Sequential Bayesian Learning: Discriminative, Generative, and Hybrid
abstract
This paper presents a novel solution to track a visual object under changes in illumination, viewpoint, pose, scale, and occlusion. Under the framework of sequential Bayesian learning, we first develop a discriminative model-based tracker with a fast relevance vector machine algorithm, and then, a generative model-based tracker with a novel sequential Gaussian mixture model algorithm. Finally, we present a three-level hierarchy to investigate different schemes to combine the discriminative and generative models for tracking. The presented hierarchical model combination contains the learner combination (at level one), classifier combination (at level two), and decision combination (at level three). The experimental results with quantitative comparisons performed on many realistic video sequences show that the proposed adaptive combination of discriminative and generative models achieves the best overall performance. Qualitative comparison with some state-of-the-art methods demonstrates the effectiveness and efficiency of our method in handling various challenges during tracking.
Yun Lei, Xiaoqing Ding, Shengjin Wang
IEEE Trans. Syst. Man Cybern. Part B2
2007 Improving Naive Bayes Text Classifier Using Smoothing Methods
Xiaoqing Ding
ECIR2
2007 Clustering Consumer Photos Based on Face Recognition
abstract
The ability of finding photos of a particular person through face recognition is a highly desired feature in indexing, searching and browsing consumer photo collections. In this research, based on an advanced face recognition engine we developed in prior work, one two-pass clustering approach is proposed which groups photos of the same person in a fully automatic way. Firstly, a similarity matrix for all detected faces is computed, with which a semi-supervised clustering is done. Next, larger clusters are selected and modeled as people frequently appearing in the image collection. Then, smaller clusters are recognized against these dominant clusters. Contextual information is used to obtain better results. The approach achieved promising accuracy when tested on an image dataset containing 2316 photos.
Liexian Gu, Tong Zhang 0007, Xiaoqing Ding
ICME3
2007 Multi-view Moving Human Detection and Correspondence Based on Object Occupancy Random Field
Xiaoqing Ding, Shengjin Wang, Youshou Wu
ISNN (3)2
2007 Discriminative Training Based Quadratic Classifier for Handwritten Character Recognition
abstract
In offline handwritten character recognition, the classifier with modified quadratic discriminant function (MQDF) has achieved good performance. The parameters of MQDF classifier are commonly estimated by the maximum likelihood (ML) estimator, which maximizes the within-class likelihood instead of directly minimizing the classification errors. To improve the performance of MQDF classifier, in this paper, the MQDF parameters are revised by discriminative training using a minimum classification error (MCE) criterion. The proposed algorithm is applied to recognizing handwritten numerals and handwritten Chinese characters, the recognition rates obtained are among the highest that have ever been reported.
Xiaoqing Ding, Hailong Liu 0006
Int. J. Pattern Recognit. Artif. Intell.2
2007 Character Independent Font Recognition on a Single Chinese Character
abstract
A novel algorithm for font recognition on a single unknown Chinese character, independent of the identity of the character, is proposed in this paper. We employ a wavelet transform on the character image and extract wavelet features from the transformed image. After a Box-Cox transformation and LDA (Linear Discriminant Analysis) process, the discriminating features for font recognition are extracted and classified through a MQDF (Modified Quadric Distance Function) classifier with only one prototype for each font class. Our experiments show that our algorithm can achieve a recognition rate of 90.28 percent on a single unknown character and 99.01 percent if five characters are used for font recognition. Compared with existing methods, all of which are based on a text block, our method can provide a higher recognition rate and is more flexible and robust, since it is based on a single unknown character. Additionally, our method demonstrates that it is possible to extract subtle yet discriminative signals embedded in a much larger noisy background.
Xiaoqing Ding
IEEE Trans. Pattern Anal. Mach. Intell.1
2006 Application of Bi-gram Driven Chinese Handwritten Character Segmentation for an Address Reading System
Xiaoqing Ding
Document Analysis Systems2
2006 Offline Handwritten Arabic Character Segmentation with Probabilistic Model
Pingping Xiu, Liangrui Peng, Xiaoqing Ding
Document Analysis Systems3
2006 Object Detection Via Fusion of Global Classifier and Part-Based Classifier
Shengjin Wang, Xiaoqing Ding
ISNN (2)3
2006 Visual Similarity Based Document Layout Analysis
Di Wen 0001, Xiaoqing Ding
J. Comput. Sci. Technol.2
2005 Handwritten Character Recognition Using Gradient Feature and Quadratic Classifier with Multiple Discrimination Schemes
abstract
In the research of statistical approach for handwritten character recognition, directional element feature (DEF) and modified quadratic discriminant function (MQDF) have been extremely successful and widely used in practical applications. In this paper, we apply several state-of-the-art techniques of handwritten character recognition on this baseline system to improve the recognition accuracy. In feature extraction stage, gradient feature is extracted to replace DEF, which provides higher resolution on both magnitude and angle of the directional strokes in character image. In classification stage, the performance of MQDF classifier is enhanced by multiple discrimination schemes, including minimum classification error (MCE) training on the classifier parameters and modified distance representation for similar characters discrimination. All these techniques we use lead to improvement on the character recognition rate. The performance of the improved recognition system has been evaluated by both handwritten digit recognition and handwritten Chinese character recognition experiments, in which very promising results are achieved.
Hailong Liu 0006, Xiaoqing Ding
ICDAR2
2005 A hybrid post-processing system for offline handwritten Chinese script recognition
Chew Lim Tan, Xiaoqing Ding
Pattern Anal. Appl.3
2005 Gabor filters-based feature extraction for character recognition
Xiaoqing Ding, Changsong Liu
Pattern Recognit.2
2004 Recognition of 3-D objects in multiple statuses based on Markov random field models
abstract
A general framework is presented to realize 3D object recognition, invariant to object scaling, deformation, rotation, occlusion, and viewpoint change. This framework utilizes densely sampled grids, with different resolutions, to represent the local information of the input image. A Markov random field (MRF) model is then created to model the geometric distribution of the object key nodes. Flexible matching, which is aimed at finding the accurate correspondence map between the key points of two images, is performed by combining the local similarities and the geometric relations together using the highest confidence first (HCF) method. Afterwards, a global similarity is calculated for object recognition. Experimental results on the Coil-100 object database are presented. The excellent recognition rates achieved in all the experiments indicate that our approach is well-suited for appearance-based recognition.
Xiaoqing Ding, Shengjin Wang
ICASSP (5)2
2004 Contextual post-processing based on the confusion matrix in offline handwritten Chinese script recognition
Chew Lim Tan, Xiaoqing Ding, Changsong Liu
Pattern Recognit.3
2003 A Cylindrical Surface Model to Rectify the Bound Document Image
abstract
We propose a novel approach on how to rectify the photo image of the bound document. The surface of the document is modeled by a cylindrical surface. By the geometry of camera image formation, the equations using the cue of directrixes to map the points on the surface in the 3D scene to the points on the image plane are achieved. Baselines of the horizontal text line are extracted as projections of directrixes to estimate the bending extent of the surface, and then the images are rectified. The proposed method needs no auxiliary device. Experimental results are presented to demonstrate the feasibility and the application of the method.
Huaigu Cao, Xiaoqing Ding, Changsong Liu
ICCV2
2003 Rectifying the Bound Document Image Captured by the Camera: A Model Based Approach
abstract
A model based approach for rectifying the camera image of the bound document has been developed, i.e., the surface of the document is represented by a general cylindrical surface. The principle of using the model to unwrap the image is discussed. Practically, the skeleton of each horizontal text line is extracted to help estimate the parameter of the model, and rectify the images. To use the model, only a few priori is required, and no more auxiliary device is necessary. Experiment results are given to demonstrate the feasibility and the stability of the method.
Huaigu Cao, Xiaoqing Ding, Changsong Liu
ICDAR2
2003 Rejection Algorithm for Mis-segmented Characters In Multilingual Document Recognition
abstract
In OCR systems the character segmentation algorithm may generate mis-segmented blocks. Feedback information from character classifier is indispensable to achieve higher character segmentation accuracy. In this paper a novel rejection algorithm is proposed to identify these mis-segmented characters more accurately. First, based on confidence evaluation of distance-based classifiers, the usual generalized confidence mapping function is modified to fit this specific purpose. Second, a novel adaptive thresholding rejection rule is proposed, which is more accurate and flexible. Experiments on Chinese, Japanese and Korean document recognition showed that new rejection algorithm evidently improved the system performance, especially for low-quality printed document recognition. 1.
Zhengang Chen, Xiaoqing Ding
ICDAR2
2003 Writer Identification Using Directional Element Features and Linear Transform
abstract
In this paper, we present an effective method to writer identification that is carried out using single Chinese character as script, which is very flexible and very easy to be used in practice. The directional element features are first extracted from the handwriting character scripts, then the dimensions of the features is reduced using PCA in order to cope with the small sample size problem. The most discriminative features are extracted from the reduced feature space using Fisher’s Linear Discriminant Analysis. The Euclidian distance is proposed for classification. Experimental result verified the effectiveness of the proposed method.
Xianliang Wang, Xiaoqing Ding, Hailong Liu 0006
ICDAR2
2003 Real-time rotation invariant face detection based on cost-sensitive AdaBoost
abstract
This paper presents a novel method of detecting faces at any degree of rotation in the image plane based on cost-sensitive AdaBoost (CS-AdaBoost) algorithm. The method first employs a cascade of very simple classifiers trained by CS-AdaBoost to determine the possible orientation of each input window and then uses an upright face detector also trained by CS-AdaBoost to verify the derotated face candidate. The training procedures for both classifiers are also given in the paper. Experimental results show that the proposed method gives higher detection ratio and real-time detection speed compared to the conventional ones.
Xiaoqing Ding
ICIP (3)2
2003 Discriminant local feature analysis of facial images
abstract
In the traditional DKL algorithm, PCA is performed before LDA to achieve stable numerical computation and good generalization in small-sample-size problems. However, PCA is based on the global information, ignoring the significant local characteristics. This paper will propose a novel algorithm called discriminant local feature analysis based on a broader understanding of LFA features. In the algorithm, LFA instead of PCA is applied before LDA. On the one hand, LFA captures local characteristics with little loss of global information. On the other hand, it presents an effective low-dimensional representation of signals, and thus reduces the dimensionality for LDA. By combining LFA and LDA, the DLFA algorithm outperforms DKL, which is showed by experiments on open-set face verification.
Qiong Yang, Xiaoqing Ding, Zhengang Chen
ICIP (2)2
2003 Robust real-time face detection based on cost-sensitive AdaBoost method
abstract
This paper presents a method of detecting faces based on cost-sensitive AdaBoost (CS-AdaBoost) algorithm. The two main differences between CS-AdaBoost algorithm and the naive AdaBoost are that (1) unequal initial weights are given to each training sample according to its misclassification cost, and (2) the weights are updated separately for positives and negatives at each boosting step. Due to these two variations, every stage of the face detector trained by CS-AdaBoost algorithm can more effectively focus on face samples than by the naive AdaBoost to achieve robust and high detection rate with modest false alarm rate, so that the final face detector can yield high detection rates, very low false positive rates, and robust performance. Experiments also demonstrate the effectiveness of our method.
Xiaoqing Ding
ICME2
2003 Unified HMM-based layout analysis framework and algorithm
Xiaoqing Ding, Youshou Wu
Sci. China Ser. F Inf. Sci.2
2002 Symmetrical PCA in face recognition
abstract
Facial symmetry is a useful natural characteristic of facial images, which can help in the development of face-oriented recognition technology and algorithms. The paper applies it to face recognition after introducing mirror images. By combining PCA with the even-odd decomposition principle, a new algorithm called symmetrical principal component analysis is proposed, in which different energy ratios of even/odd symmetrical principal components and their different sensitivities to pattern variations are employed for feature selection. This algorithm has two outstanding advantages. Firstly, it effectively improves the stability of features and remarkably raises the recognition rate. Secondly, it greatly saves computational cost as well as storage space.
Qiong Yang, Xiaoqing Ding
ICIP (2)2
2002 Automatic performance evaluation of printed Chinese character recognition systems
Chi Fang, Changsong Liu, Liangrui Peng, Xiaoqing Ding
Int. J. Document Anal. Recognit.4
2002 Combining character-based bigrams with word-based bigrams in contextual postprocessing for Chinese script recognition
abstract
It is crucial to use contextual information to improve the recognition accuracy of Chinese script in an offline, handwritten Chinese character-recognition system. However, with the increase in the number of candidates given by a character recognizer, contextual postprocessing using a word-based bigram is time-consuming. This article presents a novel contextual postprocessing method that integrates character-based bigram postprocessing with word-based bigram postprocessing in light of the complementary action between Chinese characters and Chinese words. On the basis of isolated character recognition, character-based bigram postprocessing using a forward-backward search is first executed on a big candidate set, which improves both the accuracy and efficiency of the candidate set (the cumulative accuracy of the top ten candidates is greatly boosted). Then, to further improve accuracy, word-based bigram postprocessing (WBP) is executed on a small candidate set. This method obtains high accuracy while paying attention to postprocessing speed at the same time. Experimental results for three Chinese scripts (about 66,000 characters in total) demonstrate the effectiveness of our method: character-based bigram postprocessing improves accuracy from 81.58% to 94.50%, and the cumulative accuracy of the top ten candidates rises from 94.33% to 98.25%. After WBP, 95.75% accuracy is achieved, which is equivalent to the accuracy of WBP executed on a big candidate set. However, our method is more than 100 times faster than that of WBP.
Xiaoqing Ding, Chew Lim Tan
ACM Trans. Asian Lang. Inf. Process.2
2001 Automatic Text Location in Natural Scene Images
abstract
A new connected component based segmentation algorithm which automatically extracts text regions from natural scene images is proposed in this paper. This approach utilizes a multichannel decomposition method to locate text blocks in complex backgrounds. Block alignment analysis and recognition confidence values are used in the combination and identification of the connected components. The algorithm is applied to a test image database and shows promising results.
Xiaoqing Ding, Youshou Wu
ICDAR2
2001 An Automatic Performance Evaluation Method for Document Page Segmentation
abstract
Automatic performance evaluation for a document page segmentation module is necessary, as OCR products are used to manipulate large scale of documents with complex layout, especially for newspapers. The paper presents a region-based method to evaluate the performance of a page segmentation module by analyzing geometric region relationships between the segmentation results and the preset ground-truth. The ground-truth is not only the correct answer to page segmentation, but also the comparison benchmark of the automatic evaluation, so it has more restricted geometric constraints. The region-matching algorithm is realized by searching the equal region in the segmentation results for each region in the ground-truth. The performance parameters are calculated based on the matching results. An experiment is given to test two page segmentation modules in a popular Chinese OCR product-THOCR2000, and the results show this method is effective.
Liangrui Peng, Changsong Liu, Xiaoqing Ding, Jirong Zheng
ICDAR4
2001 Character Extraction and Recognition in Natural Scene Images
abstract
With the proposal of the concept of a "smart camera", character recognition in natural scene images has become an interesting but difficult task nowadays. In this paper, we propose an algorithm for extracting characters from text regions of natural scene images with complex backgrounds. Our method first clusters the color feature vectors of the text regions into a number of color classes by applying a modified coarse-fine fuzzy c-means algorithm. Then, different slices are constructed according to these color classes. Characters are eventually extracted from the images using the information of segmentation and recognition. Some experiments have shown that this method is a promising starting point for such applications.
Xiaoqing Ding, Changsong Liu
ICDAR2
2001 Offline Handwritten Character Recognition Based on Discriminative Training of Orthogonal Gaussian Mixture Model
abstract
The statistical approach to offline handwritten character recognition, in which classifier design is very important, has been used widely. To approximate the class conditional density more precisely, it can be represented by an orthogonal Gaussian mixture model (OGMM). The parameters of the OGMM are commonly estimated by an expectation-maximization algorithm (EM), which converges to the maximum likelihood estimation (MLE). Since the MLE cannot directly minimize the classification errors as the goal of classifier design, a discriminative training method based on the minimum classification error (MCE) criterion is used to adjust the parameters of the OGMM. In order to achieve good generalization, the complexity of the classifier, namely the number of components in the OGMM, is determined by following the structure risk minimization (SRM) principle. Finally, the recognition performance is demonstrated by applying it to the handwritten numerals in the NIST database.
Xiaoqing Ding, Jiayong Zhang
ICDAR2
2001 Form Frame Line Detection with Directional Single-Connected Chain
abstract
In this paper, a novel form frame line detection algorithm is proposed based on the directional single-connected chain (DSCC). Defined as an array of black pixel run-lengths, DSCC works very well as an image structure element or vector in our vectorization algorithm. By merging multiple DSCCs under some constraints, we are able to extract the form frame lines automatically yet fast. The speed of our algorithm is comparable with some well-known projection methods. Experiments show that our algorithm is fast, resistant to moderate serious line breaks and can detect diagonal lines with any angle.
Yefeng Zheng 0001, Changsong Liu, Xiaoqing Ding, Shiyan Pan
ICDAR3
2001 Offline handwritten numeral recognition using orthogonal Gaussian mixture model
abstract
In the statistical approach to offline handwritten numeral recognition, we use the Gaussian mixture model (GMM) to approximate arbitrary class conditional probability density. For simplification, the GMM is assumed to be diagonal covariance matrices. In the case of the features of handwritten numerals being correlated statistically, a large number of mixture components are usually needed to obtain a good approximation. To solve this problem, the feature vectors are first transformed to the space spanned by the eigenvectors of the covariance matrix so that the correlation among the elements is reduced, namely orthogonal transformation. This GMM is defined as orthogonal Gaussian mixture model (OGMM). Finally, the effectiveness of this algorithm is demonstrated by applying it to the NIST database.
Xiaoqing Ding
ICIP (1)2
2001 Location and interpretation of destination addresses on handwritten Chinese envelopes
Junliang Xue, Xiaoqing Ding, Changsong Liu, Weiwei Qian
Pattern Recognit. Lett.2
2000 Analysis, Understanding and Representation of Chinese Newspaper with Complex Layout
abstract
Layout analysis, understanding and representation are important problems when transforming paper document to its electronic version. For a Chinese newspaper with a complex layout, a bottom-up algorithm of layout analysis based on nearest neighbor connect-strength and line confidence is proposed. We also propose a rule-based growing algorithm used for layout understanding. The implementation of layout representation is discussed at the same time. Using these algorithms with a Chinese OCR engine, we developed a complete system that can be used for automatic electronic publishing. The algorithms were proved to be efficient and practical by experimental results and by a practically running system.
Xiaoqing Ding
ICIP2
2000 A Method for Segmentation of Cursive Handwritings and its Application to Character Shape Extraction
abstract
Image processing and shape analysis are important and useful in many document processing applications, such as segmentation of connected characters, recognition of confusing character shapes, etc. A skeleton segmentation algorithm and a shape analysis method based on the segmentation result is presented. The motivation of the work presented in this paper is the idea that the basic elements in character shapes are line segments with uniform directional attributes. We define these line segments as image cells. An algorithm is carefully designed to segment the character skeletons into image cells robustly. Based on the image cells, a new selective attention approach for character shape extraction is proposed. To guarantee the robustness of the proposed algorithm, the shape extraction problem is decomposed into two steps: attention focus area detection and shape growing. The effectiveness of the proposed algorithms is verified by experimental results.
Jiang Gao, Xiaoqing Ding
ICIP2
2000 Offline Handwritten Chinese Character Recognition Using Optimal Sampling Features
Xiaoqing Ding
ICMI2
2000 On Improvement of Feature Extraction Algorithms for Discriminative Pattern Classification
abstract
Two new feature extraction strategies-modified multiple discriminant analysis (MMDA) and difference principal component analysis (DPCA)-are presented and derived. The proposed algorithms are especially useful in automatic feature extraction from patterns in a small category set. Experimental results for recognition of Chinese character fonts and handwritten numerals using MMDA and DPCA are presented. Compared with the traditional algorithms, MMDA and DPCA provide more effective feature metrics for pattern discrimination in some settings.
Jiang Gao, Xiaoqing Ding
ICPR2
2000 Gray-Scale Character Image Recognition Based on Fuzzy DCT Transform Features
abstract
We propose a method for recognizing gray-scale text images with noise and of extraordinary low resolution. First, a fuzzy classification of pixels in the gray-scale character images is applied to form the fuzzy attributed pixel graph using integral ratio techniques. Second, a DCT transform is applied on this fuzzy graph and parts of the frequency domain components are selected and compressed to produce the features for recognition. By using the modified quadratic discriminant function in the classifier, we achieve satisfactory recognition results with the average recognition rate of 96.82% on 15 pixel/spl times/15 pixel character images and at the speed of over 30 characters/s.
Xiaoqing Ding, Changsong Liu
ICPR2
2000 Multi-Scale Feature Extraction and Nested-Subset Classifier Design for High Accuracy Handwritten Character Recognition
abstract
Both efficient representation and robust classification are essential to high-performance cursive offline handwritten Chinese character recognition. A novel multi-scale feature extraction method is presented based on the information entropy theory. Feature detection and compression are thus combined into an integrated optimization process. A series of optimal feature-spaces are constructed at varying values of the scale parameter and the best one is obtained with the maximum LDA criterion over the scale interval. For more robust classification, we introduce a structure into the Mahalanobis distance classifier and strike the balance between machine capacity and the performance on the training data in light of the ideas of structural risk minimization. A high accuracy recognition system is developed based on the new methods and for the first time, 4 widely different databases ranging from regular to completely unconstrained with several structural distortions and stroke connections are fully tested. The accuracies of 99.S% on regular database and 88.4% on cursive one at the speed of over 40 characters/s are achieved.
Jiayong Zhang, Xiaoqing Ding, Changsong Liu
ICPR2
1999 A Robust Skew Detection Algorithm for Grayscale Document Image
abstract
A fast and robust skew detection algorithm for gray-scale images is presented. The MCCSD (modified cross-correlation skew detection) algorithm uses horizontal and vertical cross-correlation simultaneously to deal with vertically laid-out text, which is commonly used in Chinese or Japanese documents. Instead of calculating the correlation for the entire image, we use small randomly selected regions to speed up the process. The region verification stage and further processing of auxiliary peaks make our method robust and reliable. An experiment shows that the proposed method has good results in detecting skew in various kinds of pages.
Xiaoqing Ding
ICDAR2
1999 A Segmentation Algorithm for Handwritten Chinese Character Strings
abstract
A new algorithm for segmenting handwritten Chinese character strings is presented. This approach is based on a split-and-merge strategy by locating possible ligatures between Chinese characters and merging the over-split components. The strategy is proposed by considering the structural properties of Chinese character strings. To guarantee the accuracy of the above segmentation algorithm, a recognizer is involved in order to aid the segmentation process. A maximum a posteriori probability index is derived for joint optimization of the segmentation and recognition results, and a dynamic programming algorithm-a modified level-building algorithm-is used to optimize this index. The whole algorithm is applied to a Chinese bank check amount recognition task, and some promising experimental results are obtained.
Jiang Gao, Xiaoqing Ding, Youshou Wu
ICDAR2
1999 Destination Address Block Location on Handwritten Chinese Envelope
abstract
Handwritten Chinese envelopes have three obvious characteristics: handwritten text lines are the main part of the envelope image; there is no distinct border between the destination address block (DAB) and sender address block (SAB); and handwritten Chinese characters are composed of complicated strokes with more freedom. A method for segmenting the envelope image into several candidate DABs and selecting the best one as the DAB is not yet suitable for handwritten Chinese envelopes. This paper presents a new method of DAB location on envelopes based on extraction of text lines, according to the characteristics of handwritten Chinese envelopes. Our method divides the DAB location problem into several simpler sub-problems and solves them respectively. Experimental results of extraction rate for DAB and average processing time on 500 samples directly sampled from mail sorting machines shows that this method is effective and takes less CPU time.
Junliang Xue, Xiaoqing Ding, Changsong Liu, Shiyan Pan, Hongwei Kong
ICDAR2
1999 Spatio-temporal Unified Model for On-line Handwritten Chinese Character Recognition
abstract
This paper presents a novel spatio-temporal modeling method for on-line handwritten Chinese character recognition. In this method, a statistical structure model (SSM) is used to describe the structural feature of Chinese characters from a probabilistic aspect, and an improved hidden Markov model (PCHMM) is employed to capture temporal information contained in ink. These two models are combined closely leading to a powerful spatio-temporal unified model (STUM), which has shown strong description ability and resulted in superior performance in the experiments where traditional models such as HMM (Hidden Markov Model) and ARG (Attributed Relational Graph) are also introduced and compared.
Xiaoqing Ding, Youshou Wu, Zhan Lu
ICDAR2
1998 A New Nonlinear Shape Normalization Method for Off-line Handwritten Chinese Character Recognition
Youbin Chen, Xiaoqing Ding, Youshou Wu
ACCV (1)2
1998 Evaluation and Application of Recognition Confidence in OCR
Xiaoqing Ding, Youbin Chen, Youshou Wu
ACCV (1)2
1998 Combinatorial Coarse Classification Method for OLCCR
Xiaoqing Ding, Youshou Wu, Fanxia Guo
ACCV (2)2
1998 Construct H-MFNN by Complexity Reduction Approach
Chi Fang, Mingsheng Zhao, Youshou Wu, Xiaoqing Ding
ICONIP4
1998 Knowledge Extraction Based Multilayer Feedforward Neural Networks
Youshou Wu, Chi Fang, Xiaoqing Ding
ICONIP4
1998 Adaptive confidence transform based classifier combination for Chinese character recognition
Xiaoqing Ding, Youshou Wu
Pattern Recognit. Lett.2
1997 Off-line Handwritten Chinese Character Recognition Based on Crossing Line Feature
abstract
A new method to extract crossing line features for off-line handwritten Chinese character recognition is proposed in this paper. Firstly, the input pattern is nonlinearly normalized in order to compensate for shape variations. Secondly, the normalized pattern is separated into four subpatterns according to the four kinds of elementary strokes. Thirdly, the four subpatterns are uniformly divided into M/spl times/M cells respectively. In every cell, the crossing lines are counted. Then a 4M/sup 2/-dimensional feature vector is generated. An off-line handwritten Chinese character recognition system is built based on this feature. Our experiments have demonstrated the effectiveness of the method proposed in this paper.
Youbin Chen, Xiaoqing Ding, Youshou Wu
ICDAR2
1997 Handwritten Numeral Recognition Using MFNN Based Multiexpert Combination Strategy
abstract
A novel unconstrained numeral recognition system using an MFNN (multilayered feedforward neural net)-based multi-expert combination strategy is proposed. Taking advantage of every method's confidence value as well as its output label, this strategy can improve the system's performance significantly with as few as two experts. Another novel feature of the system is that the multi-expert combination problem is converted into a new classification problem, and then the MFNNs are used to handle the combination adaptively. The proposed system has achieved promising results on the NIST database.
Xiaoqing Ding, Youshou Wu
ICDAR2
1997 Recognizing on-line handwritten Chinese character via FARG matching
abstract
The paper presents a novel method for online handwritten Chinese character recognition. In our method, each category of character is described by a fuzzy attributed relational graph (FARG). A relaxation algorithm is developed to match the input pattern with every FARG. For decision making, a similarity measure is established via statistical technique to calculate the matching degree between the input pattern and referenced FARG, according to which the recognition result is determined. The principle of our method makes it very robust against stroke connection and stroke order variation as well as stroke shape deformation. A database of 22530 samples collected from 6 subjects is used to test our recognition system which can recognize 3755 categories of Chinese characters. The result shows that our method is very effective: a top 1 recognition rate of 98.8% and a top 10 of 99.7% are reached.
Xiaoqing Ding, Youshou Wu
ICDAR2
1996 Speeding up the training process of the MFNN by optimizing the hidden layers' outputs
Mingsheng Zhao, Youshou Wu, Xiaoqing Ding
Neurocomputing3
1995 Realization of a high-performance bilingual Chinese-English OCR system
abstract
This paper focuses on the realization of a bilingual Chinese-English OCR system. First, the Twice-Segment Algorithm is used for segmentation of documents with Chinese and English characters mixed. Then the comprehensive recognition method is employed to improve the robustness of Chinese character recognition. A new measurement of robustness of OCR recognition performance is also put forward here. Finally, exciting experimental results are given.
Xiaoqing Ding, Fanxia Guo, Youshou Wu
ICDAR2
1995 Description and recognition of form and automated form data entry
abstract
In this paper we present a form description method, in which frame lines are used to constitute a so-called frame template, which reflects the structure of a form either topologically or geometrically. Relevant item traversal algorithm is then proposed to locate and label form's items. We have also developed a robust and fast frame line detection method to make this form description practical for form recognition. Experimental results show our approach provides an effective way to convert printed forms into computerized format or collect information for database from printed forms.
Xiaoqing Ding, Youshou Wu
ICDAR2
1993 Stroke extraction from Chinese characters by improved SLSA algorithm
abstract
When syntactic pattern recognition approach is used to recognize handwritten Chinese characters it is inevitable that strokes from Chinese characters are extracted first. In 1987 Gu and Wang presented a stroke-extracting algorithm named straight line sequence approximation (SLSA) algorithm. SLSA does not need a thinning process, so it is very fast. However, in the crossing regions among strokes SLSA often extracts false strokes and causes stroke-splitting. The improved SLSA algorithm described in this paper aims at solving this problem.
Fu Jin, Zheng-Ming Chai, Xiaoqing Ding
VCIP3