Xu-Yao Zhang

dblp:58/9924 · also XuYao Zhang · DBLP profile ↗
← Back
14ranked-venue papers in the field
2as first author
3since 2021 · last 2025
0000-0001-9260-188XORCID · conflict

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 12 (1 first)Data Mining & Knowledge Discovery · 2 (1 first)
YearPublicationVenuePosition
2025 OracleGCD: Generalized Category Discovery for Oracle Bone Scripts
Hetao Wu, Kunchi Li, Xu-Yao Zhang, Dahan Wang
ICDAR (3)3
2021 Adaptive Scaling for Archival Table Structure Recognition
Xiao-Hui Li 0012, Xu-Yao Zhang, Cheng-Lin Liu 0001
ICDAR (1)3
2021 Document Dewarping with Control Points
Guo-Wang Xie, Xu-Yao Zhang, Cheng-Lin Liu 0001
ICDAR (1)3
2020 Dewarping Document Image by Displacement Flow Estimation with Fully Convolutional Network
Guo-Wang Xie, Xu-Yao Zhang, Cheng-Lin Liu 0001
DAS3
2019 Cross-Modal Prototype Learning for Zero-Shot Handwriting Recognition
abstract
In contrast to machine recognizers that rely on training with large handwriting data, humans can recognize handwriting accurately on learning from few samples, and can even generalize to handwritten characters from printed samples. Simulating this ability in machine recognition is important to alleviate the burden of labeling large handwriting data, especially for large category set as in Chinese text. In this paper, inspired by human learning, we propose a cross-modal prototype learning (CMPL) method for zero-shot online handwritten character recognition: for unseen categories, handwritten characters can be recognized without learning from handwritten samples, but instead from printed characters. Particularly, the printed characters (one for each class) are embedded into a convolutional neural network (CNN) feature space to obtain prototypes representing each class, while the online handwriting trajectories are embedded with a recurrent neural network (RNN). Via cross-modal joint learning, handwritten characters can be recognized according to the printed prototypes. For unseen categories, handwritten characters can be recognized by only feeding a printed sample per category. Experiments on a benchmark Chinese handwriting database have shown the effectiveness and potential of the proposed method for zero-shot handwriting recognition.
Xiang Ao 0002, Xu-Yao Zhang, Hong-Ming Yang, Cheng-Lin Liu 0001
ICDAR2
2019 CASIA-AHCDB: A Large-Scale Chinese Ancient Handwritten Characters Database
abstract
This paper introduces a Chinese Ancient Handwritten Characters Database (CASIA-AHCDB) for character recognition research. The database was built by annotating 11,937 pages of Chinese ancient handwritten documents. It consists of more than 2.2 million annotated handwritten character samples of 10,350 categories. According to the source of these documents, the database is divided into two datasets of different styles: Complete Library in Four Sections (AHCDB-style1) and Ancient Buddhist Scriptures (AHCDB-style2). Each dataset can be divided into three parts based on its applications. The first part, called basic category set, contains samples of common categories in two datasets, and is suitable for basic character recognition task. The second part, called enhanced category set, is mainly used for open-set character recognition task based on the basic character recognition. The third part, called the reserved category set, can be used in many pattern recognition tasks in the future. Based on the large category set, the various writing styles and the imbalanced sample number per category, CASIA-AHCDB can also be used for various classification and learning tasks such as transfer learning, few-shot learning. We performed experiments of basic character recognition on the basic category set, and report the results for benchmark. More techniques can be evaluated on this challenging database in the future.
Dahan Wang, Xu-Yao Zhang, Zhaoxiang Zhang 0001, Cheng-Lin Liu 0001
ICDAR4
2019 Fast Text/non-Text Image Classification with Knowledge Distillation
abstract
How to efficiently judge whether a natural image contains texts or not is an important problem. Since text detection and recognition algorithms are usually time-consuming, and it is unnecessary to run them on images that do not contain any texts. In this paper, we investigate this problem from two perspectives: the speed and the accuracy. First, to achieve high speed for efficient filtering large number of images especially on CPU, we propose using small and shallow convolutional neural network, where the features from different layers are adaptively pooled into certain sizes to overcome difficulties caused by multiple scales and various locations. Although this can achieve high speed but its accuracy is not satisfactory due to limited capacity of small network. Therefore, our second contribution is using the knowledge distillation to improve the accuracy of the small network, by constructing a larger and deeper neural network as teacher network to instruct the learning process of the small network. With the above two strategies, we can achieve both high speed and high accuracy for filtering scene text images. Experimental results on a benchmark dataset have shown the effectiveness of our method: the teacher network yields state-of-the-art performance, and the distilled small network achieves high performance while maintaining high speed which is 176 times faster on CPU and 3.8 times faster on GPU than a compared benchmark method.
Miao Zhao, Rui-Qi Wang, Xu-Yao Zhang, Linlin Huang 0001, Jean-Marc Ogier
ICDAR4
2018 Image-to-Markup Generation via Paired Adversarial Learning
Jin-Wen Wu, Yan-Ming Zhang 0001, Xu-Yao Zhang, Cheng-Lin Liu 0001
ECML/PKDD (1)4
2017 Handwriting Style Mixture Adaptation
abstract
In handwriting recognition, the test data usually come from multiple writers which are not shown in the training data. Therefore, adapting the base classifier towards the new style of each writer can significantly improve the generalization performance. Traditional writer adaptation methods usually assume that there is only one writer (one style) in the test data, and we call this situation as style-clear adaptation. However, a more common situation is that multiple handwriting styles exist in the test data, which is widely appeared in multi-font documents and handwriting data produced by the cooperation of multiple writers. We call the adaptation in this situation as style-mixture adaptation. To deal with this problem, in this paper, we propose a novel method called K-style mixture adaptation (K-SMA) with the assumption that there are totally K styles in the test data. Specifically, we first partition the test data into K groups (style clustering) according to their style consistency, which is measured by a newly designed style feature that can eliminate class (category) information and keep handwriting style information. After that, in each group, a style transfer mapping (STM) is used for writer adaptation. Since the initial style clustering may be not reliable, we repeat this process iteratively to improve the adaptation performance. The K-SMA model is fully unsupervised which do not require either the class label or the style index. Moreover, the K-SMA model can be effectively combined with the benchmark convolutional neural network (CNN) models. Experiments on the online Chinese handwriting database CASIA-OLHWDB demonstrate that K-SMA is an efficient and effective solution for style-mixture adaptation.
Hong-Ming Yang, Xu-Yao Zhang, Cheng-Lin Liu 0001
ICDAR2
2013 Extraction of Serial Numbers on Bank Notes
abstract
The study of RMB (renminbi bank note, the paper currency used in China) serial number recognition draws more and more attention in recent years, for reducing financial crime, improving financial market stability and social security. The accuracy of RMB recognition relies heavily on the extraction, which is a challenging problem due to background variations and uneven illumination. In this paper, we present a new system that extracts the RMB characters directly from scanned RMB images. First, two different techniques, namely skew correction and orientation identification are used to detect the region which contains RMB serial number. Then the detected text region is binarized by a combined thresholding technique. After that, a local contrast average method is introduced to extract the RMB characters from the binarization result. The experiments demonstrate that the proposed binarization method outperforms other well-known methods. For character extraction, we report an overlap-recall rate of 79.68% and an overlap-precision rate of 98.10% respectively.
Bo-Yuan Feng, Mingwu Ren, Xu-Yao Zhang, Ching Y. Suen
ICDAR3
2013 ICDAR 2013 Chinese Handwriting Recognition Competition
abstract
This paper describes the Chinese handwriting recognition competition held at the 12th International Conference on Document Analysis and Recognition (ICDAR 2013). This third competition in the series again used the CASIA-HWDB/OLHWDB databases as the training set, and all the submitted systems were evaluated on closed datasets to report character-level correct rates. This year, 10 groups submitted 27 systems for five tasks: classification on extracted features, online/offline isolated character recognition, online/offline handwritten text recognition. The best results (correct rates) are 93.89% for classification on extracted features, 94.77% for offline character recognition, 97.39% for online character recognition, 88.76% for offline text recognition, and 95.03% for online text recognition, respectively. In addition to the test results, we also provide short descriptions of the recognition methods and brief discussions on the results.
Qiufeng Wang 0001, Xu-Yao Zhang, Cheng-Lin Liu 0001
ICDAR3
2013 Locally Smoothed Modified Quadratic Discriminant Function
abstract
Modified quadratic discriminant function (MQDF) is a state-of-the-art classifier for handwriting recognition. However, the big gap between accuracies on training and testing sets indicates that MQDF has a good capability to fit training data but the generalization performance is not promising. To solve this problem, we propose a new model called locally smoothed modified quadratic discriminant function (LSMQDF) by smoothing the covariance matrix of each class with its nearest neighbor classes. LSMQDF can be viewed as a regularization to avoid over-fitting. The covariance matrix estimated by local smoothing is more accurate and robust. LSMQDF can be also viewed as an extension of the global smoothing method, namely regularized discriminant analysis (RDA). Experiments on both offline and online Chinese handwriting databases demonstrate that: with local smoothing, the accuracy on training set is decreased (over-fitting avoided), and the accuracy on testing set is improved significantly and consistently (generalization improved).
Xu-Yao Zhang, Cheng-Lin Liu 0001
ICDAR1
2013 Feature Transformation with Class Conditional Decorrelation
abstract
The well-known feature transformation model of Fisher linear discriminant analysis (FDA) can be decomposed into an equivalent two-step approach: whitening followed by principal component analysis (PCA) in the whitened space. By proving that whitening is the optimal linear transformation to the Euclidean space in the sense of minimum log-determinant divergence, we propose a transformation model called class conditional decor relation (CCD). The objective of CCD is to diagonalize the covariance matrices of different classes simultaneously, which is efficiently optimized using a modified Jacobi method. CCD is effective to find the common principal components among multiple classes. After CCD, the variables become class conditionally uncorrelated, which will benefit the subsequent classification tasks. Combining CCD with the nearest class mean (NCM) classification model can significantly improve the classification accuracy. Experiments on 15 small-scale datasets and one large-scale dataset (with 3755 classes) demonstrate the scalability of CCD for different applications. We also discuss the potential applications of CCD for other problems such as Gaussian mixture models and classifier ensemble learning.
Xu-Yao Zhang, Kaizhu Huang, Cheng-Lin Liu 0001
ICDM1
2011 Perceptron Learning of Modified Quadratic Discriminant Function
abstract
Modified quadratic discriminant function (MQDF) is the state-of-the-art classifier in handwritten character recognition. Discriminative learning of MQDF can further improve its performance. Recent advances justify the efficacy of minimum classification error criteria in learning MQDF (MCE-MQDF). We provide an alternative choice to MCE-MQDF based on the Perceptron learning (PL-MQDF). For better generalization performance, we propose a new dynamic margin regularization. To relieve the heavy burden in training process, active set technique is employed, which can save most of the computation with negligible loss in accuracy. In experiments on handwritten digit datasets and a large-scale Chinese handwritten character database, the proposed PL-MQDF was demonstrated superior in both error reduction and training speedup.
Tong-Hua Su, Cheng-Lin Liu 0001, Xu-Yao Zhang
ICDAR3