EDBT 2026 Demo / reviewers in the wild / expert
Baihua Xiao
dblp:13/776
· DBLP profile ↗
100ranked-venue papers
1as first author
13since 2021 · last 2025
0000-0003-3941-1141ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 53 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 44 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 6 since 2021Databases, data management, data science and information retrieval · 10Security and privacy · 3Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ProSyno: context-free prompt learning for synonym discovery
Hongyun Bao, Suncong Zheng, Yuqiao Liu 0003, Baihua Xiao, Dongyuan Lu |
Frontiers Comput. Sci. | 7 |
| 2025 | SDLFusion: A salient-aware differentiated learning network for infrared and visible image fusion
Xiaoting Fan, Shuang Liu 0001, Baihua Xiao |
Knowl. Based Syst. | 5 |
| 2025 | Cloud-Type Classification Using Multimodal Integration Transformer Based on Cloud Images and Millimeter-Wave Cloud Radar ObservationsabstractThe existing methods fail to simultaneously utilize the appearance information and the internal structure of clouds for cloud type classification, resulting in incomplete cloud representation. In this paper, we exploit cloud images and Millimeter-wave Cloud Radar (MMCR) observations for cloud type classification, and propose a novel Transformer network named Multi-modal Integration Transformer (MMITrans) to describe completed information of clouds. To this end, we design MMITrans as three subnetworks, i.e., vision subnetwork, MMCR subnetwork and multi-modal fusion network. Specifically, we extract the visual features from the cloud images through the vision subnetwork. Meanwhile, we first convert MMCR observations into several cloud-related indicators, and propose the Indicator-Tokenization to effectively tokenize them and obtain the indicator features using the MMCR subnetwork. Furthermore, we propose the Multi-modal Cross Attention in the multi-modal fusion network to sufficiently fuse the visual features and the indicator features in a multiple-input way. We perform a series of experiments on Cloud images and Millimeter-wave cloud radar observations Dataset, i.e., CMD-Beijing and CMD-Gansu, and the experimental results demonstrate the superiority of the proposed MMITrans. Shuang Liu 0001, Zeyu Zang, Zhong Zhang 0001, Shuzhen Hu, Baihua Xiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Completed Interaction Networks for Pedestrian Trajectory PredictionabstractThe social and environmental interactions, as well as the pedestrian goal are crucial for pedestrian trajectory prediction. This is because they could learn both complex interactions in the scenes and the intentions of the pedestrians. However, most existing methods either learn the one-moment social interactions, or supervise the pedestrian trajectories using long-term goal, resulting in suboptimal prediction performances. In this paper, we propose a novel network named Completed Interaction Network (CINet) to simultaneously consider the social interactions in all moments, the environmental interactions and the short-term goal of pedestrians in a unified framework for pedestrian trajectory prediction. Specifically, we propose the Spatio-Temporal Transformer Layer (STTL) to fully mine the spatio-temporal information among historical trajectories of all pedestrians in order to obtain the social interactions in all moments. Additionally, we present the Gradual Goal Module (GGM) to capture the environmental interactions under the supervision of the short-term goal, which is beneficial to understanding the intentions of the pedestrian. Afterwards, we employ the cross-attention to effectively integrate the all-moment social and environmental interactions. The experimental results on three standard pedestrian datasets, i.e., ETH/UCY, SDD and inD demonstrate that our method achieves a new state-of-the-art performance. Furthermore, the visualization results indicate that our method could predict trajectories more reasonably in complex scenarios such as sharp turns, infeasible areas and so on. Zhong Zhang 0001, Jianglin Zhou, Shuang Liu 0001, Baihua Xiao |
IEEE Trans. Multim. | 4 |
| 2024 | A comprehensive review of image retargeting
Xiaoting Fan, Zhong Zhang 0001, Baihua Xiao, Tariq S. Durrani |
Neurocomputing | 4 |
| 2024 | Completed Part Transformer for Person Re-IdentificationabstractRecently, part information of pedestrian images has been demonstrated to be effective for person re-identification (ReID), but the part interaction is ignored when using Transformer to learn long-range dependencies. In this article, we propose a novel transformer network named Completed Part Transformer (CPT) for person ReID, where we design the part transformer layer to learn the completed part interaction. The part transformer layer includes the intra-part layer and the part-global layer, where they consider long-range dependencies from the aspects of the intra-part interaction and the part-global interaction, simultaneously. Furthermore, in order to overcome the limitation of fixed number of the patch tokens in the transformer layer, we propose the Adaptive Refined Tokens (ART) module to focus on learning the interaction between the informative patch tokens in the pedestrian image, which improves the discrimination of the pedestrian representation. Extensive experimental results on four person ReID datasets, i.e., MSMT17, Market1501, DukeMTMC-reID, and CUHK03, demonstrate that the proposed method achieves a new state-of-the-art performance, e.g., it achieves 68.0% mAP and 84.6% Rank-1 accuracy on MSMT17. Zhong Zhang 0001, Di He 0008, Shuang Liu 0001, Baihua Xiao, Tariq S. Durrani |
IEEE Trans. Multim. | 4 |
| 2023 | Cross-modality person re-identification using hybrid mutual learningabstractAbstract Cross‐modality person re‐identification (Re‐ID) aims to retrieve a query identity from red, green, blue (RGB) images or infrared (IR) images. Many approaches have been proposed to reduce the distribution gap between RGB modality and IR modality. However, they ignore the valuable collaborative relationship between RGB modality and IR modality. Hybrid Mutual Learning (HML) for cross‐modality person Re‐ID is proposed, which builds the collaborative relationship by using mutual learning from the aspects of local features and triplet relation. Specifically, HML contains local‐mean mutual learning and triplet mutual learning where they focus on transferring local representational knowledge and structural geometry knowledge so as to reduce the gap between RGB modality and IR modality. Furthermore, Hierarchical Attention Aggregation is proposed to fuse local feature maps and local feature vectors to enrich the information of the classifier input. Extensive experiments on two commonly used data sets, that is, SYSU‐MM01 and RegDB verify the effectiveness of the proposed method. Zhong Zhang 0001, Sen Wang 0007, Shuang Liu 0001, Baihua Xiao, Tariq S. Durrani |
IET Comput. Vis. | 5 |
| 2023 | Unsupervised Domain Adaptation for Remote Sensing Image Segmentation Based on Adversarial Learning and Self-TrainingabstractThere is a large amount of out-of-distribution data (OOD) in remote sensing, which hinders high-accuracy segmentation models under the assumption of independent identical distribution (i.i.d.) from stable and reliable performance in real-world remote sensing applications. And Domain Adaptation (DA) is presented to seamlessly extend classifiers to the label-scarce target domain in the presence of the label-sufficient source domain with different data distributions. However, given that the domain shift, i.e. the distribution difference between the two domains, is more serious in remote sensing images, the current DA methods for image segmentation in Computer Vision (CV) typically perform unsatisfactorily in remote sensing, even suffering from the negative domain alignment. To this end, this paper proposes the Self-Training Adversarial Domain Adaptation (STADA) method for remote sensing image segmentation, which not only performs adversarial learning to extract domain-invariant features, but also implements Self-Training using pseudo-labels in the target domain denoised by the conditional adversarial loss for classifier adaptation. The ISPRS and WHU datasets are employed to conduct extensive experiments to investigate the effectiveness of STADA and the specific effect of its each DA component. And the experimental results demonstrate that STADA outperforms other state-of-the-art DA methods in the remote sensing image segmentation task. Chenbin Liang, Bo Cheng 0005, Baihua Xiao, Yunyun Dong |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | Multilevel Heterogeneous Domain Adaptation Method for Remote Sensing Image SegmentationabstractDue to more abundant data sources, more various objects of interest, and more time-consuming annotations, there is a large amount of out-of-distribution (OOD) data in the remote sensing field, on which the performance of high-accuracy image segmentation models trained under ideal experimental conditions generally degrades dramatically. Domain adaptation (DA) consequently comes into being, which aims to learn the predictor for the label-scarce target domain of interest with the help of the label-sufficient source domain in the presence of the distribution difference, namely, domain shift, between the two domains. However, the off-the-shelf DA methods for image segmentation not only struggle to cope with the more complex domain shift problems in remote sensing imagery but also almost cannot process heterogeneous data directly without information loss. While the current heterogeneous DA methods mostly still rely on some supervision information from the target domain, which is typically inaccessible in the real world. To overcome these drawbacks, we propose the multilevel heterogeneous unsupervised DA (UDA) method, termed MHDA, which unifies the instance-level DA based on cycle consistency, the feature-level DA based on contrastive learning, and the decision-level DA based on task consistency into a framework to more effectively handle the complex domain shift and heterogeneous data. After that, extensive DA experiments are conducted on the International Society for Photogrammetry and Remote Sensing (ISPRS) dataset, the BigCity dataset constructed by ourselves, and the Wuhan University (WHU) dataset, to explore the effect of each module in MHDA, the necessity of heterogeneous DA, and the effectiveness of multilevel DA. And the results demonstrate that MHDA can achieve superior performance on the remote sensing image segmentation task, compared with several state-of-the-art DA methods. Chenbin Liang, Bo Cheng 0005, Baihua Xiao, Yunyun Dong, Jinfen Chen |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Ground-Based Cloud Detection Using Multiscale Attention Convolutional Neural NetworkabstractCloud detection plays a significant role in ground-based remote sensing observation, and it is quite challenging due to the variations in illumination and cloud form, and the vague boundaries between cloud and sky. In this letter, we propose a novel deep model named multiscale attention convolutional neural network (MACNN) for ground-based cloud detection, which possesses a symmetric encoder–decoder structure. For accurate cloud detection, we design the multiscale module in MACNN to obtain different receptive fields by using different hole rates for the filters, and meanwhile, we propose the attention module in MACNN to learn the attention coefficients in order to reflect different importance of pixels. Furthermore, we release the Tianjin Normal University (TJNU) cloud detection database (TCDD) to provide a comparative study for different methods, and to the best of our knowledge, it is the largest cloud detection database. We conduct a series of experiments on the TCDD, and the experimental results demonstrate that the proposed MACNN outperforms state-of-the-art methods in five quantitative evaluation criteria. Zhong Zhang 0001, Shuzhen Yang, Shuang Liu 0001, Baihua Xiao, Xiaozhong Cao |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Cross-Domain Person Re-Identification Using Heterogeneous Convolutional NetworkabstractPerson re-identification (Re-ID) is a challenging task due to variations in pedestrian images, especially in cross-domain scenarios. The existing cross-domain person Re-ID approaches extract the feature from single pedestrian image, but they ignore the correlations among pedestrian images. In this paper, we propose Heterogeneous Convolutional Network (HCN) for cross-domain person Re-ID, which learns the appearance information of pedestrian images and the correlations among pedestrian images simultaneously. To this end, we first utilize Convolutional Neural Network (CNN) to extract the appearance features for pedestrian images. Then we construct a graph in the target dataset where the appearance features are treated as the nodes and the similarity represents the linkage between the nodes. Afterwards, we propose Dual Graph Convolution (DGConv) to explicitly learn the correlation information from the similar and dissimilar samples, which could avoid the over-smoothing caused by the fully connected graph. Furthermore, we design HCN as a multi-branch structure to mine the structural information of pedestrians. We conduct extensive evaluations for HCN on three datasets, i.e. Market-1501, DukeMTMC-reID and MSMT17, and the results demonstrate that HCN is superior to the state-of-the-art methods. Zhong Zhang 0001, Shuang Liu 0001, Baihua Xiao, Tariq S. Durrani |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | GCN-Based Semantic Segmentation Method for Mine Information Extraction in GAOFEN-1 ImageryabstractMine information extraction is of great significance to the construction of ecological civilization, the dynamic monitoring of mine development and the scientific management of mineral resources. With the emergence of high spatial resolution remote sensing imagey, traditional machine learning method gradually cannot meet the increasing demands of image interpretation. CNN-based semantic segmentation method provides a great solution for this issue. With the deepening of network layers, more the high-level features can be obtained, which brings the outstanding performance of many computer vision tasks, but also leads to the loss of structural information, which is crucial for mine information extraction. Therefore, in order to improve these drawbacks, we proposed a novel network based on the classical semantic segmentation network, SegNet, and Graph Convolutional Network (GCN) that makes our method more sensitive to structural information. Then, taking the iron mine located in Qian'an City, Hebei Province as experimental area, we employed our method to extract five mainly mine objects: stopes, ore heap, waste dump, tailings reservoir and concentration based on GF-1 imagery. Compared with SegNet, the mIoU of our method was improved by about 5% on our dataset and was improved by about 2.2% on PASCAL VOC2012 dataset. Chenbin Liang, Baihua Xiao, Bo Cheng 0005 |
IGARSS | 2 |
| 2021 | Unconstrained end-to-end text reading with feature rectification
Yanna Wang, Chunheng Wang, Baihua Xiao, Cunzhao Shi |
Pattern Recognit. Lett. | 4 |
| 2020 | DetectGAN: GAN-based text detector for camera-captured document images
Jinyuan Zhao, Yanna Wang, Baihua Xiao, Cunzhao Shi, Fuxi Jia, Chunheng Wang |
Int. J. Document Anal. Recognit. | 3 |
| 2020 | Selective feature connection mechanism: Concatenating multi-layer CNN features with a feature selector
Yanna Wang, Chunheng Wang, Cunzhao Shi, Baihua Xiao |
Pattern Recognit. Lett. | 5 |
| 2020 | Adversarial learning based attentional scene text recognizer
Jinyuan Zhao, Yanna Wang, Baihua Xiao, Cunzhao Shi, Jingzhong Jiang, Chunheng Wang |
Pattern Recognit. Lett. | 3 |
| 2020 | Fuzzy Multilayer Clustering and Fuzzy Label Regularization for Unsupervised Person ReidentificationabstractUnsupervised person reidentification has received more attention due to its wide real-world applications. In this paper, we propose a novel method named fuzzy multilayer clustering (FMC) for unsupervised person reidentification. The proposed FMC learns a new feature space using a multilayer perceptron for clustering in order to overcome the influence of complex pedestrian images. Meanwhile, the proposed FMC generates fuzzy labels for unlabeled pedestrian images, which simultaneously considers the membership degree and the similarity between the sample and each cluster. We further propose the fuzzy label regularization (FLR) to train the convolutional neural network (CNN) using pedestrian images with fuzzy labels in a supervised manner. The proposed FLR could regularize the CNN training process and reduce the risk of overfitting. The effectiveness of our method is validated on three large-scale person reidentification databases, i.e., Market-1501, DukeMTMC-reID, and CUHK03. Zhong Zhang 0001, Meiyan Huang, Shuang Liu 0001, Baihua Xiao, Tariq S. Durrani |
IEEE Trans. Fuzzy Syst. | 4 |
| 2019 | Document image binarization with cascaded generators of conditional generative adversarial networks
Jinyuan Zhao, Cunzhao Shi, Fuxi Jia, Yanna Wang, Baihua Xiao |
Pattern Recognit. | 5 |
| 2019 | A Selection Criterion for the Optimal Resolution of Ground-Based Remote Sensing Cloud Images for Cloud ClassificationabstractIn ground-based remote sensing cloud image observation, images with the highest possible resolution are captured to obtain sufficient information about clouds. However, when features are extracted and classification is performed on the basis of the original images, a high-resolution probably means a high (or even more, unacceptable) computation cost. In practical application, a simple and commonly adopted method is to appropriately resize the original image to a version with a decreased resolution. An inevitable problem is whether useful information is lost in this resizing operation. This paper demonstrates that information loss is inevitable and poor classification results may be obtained from the analysis of local binary pattern (LBP) histogram features. However, this problem has been always neglected in previous studies, and the original image is arbitrarily resized without any criterion. In particular, the histogram features based on LBPs actually reflect the distribution of features. Thus, a criterion based on the Kullback-Leibler divergence between LBP histograms from the original and resized images and a penalty term imposed on the resolution are proposed to select the resolution of the resized image. The optimal resolution of the resized image can be selected by minimizing this criterion. Furthermore, experiments based on three ground-based remote sensing cloud image data sets with different original resolutions validate this criterion by analyzing the LBP histogram features. Yu Wang 0045, Chunheng Wang, Cunzhao Shi, Baihua Xiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2019 | Unsupervised Semantic-Based Aggregation of Deep Convolutional FeaturesabstractIn this paper, we propose a simple but effective semantic-based aggregation (SBA) method. The proposed SBA utilizes the discriminative filters of deep convolutional layers as semantic detectors. Moreover, we propose the effective unsupervised strategy to select some semantic detectors to generate the "soft region proposals," which highlight certain discriminative pattern of objects and suppress the noise of background. The final global SBA representation could then be acquired by aggregating the regional representations weighted by the selected "soft region proposals" corresponding to various semantic content. Our unsupervised SBA is easy to generalize and achieves excellent performance on various tasks. We conduct comprehensive experiments and show that our unsupervised SBA outperforms the state-of-the-art unsupervised and supervised aggregation methods on image retrieval, place recognition, and cloud classification. Chunheng Wang, Cheng-Zuo Qi, Cunzhao Shi, Baihua Xiao |
IEEE Trans. Image Process. | 5 |
| 2019 | Iterative Manifold Embedding Layer Learned by Incomplete Data for Large-Scale Image RetrievalabstractExisting manifold learning methods are not appropriate for image retrieval tasks, because most of them are unable to process query images and they have much greater computational cost especially for large-scale database. Therefore, we propose the iterative manifold embedding (IME) layer, of which the weights are learned offline by an unsupervised strategy, to explore the intrinsic manifolds by incomplete data. On the large-scale database that contains 27 000 images, the IME layer is more than 120 times faster than other manifold learning methods to embed the original representations at query time. We embed the original descriptors of database images that lie on manifold in a high-dimensional space into manifold-based representations iteratively to generate the IME representations in an offline learning stage. According to the original descriptors and the IME representations of database images, we estimate the weights of the IME layer by ridge regression. In the online retrieval stage, we employ the IME layer to map the original representation of a query image with an ignorable time cost (2 ms per image). We experiment on five public standard datasets for image retrieval. The proposed IME layer significantly outperforms the related dimension reduction methods and manifold learning methods. Without postprocessing, our IME layer achieves a boost in the performance of state-of-the-art image retrieval methods with postprocessing on most datasets, and needs less computational cost. The code is available at https://github.com/XJhaoren/IME_layer. Chunheng Wang, Cheng-Zuo Qi, Cunzhao Shi, Baihua Xiao |
IEEE Trans. Multim. | 5 |
| 2019 | Multi-Kernel Coupled Projections for Domain Adaptive Dictionary LearningabstractDictionary learning has produced state-of-the-art results in various classification tasks. However, if the training data have a different distribution than the testing data, the learned sparse representation might not be optimal. Recently, several domain-adaptive dictionary learning (DADL) methods and kernels have been proposed and have achieved impressive performance. However, the performance of these single kernel-based methods heavily depends heavily on the choice of the kernel, and the question of how to combine multiple kernel learning (MKL) with the DADL framework has not been well studied. Motivated by these concerns, in this paper, we propose a multi-kernel domain-adaptive sparse representation-based classification (MK-DASRC) and then use it as a criterion to design a multi-kernel sparse representation-based domain-adaptive discriminative projection method, in which the discriminative features of the data in the two domains are simultaneously learned with the dictionary. The purpose of this method is to maximize the between-class sparse reconstruction residuals of data from both domains, and minimize the within-class sparse reconstruction residuals of data in the low-dimensional subspace. Thus, the resulting representations can satisfactorily fit MK-DASRC and simultaneously display discriminability. Extensive experimental results on a series of benchmark databases show that our method performs better than the state-of-the-art methods. Yuhui Zheng, Guoqing Zhang 0002, Baihua Xiao, Fu Xiao 0001, Jianwei Zhang 0005 |
IEEE Trans. Multim. | 4 |
| 2018 | Unsupervised Part-Based Weighting Aggregation of Deep Convolutional Features for Image RetrievalabstractIn this paper, we propose a simple but effective semantic part-based weighting aggregation (PWA) for image retrieval. The proposed PWA utilizes the discriminative filters of deep convolutional layers as part detectors. Moreover, we propose the effective unsupervised strategy to select some part detectors to generate the "probabilistic proposals," which highlight certain discriminative parts of objects and suppress the noise of background. The final global PWA representation could then be acquired by aggregating the regional representations weighted by the selected "probabilistic proposals" corresponding to various semantic content. We conduct comprehensive experiments on four standard datasets and show that our unsupervised PWA outperforms the state-of-the-art unsupervised and supervised aggregation methods. Cunzhao Shi, Cheng-Zuo Qi, Chunheng Wang, Baihua Xiao |
AAAI | 5 |
| 2018 | An Effective Binarization Method for Disturbed Camera-Captured Document ImagesabstractMany researchers make numerous work on document image binarization. However, the binarization results of camera-captured document images remain to be improved due to many disturbances such as creases, noises and shadows. To binarize these images effectively, this paper proposes an adaptive local thresholding method which takes advantages of multi-level multi-scale local statistical information. By using the context information of multiple scales, the pixels in the image are classified by coarse to fine. The majority of background areas were removed by multiscale analysis of variance. For the text area, the binarization threshold is dynamically adjusted according to the estimated clarity. Our method can make the grayscale image binarization directly, without adding any postprocessing operation. The experimental results show that our method can significantly improve the performance of OCR system and is also suitable for degraded document images. Jinyuan Zhao, Cunzhao Shi, Fuxi Jia, Yanna Wang, Baihua Xiao |
ICFHR | 5 |
| 2018 | Joint Encoding LBP Features from Infrared and Visible-Light Cloud Image Observations for Ground-Based Cloud ClassificationabstractCloud type classification based on ground-based cloud image observations is an important task in atmospheric research. Currently, two kinds of cloud image observations with infrared and visible light images are widely used for cloud classification. However, they are only independently analyzed and simply compared in the current study. The useful information from these two kinds of images is not fully utilized and integrated. The classification performance could be improved if taking full advantage of the complementary information of these two observations. Thus, first, a database containing these two kinds of cloud images with same temporal resolution is released in this study. Then, a two-observation joint encoding strategy of LBP (local binary pattern) features is proposed to implement cloud classification by encoding the joint distribution of LBP patterns in different observations, which captures the correlation between two observations. Experimental results based on this database show the significant superiority of the proposed method compared to the results based on the single observation. Yu Wang 0045, Chunheng Wang, Cunzhao Shi, Baihua Xiao |
IGARSS | 4 |
| 2018 | CRF based text detection for natural scene images using convolutional neural network and context information
Yanna Wang, Cunzhao Shi, Baihua Xiao, Chunheng Wang, Cheng-Zuo Qi |
Neurocomputing | 3 |
| 2018 | Degraded document image binarization using structural symmetry of strokes
Fuxi Jia, Cunzhao Shi, Kun He 0002, Chunheng Wang, Baihua Xiao |
Pattern Recognit. | 5 |
| 2017 | Grayscale-Projection Based Optimal Character Segmentation for Camera-Captured Faint Text RecognitionabstractThe faint text document images possess shallow characters inherently and the camera-captured form introduces more degradations such as low-resolution, non-uniform illumination and out-of-focus blur, which make the text binarization very difficult. In this paper, we propose a grayscale-projection based optimal character segmentation method for camera-captured faint text recognition. Instead of extracting the character candidates, we use the gradient projection to extract a series of segmentation candidates which contain inter-character gaps and intra-character gaps as well. In order to select the optimal segmentation path from all possible situations, we construct a segmentation tree and set a evaluation score for each path. The score integrates the information of single point projection, overall distribution and recognition probability. Finally the optimal segmentation path is obtained by selecting the path with the highest score. We collect a faint text recognition dataset and evaluate our method on it. Experimental results show that our method outperforms the binary-projection method and the convolutional recurrent neural network approach in terms of text segmentation and recognition accuracy. Fuxi Jia, Cunzhao Shi, Yanna Wang, Chunheng Wang, Baihua Xiao |
ICDAR | 5 |
| 2017 | Learning Spatially Embedded Discriminative Part Detectors for Scene Character RecognitionabstractRecognizing scene character is extremely challenging due to various interference factors such as character translation, blur and uneven illumination, etc. Considering that characters are composed of a series of parts and different parts attract diverse attentions when people observe a character, we should assign different importance to each part to recognize scene character. In this paper, we propose a discriminative character representation by aggregating the responses of the spatially embedded salient part detectors. Specifically, we first extract the convolution activations from the pre-trained convolutional neural network (CNN). These convolutional activations are considered as the local descriptors of the character parts. Then we learn a set of part detectors and pick the distinctive convolutional activations which respond to the salient parts. Moreover, to alleviate the effect of character translation, rotation and deformation, etc, we assign a response region for each part detector and search the maximal response in this region. Finally, we aggregate the maximal outputs of all the salient part detectors to represent character. The experiments on three datasets show the effectiveness of the proposed method for scene character recognition. Yanna Wang, Cunzhao Shi, Baihua Xiao, Chunheng Wang |
ICDAR | 3 |
| 2017 | Spatial weighted fisher vector for image retrievalabstractSeveral recent works interpret convolutional features produced by deep convolutional neural networks as local descriptors. Existing high-dimensional aggregation based methods, e.g., Fisher Vector (FV) obtain inferior performance to pooling based methods in most situations, and we observe that it is mainly caused by the ignorance of spatial weights. In this paper, we propose a novel method named spatial weighted Fisher Vector (SWFV) to enhance the representation of FV by injecting the spatial weight map to FV. In addition, we further analyze the distribution of spatial weights and propose truncated spatial weighted FV (TSWFV). Experimental results on two benchmark datasets demonstrate that the two proposed methods achieve competitive results compared with other global representation based methods. Cheng-Zuo Qi, Cunzhao Shi, Chunheng Wang, Baihua Xiao |
ICME | 5 |
| 2017 | Ground-Based Cloud Detection Using Graph Model Built Upon SuperpixelsabstractCloud detection plays an important role in climate models, climate predictions, and meteorological services. Although researchers have given increasing efforts on cloud detection, the performance is still unsatisfactory due to the diverse nature of clouds. Considering the fact that one source of information (color or texture) is not enough to segment cloud from clear sky, in this letter, we propose a novel ground-based cloud detection method using graph model (GM) built upon superpixels to integrate multiple sources of information. First, we use the superpixel segmentation to divide the image into a series of subregions according to the color similarity and spatial continuity. Next, adjacent superpixels are merged according to their similarity of extracted features. Finally, we build a GM on the merged superpixels by considering each superpixel as a node and adding edges between neighboring ones. The unary cost is set according to the classification score of Random Forests, while pairwise cost reflects the penalties for color and texture discontinuity between neighboring components. The final segmentation could be acquired by minimizing the cost function. Moreover, the algorithm is computationally efficient as we use the superpixels rather than raw pixels as computation units. Experimental results demonstrate the effectiveness and efficiency of the proposed method for cloud detection. Cunzhao Shi, Yu Wang 0045, Chunheng Wang, Baihua Xiao |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2017 | Deep Convolutional Activations-Based Features for Ground-Based Cloud ClassificationabstractGround-based cloud classification is crucial for meteorological research and has received great concern in recent years. However, it is very challenging due to the extreme appearance variations under different atmospheric conditions. Although the convolutional neural networks have achieved remarkable performance in image classification, no one has evaluated their suitability for cloud classification. In this letter, we propose to use the deep convolutional activations-based features (DCAFs) for ground-based cloud classification. Considering the unique characteristic of cloud, we believe the local rich texture information might be more important than the global layout information and, thus, give a comprehensive evaluation of using both shallow convolutional layers-based features and DCAFs. Experimental results on two challenging public data sets demonstrate that although the realization of DCAF is quite straightforward without any use-dependent tricks, it outperforms conventional hand-crafted features considerably. Cunzhao Shi, Chunheng Wang, Yu Wang 0045, Baihua Xiao |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2017 | Measure for the Difference Between LBP Features Extracted From Original and Resized Cloud Images With Varying ResolutionsabstractCurrently, ground-based cloud images taken by using a whole-sky imager are especially popular in the field of meteorology because of their high resolution and accurate cloud information. Cloud images are natural texture images, and thus texture features based on local binary patterns (LBPs) are widely used to analyze texture images. However, the high-computation cost of extracting LBP features from high-resolution cloud texture images may make this technique unacceptable in practical image processing. A commonly adopted method is to resize the original image to an appropriate version with a decreased resolution. But this process will inevitably result in information loss. Accordingly, a measure based on the Kullback-Leibler (KL) divergence of the difference between LBP histogram features extracted from the original and resized images with varying resolutions is reported in this letter. Furthermore, a confidence interval technique is introduced to validate the significance of such difference. Experiments based on real ground-based cloud images show the measurement results of KL divergence in LBP features extracted from original and resized images. The experimental results indicate that images should be resized with caution when performing image processing. Yu Wang 0045, Cunzhao Shi, Chunheng Wang, Baihua Xiao |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2017 | Fisher vector for scene character recognition: A comprehensive evaluation
Cunzhao Shi, Yanna Wang, Fuxi Jia, Kun He 0002, Chunheng Wang, Baihua Xiao |
Pattern Recognit. | 6 |
| 2017 | Learning completed discriminative local features for texture classification
Zhong Zhang 0001, Shuang Liu 0001, Xing Mei, Baihua Xiao |
Pattern Recognit. | 4 |
| 2017 | Multi-order co-occurrence activations encoded with Fisher Vector for scene character recognition
Yanna Wang, Cunzhao Shi, Chunheng Wang, Baihua Xiao, Cheng-Zuo Qi |
Pattern Recognit. Lett. | 4 |
| 2017 | Logo Retrieval Using Logo Proposals and Adaptive Weighted PoolingabstractThis letter presents a novel approach for logo retrieval. Considering the fact that logo only occupies a small portion of an image, we apply Faster R-CNN to detect logo proposals first, and then use a two-step pooling strategy with adaptive weight to obtain an accurate global signature. The adaptive weighted pooling method can effectively balance the recall and precision of proposals by incorporating the probability of each proposal being a logo. Experimental results show that the proposed method interprets the similarity between query and database image more accurately and achieves state of the art performance. Cheng-Zuo Qi, Cunzhao Shi, Chunheng Wang, Baihua Xiao |
IEEE Signal Process. Lett. | 4 |
| 2016 | Document Image Binarization Using Structural Symmetry of StrokesabstractIn this paper, a novel local threshold binarization method using structural symmetry of strokes is proposed. Different from most existing local threshold methods which use the whole region to compute the threshold, we estimate the local threshold by only using the structural symmetric pixels (SSP) of the region so as to suppress the non-text pixels and maintain the text ones as well. The SSP is defined as those pixels around strokes whose gradient magnitudes are big enough and directions are opposite. As the gradient map is our basis for computing the SSP, we further propose to estimate background surface first and extract potential SSP in the compensated image so as to deal with degradations of document images such as uneven illumination, low contrast and stain. To prove the effectiveness of our method, tests on two public document image datasets are preformed and the experimental results show that our method outperforms other local threshold binarization approaches on both F-measure and PSNR. Fuxi Jia, Cunzhao Shi, Kun He 0002, Chunheng Wang, Baihua Xiao |
ICFHR | 5 |
| 2016 | OTSU guided adaptive binarization of CAPTCHA image using gamma correctionabstractGamma correction, a nonlinear operation, has long been used to code and decode luminance or tristimulus values in video or still image systems [1]. In this paper, we make the following observations: for CAPTCHA images which could not be well binarized using the threshold of OTSU, there exists a gamma corrected image which could be well segmented by the OTSU threshold and the value of the best gamma could be revealed by observing the maximal inter-class variance (MICV) values of different images transformed by different values of gamma. Concretely, we convert the R, G, B channels of the original CAPTCHA image with different gamma values and transform the color images to gray-level images. Each gray-level image could be then segmented by the threshold acquired by OTSU. By linking each gamma value with the corresponding maximal inter-class variance value, we could draw a changing curve of variance values versus gamma. The best gamma could be acquired by finding the point whose related MICV starts to change slowly. Moreover, the polarity of the image could also be revealed by the changing trend of the curve. Experimental results on different categories of CAPTCHA images demonstrate the effectiveness of the observations for binarizing the CAPTCHA images and telling the polarity as well. Cunzhao Shi, Yanna Wang, Baihua Xiao, Chunheng Wang |
ICPR | 3 |
| 2016 | Multiple Continuous Virtual Paths Based Cross-View Action RecognitionabstractIn this paper, we propose a novel method for cross-view action recognition via multiple continuous virtual paths which connect the source view and the target view. Each point on one virtual path is a virtual view which is obtained by a linear transformation of an action descriptor. All the virtual views are concatenated into an infinite-dimensional feature to characterize continuous changes from the source to the target view. To utilize these infinite-dimensional features directly, we propose a virtual view kernel (VVK) to compute the similarity between two infinite-dimensional features, which can be readily used to construct any kernelized classifiers. In addition, a constraint term is introduced to fully utilize the information contained in the unlabeled samples which are easier to obtain from the target view. The rationality behind the constraint is that any action video belongs to only one class. To further explore complementary visual information, we utilize multiple continuous virtual paths. The original source and target views are projected to different auxiliary source and target views using the random projection technique. Then we fuse all the VVKs generated from all pairs of auxiliary views. Our method is verified on the IXMAS and MuHAVi datasets, and the experimental results demonstrate that our method achieves better performance than the state-of-the-art methods. Zhong Zhang 0001, Shuang Liu 0001, Chunheng Wang, Baihua Xiao, Wen Zhou 0002 |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2015 | MRF based text binarization in complex images using stroke featureabstractThis paper presents a novel binarization technique for text images based on Markov Random Field (MRF) framework. We regard stroke as an obvious feature of text to produce clustering result, which will be optimized by MRF model combining color, texture, context features to get the final binarization. The main innovations of our method are: (1) the integrated image is split into sub-images on which we can automatically acquire seed pixels of foreground and background using stroke feature; and (2) diverse weights are attached to seed pixels according to their location information, then highly confident cluster centers of sub-image can be acquired by gathering weighted seeds. The experimental results show that our method is robust and accurate on both video and scene images. Yanna Wang, Cunzhao Shi, Baihua Xiao, Chunheng Wang |
ICDAR | 3 |
| 2015 | Scene text recognition by learning co-occurrence of strokes based on spatiality embedded dictionaryabstractText information contained in scene images is very helpful for high‐level image understanding. In this study, the authors propose to learn co‐occurrence of local strokes for scene text recognition by using a spatiality embedded dictionary (SED). Unlike spatial pyramid partitioning images into grids to incorporate spatial information, the authors SED associates every codeword with a particular response region and introduces more precise spatial information for robust character recognition. After localised soft coding and max pooling of the first layer, a sparse dictionary is learned to model co‐occurrence of several local strokes, which further improves classification performance. Experimental results on two scene character recognition datasets ICDAR2003 and CHARS74 K demonstrate that their character recognition method outperforms state‐of‐the‐art methods. Besides, competitive word recognition results are also reported for four benchmark word recognition datasets ICDAR2003, ICDAR2011, ICDAR2013 and street view text when combining their character recognition method with a conditional random field language model. Song Gao 0009, Chunheng Wang, Baihua Xiao, Cunzhao Shi, Wen Zhou 0002, Zhong Zhang 0001 |
IET Comput. Vis. | 3 |
| 2015 | Cross-view face recognition via structured dictionary based domain shiftabstractView variation is a major challenge in face recognition. In this study, the authors propose a novel cross‐view face recognition method by seeking potential intermediate domains between the source and target views to model the connection of varying‐views faces. Specifically, each intermediate domain is associated with a dictionary subspace. Learning proceeds in two phases. First, the authors discriminatively train a sub‐dictionary for each subclass of data, which then compose a structured dictionary of powerful reconstructive and discriminative capability on the source data. Secondly, the authors gradually adapt the source domain dictionary to the target domain by incrementally reducing the reconstruction error on the target data, which forms a smooth transition path connecting the source and target domains. Instead of updating the structured dictionary integrally, the authors develop a refined sub‐dictionary‐based updating algorithm, which makes the intermediate dictionaries fit on the target data better and faster. Finally, the authors apply invariant sparse codes across the source, intermediate and target domains to render domain‐shared representations, where the sample differences caused by view changes are reduced. Experiments on the CMU‐PIE and Multi‐PIE dataset demonstrate the effectiveness of the proposed method. Xue Chen 0002, Chunheng Wang, Baihua Xiao, Xinyuan Cai |
IET Comput. Vis. | 3 |
| 2015 | A character image restoration method for unconstrained handwritten Chinese character recognition
Yunxue Shao, Chunheng Wang, Baihua Xiao |
Int. J. Document Anal. Recognit. | 3 |
| 2015 | Ground-Based Cloud Detection Using Automatic Graph CutabstractGround-based cloud detection plays an essential role in meteorological research, and object segmentation techniques have recently been introduced to solve this issue. As a kind of object segmentation technique, interactive graph cut has emerged as a very powerful tool due to its effective segmentation ability. However, it requires users to provide labels for certain pixels as “object” or “background,” which inevitably prohibits automatic cloud detection in large-scale applications. In this letter, we focus on the issue of automatic cloud detection and propose a novel algorithm named as automatic graph cut. We treat clouds as a special kind of object and eliminate human labeling by two procedures. First, we adaptively compute the thresholds for each cloud image which automatically label some pixels as “cloud” or “clear sky” with high confidence. Then, those labeled pixels serve as hard constraint seeds for the following graph cut algorithm. The experimental results show that the proposed algorithm not only achieves better results than the state-of-the-art cloud detection algorithms but also achieves comparable results with the interactive segmentation algorithm. Shuang Liu 0001, Zhong Zhang 0001, Baihua Xiao, Xiaozhong Cao |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2015 | Automatic Cloud Detection for All-Sky Images Using Superpixel SegmentationabstractCloud detection plays an essential role in meteorological research and has received considerable attention in recent years. However, this issue is particularly challenging due to the diverse characteristics of clouds. In this letter, a novel algorithm based on superpixel segmentation (SPS) is proposed for cloud detection. In our proposed strategy, a series of superpixels could be obtained adaptively by SPS algorithm according to the characteristics of clouds. We first calculate a local threshold for each superpixel and then determine a threshold matrix for the whole image. Finally, cloud can be detected by comparing with the obtained threshold matrix. Experimental results show that our proposed algorithm achieves better performance than the current cloud detection algorithms. Shuang Liu 0001, Zhong Zhang 0001, Chunheng Wang, Baihua Xiao |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2015 | Robust relative attributes for human action recognition
Zhong Zhang 0001, Chunheng Wang, Baihua Xiao, Wen Zhou 0002, Shuang Liu 0001 |
Pattern Anal. Appl. | 3 |
| 2015 | Stroke Detector and Structure Based Models for Character Recognition: A Comparative StudyabstractCharacters, which are man-made symbols composed of strokes arranged in a certain structure, could provide semantic information and play an indispensable role in our daily life. In this paper, we try to make use of the intrinsic characteristics of characters and explore the stroke and structure-based methods for character recognition. First, we introduce two existing part-based models to recognize characters by detecting the elastic strokelike parts. In order to utilize strokes of various scales, we propose to learn the discriminative multi-scale stroke detector-based representation (DMSDR) for characters. However, the part-based models and DMSDR need to manually label the parts or key points for training. In order to learn the discriminative stroke detectors automatically, we further propose the discriminative spatiality embedded dictionary learning-based representation (DSEDR) for character recognition. We make a comparative study of the performance of the tree-structured model (TSM), mixtures-of-parts TSM, DMSDR, and DSEDR for character recognition on three challenging scene character recognition (SCR) data sets as well as two handwritten digits recognition data sets. A series of experiments is done on these data sets with various experimental setup. The experimental results demonstrate the suitability of stroke detector-based models for recognizing characters with deformations and distortions, especially in the case of limited training samples. Cunzhao Shi, Song Gao 0009, Meng-Tao Liu, Cheng-Zuo Qi, Chunheng Wang, Baihua Xiao |
IEEE Trans. Image Process. | 6 |
| 2014 | Still-to-Video face recognition via weighted scenario oriented discriminant analysisabstractIn Still-to-Video (S2V)face recognition, only a few high resolution images are enrolled for each subject, while the probe is videos of complex variations. As faces present distinct characteristics under different scenarios, recognition in the original space is obviously inefficient. In this paper, we propose a novel discriminant analysis method to learn separate mappings for different scenarios (still, video), and further pursue a common discriminant space based on these mappings. Concretely, by modeling each video as a set of local models, we form the scenario-oriented mapping learning as an Image-Model discriminant analysis framework. The learning objective is formulated by incorporating the intra-class compactness and inter-class separability for good discrimination. Moreover, a weighted learning scheme is introduced to concentrate on the discriminating information of the most confusing samples and then further enhance the performance. Experiments on the COX-S2V dataset demonstrate the effectiveness of the proposed method. Xue Chen 0002, Chunheng Wang, Baihua Xiao |
IJCB | 3 |
| 2014 | Scenario oriented discriminant analysis for still-to-video face recognitionabstractIn the Still-to-Video (S2V) face recognition, each subject is enrolled with only few high resolution images, while the probe is video clips of complex variations. As faces present distinct characteristics under different scenarios, recognition in the original space is obviously inefficient. Therefore, in this paper, we propose a novel discriminant analysis method to learn separate mappings for different scenario patterns (still, video), and further pursue a common discriminant space for the cross-scenario samples. To maximize the intra-individual correlation of samples in the mapping space, we formulate the learning objective by incorporating the intra-class compactness and the inter-class dispersion. The gradient descend algorithm is used to get the optimal solution. Experimental results on the COX-S2V dataset demonstrate the effectiveness of the proposed method and remarkable superiority over state-of-art methods. Xue Chen 0002, Chunheng Wang, Baihua Xiao, Xinyuan Cai |
ICIP | 3 |
| 2014 | Learning associate appearance manifolds for cross-pose face recognitionabstractPose variation is a major challenge in face recognition. In this paper, we propose a novel cross-pose face recognition method by learning associate appearance manifolds to model the connection of faces under different poses. The associate manifolds are built on an auxiliary set, in which each identity contains cross-pose face images. The basic assumption is that cross-pose face images from two similar identities can be projected onto similar appearance manifolds by pose-specific transforms. We first associate the input faces with alike identities from the auxiliary set. Then the manifolds of cross-pose faces in the training set are confined close to that of the associate identities in the auxiliary set. Thus, the connection of cross-pose faces is well modeled by the associate appearance manifolds on the auxiliary set. Formally, we formulate the assumption as a manifold-based distance minimization problem, so as to learn the optimal transforms. Experiments on the Multi-PIE dataset demonstrate the effectiveness of the proposed method. Xue Chen 0002, Chunheng Wang, Baihua Xiao, Xinyuan Cai |
ICIP | 3 |
| 2014 | Learning co-occurrence strokes for scene character recognition based on spatiality embedded dictionaryabstractRobust scene-text-extraction system can be used in lots of areas. In this work, we propose to learn co-occurrence of local strokes for robust character recognition by using a spatiality embedded dictionary (SED). Different from spatial pyramid partitioning images into grids to incorporate spatial information, our SED associates every codeword with a particular response region and introduces more precise spatial information for character recognition. After localized soft coding and max pooling of the first layer, a sparse dictionary is learned to model co-occurrence of several local strokes, which further improves classification performance. Experiment on benchmark datasets demonstrates the effectiveness of our method and the results outperform state-of-the-art algorithms. Song Gao 0009, Chunheng Wang, Baihua Xiao, Cunzhao Shi, Wen Zhou 0002, Zhong Zhang 0001 |
ICIP | 3 |
| 2014 | Stroke Bank: A High-Level Representation for Scene Character RecognitionabstractText information contained in scene images is very useful for image understanding. In this paper, we propose a high-level representation named stroke bank for scene character recognition. Inspired by the work of object bank, we train stroke detectors and use detectors' maximal output as features. Specifically, we collect training samples for stroke detectors based on labeled key points. We also propose to restrict classification areas of each stroke detector to particular local regions, which alleviates computation burden and retains discrimination power at the same time. Experiments on benchmark datasets demonstrate the effectiveness of our method and the results outperform state-of-the-art algorithms. Song Gao 0009, Chunheng Wang, Baihua Xiao, Cunzhao Shi, Zhong Zhang 0001 |
ICPR | 3 |
| 2014 | Human action recognition using weighted poolingabstractPooling strategies, such as max pooling and sum pooling, have been widely used to obtain the global representations for action videos. However, these pooling strategies have several disadvantages. First, they are easily affected by unwanted background local features, the absence of discriminative local features and the times of actions periodically performed by actors. Second, most pooling strategies only use local features to build the global representation that captures little mid‐level features for action representation. In this study, the authors propose a novel weighted pooling strategy based on actionlets representation for action recognition. The actionlets are defined as the movements of large bodies such as legs, arms and head, which capture rich mid‐level features for action representation. Besides, the authors’ method also incorporates the distribution information of actionlets into pooling procedure. Specifically, a pooling weight, which determines the importance of actionlet on the final video representation, is assigned to each actionlet. To learn the weight, they propose a novel discriminative learning algorithm to capture the discriminative information for pooling operation. They evaluate their weighted pooling on three datasets: KTH actions dataset, UCF sports dataset and Youtube actions dataset. Experimental results show the effectiveness of the proposed method. Wen Zhou 0002, Chunheng Wang, Baihua Xiao, Zhong Zhang 0001 |
IET Comput. Vis. | 3 |
| 2014 | End-to-end scene text recognition using tree-structured models
Cunzhao Shi, Chunheng Wang, Baihua Xiao, Song Gao 0009 |
Pattern Recognit. | 3 |
| 2014 | Action recognition via structured codebook construction
Wen Zhou 0002, Chunheng Wang, Baihua Xiao, Zhong Zhang 0001 |
Signal Process. Image Commun. | 3 |
| 2014 | SLD: A Novel Robust Descriptor for Image MatchingabstractImage matching based on local features is a challenging task because it is difficult to build a robust local descriptor which is invariant to large variations in scale, viewpoints, illumination and rotation. To address these issues, Scale Invariant Feature Transform (SIFT) descriptor has been proposed to build a robust and distinctive local descriptor. However, it is not fully affine invariant. In this letter, we propose a novel robust descriptor: Sampling based Local Descriptor (SLD) to perform reliable image matching under large variations in scale, viewpoints, illumination and rotation. We build the descriptor based on elliptical sampling which samples image pixels according to the elliptic equations. The main advantage of elliptical sampling is that two controllable parameters of elliptical sampling can generate descriptors with different viewpoints and rotations. Besides, the descriptor has two notable properties: 1) it is fully invariant to affine changes; 2) it enables fast matching process because we only need to search two controllable parameters for elliptical sampling, which is more efficient than other affine invariant descriptors. We test the proposed descriptor on standard benchmark for evaluation. Experimental results show the robustness of the proposed method under large variations in illumination, viewpoints and scale. Wen Zhou 0002, Chunheng Wang, Baihua Xiao, Zhong Zhang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2014 | Scene Text Recognition Using Structure-Guided Character Detection and Linguistic KnowledgeabstractScene text recognition has inspired great interests from the computer vision community in recent years. In this paper, we propose a novel scene text-recognition method integrating structure-guided character detection and linguistic knowledge. We use part-based tree structure to model each category of characters so as to detect and recognize characters simultaneously. Since the character models make use of both the local appearance and global structure informations, the detection results are more reliable. For word recognition, we combine the detection scores and language model into the posterior probability of character sequence from the Bayesian decision view. The final word-recognition result is obtained by maximizing the character sequence posterior probability via Viterbi algorithm. Experimental results on a range of challenging public data sets (ICDAR 2003, ICDAR 2011, SVT) demonstrate that the proposed method achieves state-of-the-art performance both for character detection and word recognition. Cunzhao Shi, Chunheng Wang, Baihua Xiao, Song Gao 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | Cross-View Action Recognition Using Contextual Maximum Margin ClusteringabstractRecently, maximum margin clustering (MMC) has been proposed for a cross-view action recognition. However, such a method neglects the temporal relationship between contiguous frames in the same action video. In this paper we propose a novel method called contextual maximum margin clustering (CMMC) to tackle cross-view action recognition. In CMMC, we add temporal regularization to give a high penalty when the contiguous frames are dissimilar. Thus, the CMMC not only achieves the goal of finding maximum margin hyperplanes, but also explicitly considers the temporal information among contiguous frames. Our method is verified on the IXMAS dataset and the experimental results demonstrate that our method can achieve better performance than the state-of-the-art methods. Zhong Zhang 0001, Chunheng Wang, Baihua Xiao, Wen Zhou 0002, Shuang Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2013 | Scene Text Recognition Using Part-Based Tree-Structured Character DetectionabstractScene text recognition has inspired great interests from the computer vision community in recent years. In this paper, we propose a novel scene text recognition method using part-based tree-structured character detection. Different from conventional multi-scale sliding window character detection strategy, which does not make use of the character-specific structure information, we use part-based tree-structure to model each type of character so as to detect and recognize the characters at the same time. While for word recognition, we build a Conditional Random Field model on the potential character locations to incorporate the detection scores, spatial constraints and linguistic knowledge into one framework. The final word recognition result is obtained by minimizing the cost function defined on the random field. Experimental results on a range of challenging public datasets (ICDAR 2003, ICDAR 2011, SVT) demonstrate that the proposed method outperforms state-of-the-art methods significantly both for character detection and word recognition. Cunzhao Shi, Chunheng Wang, Baihua Xiao, Song Gao 0009, Zhong Zhang 0001 |
CVPR | 3 |
| 2013 | Cross-View Action Recognition via a Continuous Virtual PathabstractIn this paper, we propose a novel method for cross-view action recognition via a continuous virtual path which connects the source view and the target view. Each point on this virtual path is a virtual view which is obtained by a linear transformation of the action descriptor. All the virtual views are concatenated into an infinite-dimensional feature to characterize continuous changes from the source to the target view. However, these infinite-dimensional features cannot be used directly. Thus, we propose a virtual view kernel to compute the value of similarity between two infinite-dimensional features, which can be readily used to construct any kernelized classifiers. In addition, there are a lot of unlabeled samples from the target view, which can be utilized to improve the performance of classifiers. Thus, we present a constraint strategy to explore the information contained in the unlabeled samples. The rationality behind the constraint is that any action video belongs to only one class. Our method is verified on the IXMAS dataset, and the experimental results demonstrate that our method achieves better performance than the state-of-the-art methods. Zhong Zhang 0001, Chunheng Wang, Baihua Xiao, Wen Zhou 0002, Shuang Liu 0001, Cunzhao Shi |
CVPR | 3 |
| 2013 | Adaptive Scene Text Detection Based on Transferring AdaboostabstractDetecting text in scene images is very challenging due to complex backgrounds, various fonts and different illumination conditions. Without prior knowledge, a detector previously trained using lots of samples still perform badly on a test image because of the disparities in distributions between the training samples and the testing ones. In this paper, we propose to adapt a pre-trained generic scene text detector towards new scenes by transfer learning. In particular, we choose cascade Adaboost as the detector style and try to re-weight pre-selected features according to their abilities to classify high confidence samples. The proposed adaptation mechanism has been evaluated on ICDAR 2011 scene text detection competition dataset and the encouraging experiments results can be compared with the latest published algorithms. Song Gao 0009, Chunheng Wang, Baihua Xiao, Cunzhao Shi, Zhijian Lv, Yanqin Shi |
ICDAR | 3 |
| 2013 | Coupled latent least squares regression for heterogeneous face recognitionabstractOne of the most difficult challenges in automatic face recognition is computing facial similarity between two images captured in different modalities, called heterogeneous face recognition. In this paper, we propose a novel method, named as coupled latent least squares regression, to improve the heterogeneous face recognition performance. The basic assumption is that the images of one person captured in different modalities can be viewed as modality-specific transforms of a latent ideal object. We formulate this assumption in the least squares regression framework, so as to learn the coupled transforms for different modalities. In particular, the local consistency information in the each modality is considered as a constraint to improve the generalization. Extensive experiments on two cases of heterogeneous face recognition (visible light vs. near infrared, and photo vs. sketch) validate the efficiency of the proposed method. Xinyuan Cai, Chunheng Wang, Baihua Xiao, Xue Chen 0002, Zhijian Lv, Yanqin Shi |
ICIP | 3 |
| 2013 | Modular hierarchical feature learning with deep neural networks for face verificationabstractFeature representations play a crucial role in modern face recognition systems. Most hand-crafted image descriptors usually provide low-level information. In this paper, we propose a novel feature learning method based on deep neural networks to obtain high-level, hierarchical representations for face verification. Learning proceeds in two phases. In the pre-training phase, we train Restricted Boltzmann Machine(RBM) networks for each modular region in the image separately. In the fine-tuning phase, in order to develop good discriminative ability, we stack the RBM networks of each region in deep architecture and combine deep learning with side information constraints in the whole image scale. Finally, we formulate the proposed method as an appropriate optimization problem and adopt gradient descent algorithm to get the optimal solution. We evaluate our method on the LFW dataset. Representations learned from the networks achieve comparable performance (93.11%) to the state-of-art method. Xue Chen 0002, Baihua Xiao, Chunheng Wang, Xinyuan Cai, Zhijian Lv, Yanqin Shi |
ICIP | 2 |
| 2013 | Regularized Latent Least Square Regression for Cross Pose Face Recognition
Xinyuan Cai, Chunheng Wang, Baihua Xiao, Xue Chen 0002 |
IJCAI | 3 |
| 2013 | Visual word density-based nonlinear shape normalization method for handwritten Chinese character recognition
Yunxue Shao, Chunheng Wang, Baihua Xiao |
Int. J. Document Anal. Recognit. | 3 |
| 2013 | Fast self-generation voting for handwritten Chinese character recognition
Yunxue Shao, Chunheng Wang, Baihua Xiao |
Int. J. Document Anal. Recognit. | 3 |
| 2013 | Tensor Ensemble of Ground-Based Cloud Sequences: Its Modeling, Classification, and SynthesisabstractSince clouds are one of the most important meteorological phenomena related to the hydrological cycle and affect Earth radiation balance and climate changes, cloud analysis is a crucial issue in meteorological research. Most researchers only consider the classification task of cloud images while less attention has been paid to the synthesis one. In addition, all the existing research on cloud identification from sky images is based on single cloud images. However, the cloud-measuring devices on the ground actually take one image of the clouds every few minutes and collect a series of cloud images. Thus, the existing methods neglect the temporal information exhibited by contiguous cloud images. To overcome this drawback, in this letter we treat ground-based cloud sequences (GCSs) as dynamic texture. We then propose the Tensor Ensemble of Ground-based Cloud Sequences (eTGCS) model which represents the ensemble of GCSs in a tensor manner. In the eTGCS model, all GCSs form a single tensor, and each GCS is a subtensor of the single tensor. There are two main characteristics of the eTGCS model: 1) All GCSs share an identical mode subspace, which makes the classification convenient, and 2) a new GCS can be synthesized as long as the parameters of the eTGCS model are used. Therefore, less storage space is required. Comprehensive experiments are conducted to prove the superiority of our eTGCS model. The classification accuracy achieves 92.31%, and the synthesized GCSs are similar to the original ones in visual appearance. Shuang Liu 0001, Chunheng Wang, Baihua Xiao, Zhong Zhang 0001, Xiaozhong Cao |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2013 | Scene text detection using graph model built upon maximally stable extremal regions
Cunzhao Shi, Chunheng Wang, Baihua Xiao, Song Gao 0009 |
Pattern Recognit. Lett. | 3 |
| 2013 | Attribute Regularization Based Human Action RecognitionabstractRecently, attributes have been introduced as a kind of high-level semantic information to help improve the classification accuracy. Multitask learning is an effective methodology to achieve this goal, which shares low-level features between attributes and actions. Yet such methods neglect the constraints that attributes impose on classes, which may fail to constrain the semantic relationship between the attributes and actions. In this paper, we explicitly consider such attribute-action relationship for human action recognition, and correspondingly, we modify the multitask learning model by adding attribute regularization. In this way, the learned model not only shares the low-level features, but also gets regularized according to the semantic constrains. In addition, since attribute and class label contain different amounts of semantic information, we separately treat attribute classifiers and action classifiers in the framework of multitask learning for further performance improvement. Our method is verified on three challenging datasets (KTH, UIUC, and Olympic Sports), and the experimental results demonstrate that our method achieves better results than that of previous methods on human action recognition. Zhong Zhang 0001, Chunheng Wang, Baihua Xiao, Wen Zhou 0002, Shuang Liu 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2012 | Multi-scale Fusion of Texture and Color for Background ModelingabstractBackground modeling from a stationary camera is a crucial component in video surveillance. Traditional methods usually adopt single feature type to solve the problem, while the performance is usually unsatisfactory when handling complex scenes. In this paper, we propose a multi-scale strategy, which combines both texture and color features, to achieve a robust and accurate solution. Our contributions are two folds: one is that we propose a novel textureoperator named Scale-invariant Center-symmetric Local Ternary Pattern, which is robust to noise and illumination variations, the other is that a multi-scale fusion strategy is proposed for the issue. Our method is verified on several complex real world videoswith illumination variation, soft shadows and dynamic backgrounds. We compare our method with four state-of-the-art methods, and the experimental results clearly demonstrate that our method achievesthe highest classification accuracy in complex real world videos. Zhong Zhang 0001, Chunheng Wang, Baihua Xiao, Shuang Liu 0001, Wen Zhou 0002 |
AVSS | 3 |
| 2012 | Human Action Recognition with Attribute RegularizationabstractRecently, attributes have been introduced to help object classification. Multi-task learning is an effective methodology to achieve this goal, which shares low-level features between attribute and object classifiers. Yet such a method neglects the constraints that attributes impose on classes which may fail to constrain the semantic relationship between the attribute and object classifiers. In this paper, we explicitly consider such attribute-object relationship, and correspondingly, we modify the multi-task learningmodel by adding attribute regularization. In this way, the learned model not only shares the low-level features, but also gets regularized according to the semantic constrains. Our method is verified on two challenging datasets (KTH and Olympic Sports), andthe experimental results demonstrate that our method achieves better results than previous methods in human action recognition. Zhong Zhang 0001, Chunheng Wang, Baihua Xiao, Wen Zhou 0002, Shuang Liu 0001 |
AVSS | 3 |
| 2012 | Sparse representation for face recognition based on discriminative low-rank dictionary learningabstractIn this paper, we propose a discriminative low-rank dictionary learning algorithm for sparse representation. Sparse representation seeks the sparsest coefficients to represent the test signal as linear combination of the bases in an over-complete dictionary. Motivated by low-rank matrix recovery and completion, assume that the data from the same pattern are linearly correlated, if we stack these data points as column vectors of a dictionary, then the dictionary should be approximately low-rank. An objective function with sparse coefficients, class discrimination and rank minimization is proposed and optimized during dictionary learning. We have applied the algorithm for face recognition. Numerous experiments with improved performances over previous dictionary learning methods validate the effectiveness of the proposed algorithm. Chunheng Wang, Baihua Xiao, Wen Zhou 0002 |
CVPR | 3 |
| 2012 | Adaptive Graph Cut Based Binarization of Video Text ImagesabstractInteractive image segmentation which needs the user to give certain hard constraints has shown promising performance for object segmentation. In this paper, we consider characters in text image as a special kind of object, and propose an adaptive graph cut based text binarization method to segment text from background. The main contributions of the paper lie in: 1) in order to make the binarization local adaptive with uneven background, the text region image is firstly roughly split into several sub-images on which graph cut is applied, and 2) considering the unique characteristics of the text, we propose to automatically classify some pixels as text or background with high confidence, severed as hard constraints seeds for graph cut to extract text from background by spreading the seeds into the whole sub-image. The experimental results show that our approach could get better performance in both character extraction accuracy and recognition accuracy. Cunzhao Shi, Baihua Xiao, Chunheng Wang |
Document Analysis Systems | 2 |
| 2012 | Graph-Based Background Suppression for Scene Text DetectionabstractDetecting text in video or natural scene image is quite challenging due to the complex background, various fonts and illumination conditions. The preprocessing period, which suppresses the nontext areas so as to highlight the text areas, is the basis for further text detection. In this paper, a novel graph-based background suppression method for scene text detection is proposed. Considering each pixel as a node in the graph, our approach incorporates pixel-level and context-level features into a graph. Various factors contribute to the unary and pair wise cost function which is optimized via max-flow/min-cut algorithm [16] to get a binary image whose nontext pixels are suppressed so that text pixels are highlighted. Furthermore, the proposed background suppression method could be easily combined with other detection methods to improve the performance. Experimental results on ICDAR 2011 competition dataset show promising performance. Cunzhao Shi, Baihua Xiao, Chunheng Wang |
Document Analysis Systems | 2 |
| 2012 | A New Method for Text Verification Based on Random ForestsabstractText in image or video frames contains a lot of high-level semantics which can be useful for multimedia indexing, management. Coarse text detection results may contain many false alarms, which makes it necessary to eliminate the false alarms for further recognition. As text has distinct textural features, texture-based classifier such as SVM, MLP and Adaboost has been used to classify the detection regions as text or non-text region. In this paper, a random forests based method for text verification is proposed. The reason of choosing random forests lies in: 1) its ability of maintaining accuracy in small labeled dataset and 2) its good performance in unbalanced dataset as in the case of unbalanced text and non-text distribution. Furthermore, we propose to merge different random forests trained with different kinds of features to improve the accuracy of classification. The comprehensive experimental results show that our methods are effective. Chunheng Wang, Baihua Xiao, Cunzhao Shi |
ICFHR | 3 |
| 2012 | A New Text Extraction Method Incorporating Local InformationabstractText detection and extraction in images with complex background can provide useful information for video annotation and indexing. More attention is paid to text detection for its importance, but text extraction is necessary for the text recognition, and it can test the validity of text detection. In this paper, we conclude text extraction is to segment the image and to remove noises, and then a robust text extraction method incorporating local information is proposed. First, we get the gray image from the original image and reprocess the gray image with edge enhancement. Then a binarization method incorporating local information is used to segment the gray image, by which the text-noises are removed and a binary image is obtained. Finally, the connected component analysis based on the character's density and geometric feature is performed on the binary image, by which background-noises are removed. The preliminary experiments show some promising results. Chunheng Wang, Baihua Xiao, Cunzhao Shi |
ICFHR | 3 |
| 2012 | Environment coupled metrics learning for unconstrained face verificationabstractMaking recognition more reliable under unconstrained environment is one of the most important challenges for realworld face recognition. In this paper, we propose a novel approach for unconstrained face verification. First, we use a spectral-clustering method based on Structural Similarity index to estimate the captured environments of facial images. Then for each pair of environments, we learn two coupled metrics, such that facial images captured in different environments can be transformed into a media subspace, and high recognition performance can be achieved. The coupled transformations are jointly determined by solving an optimization problem in the multi-task learning framework. Experimental results on the benchmark dataset (LFW) show the effectiveness of the proposed method in face verification across varying environments. Xinyuan Cai, Chunheng Wang, Baihua Xiao, Xue Chen 0002 |
ICIP | 3 |
| 2012 | Soft-signed sparse coding for ground-based cloud classification
Shuang Liu 0001, Chunheng Wang, Baihua Xiao, Zhong Zhang 0001, Yunxue Shao |
ICPR | 3 |
| 2012 | Contextual Fisher kernels for human action recognition
Zhong Zhang 0001, Chunheng Wang, Baihua Xiao, Wen Zhou 0002, Shuang Liu 0001 |
ICPR | 3 |
| 2012 | Learning weighted features for human action recognition
Wen Zhou 0002, Chunheng Wang, Baihua Xiao, Zhong Zhang 0001 |
ICPR | 3 |
| 2012 | Human action recognition by bagging data dependent representation
Wen Zhou 0002, Chunheng Wang, Baihua Xiao, Zhong Zhang 0001 |
ICPR | 3 |
| 2012 | Sparse representation based on matrix rank minimization and k-means clustering for recognitionabstractIn this paper, we propose a sparse coding algorithm based on matrix rank minimization and k-means clustering and for recognition. We consider the problem of removing the noise in the training samples and generating more samples at the same time. To accomplish this, we extended the matrix rank minimization problem to cope with complex data. Samples from the same class are segmented into several groups by k-means clustering algorithm, and matrix rank minimization is applied on the clustered data to separate the noises and recover the low-rank structures in the grouped data. An over-complete dictionary is constructed by connecting the low-rank structures and the training samples together to keep the samples diversity. Sparse representation is operated based on this over-complete dictionary for recognition. Furthermore, a parameter is introduced to adjust the weighting of the coefficients that code the noises. We apply the proposed algorithm for character and face recognition. Experiments with improved performances validate the effectiveness of the proposed algorithm. Chunheng Wang, Baihua Xiao |
IJCNN | 3 |
| 2012 | Deep nonlinear metric learning with independent subspace analysis for face verificationabstractFace verification is the task of determining by analyzing face images, whether a person is who he/she claims to be. It is a very challenge problem, due to large variations in lighting, background, expression, hairstyle and occlusion. The crucial problem is to compute the similarity of two face vectors. Metric learning has provides a viable solution to this problem. Until now, many metric learning algorithms have been proposed, but they are usually limited to learning a linear transformation (i.e. finding a global Mahalanobis metric). In this brief, we propose a nonlinear metric learning method, which learns an explicit mapping from the original space to an optimal subspace, using deep Independent Subspace Analysis network. Compared to kernel methods, which can also learn nonlinear transformations, our method is a deep and local learning architecture, and therefore exhibits more powerful ability to learn the nature of highly variable dataset. We evaluate our method on the LFW benchmark, and results show very comparable performance to the state-of-art methods (achieving 92.28% accuracy), while maintaining simplicity and good generalization ability. Xinyuan Cai, Chunheng Wang, Baihua Xiao, Xue Chen 0002 |
ACM Multimedia | 3 |
| 2012 | Action Recognition Using Context-Constrained Linear CodingabstractAlthough traditional bag-of-words model has shown promising results for action recognition, it takes no consideration of the relationship among spatio–temporal points; furthermore, it also suffers serious quantization error. In this letter, we propose a novel coding strategy called context-constrained linear coding (CLC) to overcome these limitations. We first calculate the contextual distance between local descriptors and each codeword by considering the spatio–temporal contextual information. Then, linear coding using contextual distance is adopted to alleviate the quantization error. Our method is verified on two challenging databases (KTH and UCF sports), and the experimental results demonstrate that our method achieves better results than previous methods in action recognition. Zhong Zhang 0001, Chunheng Wang, Baihua Xiao, Wen Zhou 0002, Shuang Liu 0001 |
IEEE Signal Process. Lett. | 3 |
| 2011 | Modified Two-Class LDA Based Compound Distance for Similar Handwritten Chinese Characters DiscriminationabstractThis paper proposes a modified two-class LDA based compound distance for similar handwritten Chinese characters discrimination. First the definition of the Intersecting Subspace (IS) between two classes and the modified between-class scatter matrix is given. Then we prove that the modified between-class scatter matrix can supply additional information. Our experiments demonstrate that the additional information can be used to discriminate points in the IS and the proposed method outperforms the previous LDA based method. Yunxue Shao, Chunheng Wang, Baihua Xiao, Rongguo Zhang |
ICDAR | 3 |
| 2011 | Multiple Instance Learning Based Method for Similar Handwritten Chinese Characters DiscriminationabstractThis paper proposes a Multiple Instance Learning based method for similar handwritten Chinese characters discrimination. The similar handwritten Chinese characters recognition problem is first defined as a Multiple-instance learning problem. Then the problem is solved by the AdaBoost framework. The proposed method selects some self-adapting critical regions as weak classifiers, and therefore it is more suitable for the wide variability of writing styles. Our experimental results demonstrate that the proposed method outperforms the other state-of-the-art methods. Yunxue Shao, Chunheng Wang, Baihua Xiao, Rongguo Zhang |
ICDAR | 3 |
| 2010 | Globally-Preserving Based Locally Linear EmbeddingabstractThe locally linear embedding (LLE) algorithm is considered as a powerful method for the problem of nonlinear dimensionality reduction. In this paper, a new method called globally-preserving based LLE (GPLLE) is proposed. It not only preserves the local neighborhood, but also keeps those distant samples still far away, which solves the problem that LLE may encounter, i.e. LLE only makes local neighborhood preserving, but can't prevent the distant samples from nearing. Moreover, GPLLE can estimate the intrinsic dimensionality d of the manifold structure. The experiment results show that GPLLE always achieves better classification performances than LLE based on the estimated d. Kanghua Hui, Chunheng Wang, Baihua Xiao |
ICPR | 3 |
| 2010 | A New Biologically Inspired Feature for Scene Image ClassificationabstractScene classification is a hot topic in pattern recognition and computer vision area. In this paper, based on the past research on vision neuroscience, we proposed a new biologically inspired feature method for scene image classification. The new feature accounts for the visual processing from simple cell to complex cell in V1 area, and also the spatial layout for scene gist signature. It provides a different line and model revision to consider some nonlinearities inV1 area. We compare it with traditional HMAX model and recently proposed ScSPM model, and experiment on a popular 15 scenes dataset. We show that our proposed method has many important differences and merits. The experiment results also show that our method outperforms the state-of-the-art like ScSPM and KSPM model. Aiwen Jiang, Chunheng Wang, Baihua Xiao, Ruwei Dai |
ICPR | 3 |
| 2010 | Data Transformation of the Histogram Feature in Object DetectionabstractDetecting objects in images is very important for several application domains in computer vision. This paper presents an experimental study on data transformation of the feature vector in object detection. We use the modified Pyramid of Histograms of Orientation Gradients descriptor and the SVM classifier to form an object detection model. We apply a simple transformation to the histogram features before training and testing. This transformation equals a small change in the kernel function for Support Vector Machines. This change is much quicker than the χ2kernel, but obtains better results. Experimental evaluations on the UIUC Image Database and TU Darmstadt Database show that the transformed features perform better than the raw features, and this transformation improves the linear separability of the histogram feature. Rongguo Zhang, Baihua Xiao, Chunheng Wang |
ICPR | 2 |
| 2010 | Conditional random field for text segmentation from images with complex background
Minhua Li, Meng Bai, Chunheng Wang, Baihua Xiao |
Pattern Recognit. Lett. | 4 |
| 2008 | A no reference image quality assessment method for JPEG2000abstractThis paper presents a novel no reference method to assess image quality. Firstly, the image is divided into many blocks. Textured blocks are selected and their amplitude fall-off curves are employed for quality prediction based on natural scene statistics. Secondly, projections of wavelet coefficients between adjacent scales with the same orientation are utilized to measure the positional similarity. At last, general regression neural network is adopted to conduct quality prediction according to features from above two aspects. The performance of our method is evaluated on a public data set and experimental results confirm its effectiveness. Jingchao Zhou, Baihua Xiao, Qiudan Li |
IJCNN | 2 |
| 2007 | Usage-Oriented Performance Evaluation for Text Localization AlgorithmsabstractThe localization of texts in image/video is the first step in a text processing system. Its effect will do great impact on the following processing steps. Although many studies have been done on text localization algorithms, there is not a universally accepted performance evaluation method. In this paper we propose two sets of metrics to evaluate the performance of text localization algorithms in different usage conditions. The metrics also consider the text distribution characteristics, and the difficulties of the underlying task. Some experiments on the proposed metrics are also given. Yichao Ma, Chunheng Wang, Baihua Xiao, Ruwei Dai |
ICDAR | 3 |
| 2007 | Integrated Segmentation and Recognition of Mixed Chinese/English DocumentabstractThis paper presents a general frame to integrate segmentation and recognition and gives a novel method to identify lingual attribute of mixed Chinese/English characters. The outstanding performance of this method is as follows. First, a text- line rather than a character segment is regarded as a process unit. Second, multi-feature is adopted based on multi-phase segmentation. Third, two types of feedbacks, including from character recognition and from character feature statistic within a text-line, are adopted throughout the whole segmentation and recognition. Fourth, it is adaptive to the quality and genre of documents. Baihua Xiao, Chunheng Wang, Ruwei Dai |
ICDAR | 2 |
| 2007 | Chinese character recognition: history, status and prospects
Ruwei Dai, Baihua Xiao |
Frontiers Comput. Sci. China | 3 |
| 2006 | A Novel Multistage Classification Strategy for Handwriting Chinese Character Recognition Using Local Linear Discriminant Analysis
Baihua Xiao, Chunheng Wang, Ruwei Dai |
ICONIP (2) | 2 |
| 2006 | CWME: A Framework of Group Support System for Emergency Responses
Yaodong Li, Huiguang He, Baihua Xiao, Chunheng Wang, Fei-Yue Wang 0001 |
ISI | 3 |
| 2004 | Parallel compact integration in handwritten Chinese character recognition
Chunheng Wang, Baihua Xiao, Ruwei Dai |
Sci. China Ser. F Inf. Sci. | 2 |
| 2000 | A New Integration Scheme with Multi-Layer Perceptron Networks for Handwritten Chinese Character RecognitionabstractIn this paper, a new integration scheme with multilayer perceptron (MLP) networks is proposed to solve handwritten Chinese character recognition problem. The idea of meta-synthesis is emphasized in this scheme, human intelligence and computer capabilities are combined together through a procedure of two-step supervised learning. Compared with previous integration schemes, this scheme has much better performance and provides a promising way of applying MLP to large vocabulary classification. Chunheng Wang, Baihua Xiao, Ruwei Dai |
ICPR | 2 |
| 2000 | Adaptive Combination of Classifiers and its Application to Handwritten Chinese Character RecognitionabstractMotivated by the idea of metasynthesis, a new adaptive classifier combination approach is proposed in this paper. Compared with previous integration methods, parameters of the proposed combination approach are dynamically acquired by a coefficient predictor based on neural network and vary, with the input pattern. It is also shown that many existing integration schemes can be considered as special cases of the proposed method. This approach is tested in application on handwritten Chinese character recognition. The experimental results demonstrate that this method can result in substantial improvement in overall performance. Baihua Xiao, Chunheng Wang, Ruwei Dai |
ICPR | 1 |