VLDB 2026 Research / reviewers in the wild / expert
Hongxi Wei
dblp:05/10400
· DBLP profile ↗
43ranked-venue papers
16as first author
24since 2021 · last 2026
0000-0002-2570-4544ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 13 first-author · 14 since 2021Databases, data management, data science and information retrieval · 16 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A MambaVision-Based Cross-Modal Feature Enhancement Network for Scene Text Super-Resolution
Ruichang Zhu, Hongxi Wei |
ICDAR (2) | 2 |
| 2025 | GL-GAN: Perceiving and Integrating Global and Local Styles for Handwritten Text Generation with MambaabstractHandwritten text generation (HTG) aims to synthesize handwritten samples by imitating a specific writer, which has a wide range of applications and thus has significant research value. However, current studies on HTG are confronted with a main bottleneck: dominant models lack the ability to perceive and integrate handwriting styles, which affects the realism of the synthesized samples. In this paper, we propose GL-GAN, which effectively captures and integrates global and local styles. Specifically, we propose a Hybrid Style Encoder (HSE) that combines a state space model (SSM) and convolution to capture multilevel style features through various receptive fields. The captured style features are then fed to the proposed Dynamic Feature Enhancement Module (DFEM), which integrates these features by adaptively modeling the entangled relationships between multilevel styles and removing redundant details. Extensive experiments on two widely used handwriting datasets demonstrate that our GL-GAN is an effective HTG model and outperforms state-of-the-art models remarkably. Our code is publicly available at:https://github.com/Fyzjym/GL-GAN. Hongxi Wei, Shiwen Sun |
COLING | 2 |
| 2025 | MRANet: An Encoder-Decoder Network with Multi-Scale Residual Atrous-Spatial Pyramid Pooling for Seismic Phase PickingabstractSeismic phase picking is one of the critical challenges in seismic data processing. With the advancement of deep learning, numerous neural network architectures have been employed to explore the correlations between seismic waveforms and the underlying information. However, existing methods predominantly rely on convolutional neural network (CNN)-based feature extraction techniques, which often overlook the denoising of input data. Additionally, CNN-based feature extraction methods face limitations in seismic phase picking tasks, particularly in capturing long-range dependencies. Inspired by the Atrous-Spatial Pyramid Pooling (ASPP) framework, this paper proposes an attention mechanism-based multi-scale residual Atrous-Spatial Pyramid Pooling module (ResASPP), which effectively captures detailed information across different scales, enabling efficient extraction of seismic signal features. Based on this module, a multi-scale residual ASPP neural network (MRANet) with multi-scale feature extraction capabilities is constructed for seismic phase picking tasks. Moreover, to improve the quality of seismic data and enhance interpretability, denoising techniques are applied to the input data. Extensive experiments conducted on the DiTing dataset demonstrate that our model achieves higher F1 scores (89.45% for P-waves and 80.40% for S-waves), outperforming existing state-of-the-art models. Hongxi Wei, Yingyue Jing, Lingguo Meng |
ICASSP | 2 |
| 2025 | SmartExp: An Adaptive Data Expansion Strategy for Improving Handwritten Text RecognitionabstractConstructing a highly accurate handwritten OCR system requires large amounts of high-quality training data, yet data collection is labor-intensive and costly. With the advance of generative models, high-quality synthetic images have been applied to enhance handwritten text recognition (HTR) models, a technique termed data expansion. However, we find that the generated data may sometimes provide only insignificant improvements to HTR. We investigate the causes behind the small gains from the perspectives of both data expansion and data augmentation. We reveal a balancing mechanism in HTR that stronger data expansion should be paired with weaker data augmentation. Informed by these insights, we propose SmartExp, a purely data-driven strategy without introducing any extra data or computational costs. Extensive experiments on three widely used benchmark datasets demonstrate that our SmartExp strategy can significantly improve various HTR models. Our code is publicly available at: https://github.com/Fyzjym/SmartExp. Hongxi Wei, Shiwen Sun |
ICASSP | 2 |
| 2025 | AS-Net: Adaptive Style-aware Network for Handwritten Text GenerationabstractHandwritten text generation (HTG) is a challenging task due to the vast diversity of handwriting styles. In addition, even the same writer can show subtle differences when writing the same character, further aggravating the difficulties of HTG. In this paper, we propose a novel Adaptive Style-aware Network (AS-Net) to address the challenging HTG task. Specifically, we propose a Style-induced Context-aware Generator (SCG) to capture handwriting styles and generate new images. The SCG learns a lightweight neural network to generate a style token (vector) for each image. Then, the context tokens with text information are element-wise multiplied by the style tokens to construct mixed tokens. Finally, we employ multiple attention mechanisms to capture dependencies among style features and entanglement with style and text features. Extensive experiments on two widely used benchmark datasets demonstrate that our AS-Net is an effective HTG model and outperforms state-of-the-art methods markedly. Our code is publicly available at: https://github.com/Fyzjym/AS-Net. Hongxi Wei |
ICASSP | 2 |
| 2025 | When CLIP Meets PHOC: A Dual-Branch Network for Historical Document Image RetrievalabstractIn this paper, we leverage Contrastive Language-Image Pre-training (CLIP) for Historical Document Image Retrieval (HDIR). We are largely inspired by recent advances on CLIP and its exceptional generalization capabilities, but for the first time, we tailor it to benefit HDIR. We put forward a dual-branch joint learning network based on CLIP and Pyramid Histograms of Characters (PHOC), termed CPNet, which supports out-of-vocabulary word spotting for both query-by-example and query-by-string. To achieve this, we introduce two key components: (i) a rule-based PHOC to enhance text representations by capturing detailed character-level information, and (ii) a visual prompt that allows the CLIP image encoder to adapt and extend its knowledge to our task through additional learnable prompts. Extensive experiments on benchmark datasets, including Kanjur and Geser, demonstrate that the proposed CPNet sets a new state-of-the-art in performance for HDIR, highlighting its potential for advancing research in other text-related retrieval tasks. Hongxi Wei |
ICASSP | 2 |
| 2025 | DCC: Plug-and-Play Dynamic Category Compression for Enhanced Handwritten Text Generation
Hongxi Wei, Shiwen Sun |
ICDAR (1) | 2 |
| 2025 | VMF-Net: Visual-Aware Multi-representation Fusion Network for Artifact-Free Handwritten Mathematical Expressions Generation
Hongxi Wei |
ICDAR (3) | 2 |
| 2025 | SFRD: Handwritten Mathematical Expressions Generation by Spatial-Aware Feature Refinement Diffusion
Hongxi Wei, Shiwen Sun |
ICDAR (4) | 2 |
| 2025 | MambaHash: Visual State Space Deep Hashing Model for Large-Scale Image RetrievalabstractDeep image hashing aims to enable effective large-scale image retrieval by mapping the input images into simple binary hash codes through deep neural networks. More recently, Vision Mamba with linear time complexity has attracted extensive attention from researchers by achieving outstanding performance on various computer tasks. Nevertheless, the suitability of Mamba for large-scale image retrieval tasks still needs to be explored. Towards this end, we propose a visual state space hashing model, called MambaHash. Concretely, we propose a backbone network with stage-wise architecture, in which grouped Mamba operation is introduced to model local and global information by utilizing Mamba to perform multi-directional scanning along different groups of the channel. Subsequently, the proposed channel interaction attention module is used to enhance information communication across channels. Finally, we meticulously design an adaptive feature enhancement module to increase feature diversity and enhance the visual representation capability of the model. We have conducted comprehensive experiments on three widely used datasets: CIFAR-10, NUS-WIDE and IMAGENET. The experimental results demonstrate that compared with the state-of-the-art deep hashing methods, our proposed MambaHash has well efficiency and superior performance to effectively accomplish large-scale image retrieval tasks. Source code is available https://github.com/shuaichaochao/MambaHash.git Hongxi Wei |
ICMR | 2 |
| 2024 | HENet: Hyperbolic-Based Encoder-Decoder Network for Word Spotting in Historical Mongolian DocumentsabstractIn the domain of historical Mongolian document image retrieval (HMDIR), word spotting poses a inherent challenge due to the frequent appearance of out-of-vocabulary (OOV) words. Existing methods have mainly focused on query-by-example (QBE), neglecting the query-by-string (QBS) approach. Meanwhile, the hierarchical structure of word makes Euclidean space not the optimal choice for representing complex structured data. To address the aforementioned problems, we propose a novel method that leverages a shared hyperbolic space to effectively align text strings and word images. Specifically, we use the Pyramidal Histogram of Characters (PHOC) for text string embeding, and a robust encoder-decoder architecture for word image embedding, then map their embeddings in the shared hyperbolic space. Moreover, we propose a new dataset of historical Mongolian documents called Geser, which includes 143,508 word images and 10,951 vocabularies. Extensive experiments conducted on two datasets of historical Mongolian documents with an OOV partitioning scheme (Kanjur and Geser), demonstrate that our proposed method surpasses state-of-the-art methods and achieves outstanding results on Geser. Hongxi Wei, Xiandong Chen |
ICASSP | 2 |
| 2024 | LABT: A Sequence-to-Sequence Model for Mongolian Handwritten Text Recognition with Local Aggregation BiLSTM and Transformer
Hongxi Wei, Shiwen Sun |
ICDAR (2) | 2 |
| 2024 | Recognition and Link Prediction of Onomatopoeia Texts with Arbitrary Shapes
Hongxi Wei |
ICDAR (3) | 2 |
| 2024 | Deepfake In-Air Signature Verification via Two-Channel Model
Hongxi Wei |
ICDAR (2) | 2 |
| 2024 | HybridHash: Hybrid Convolutional and Self-Attention Deep Hashing for Image RetrievalabstractDeep image hashing aims to map input images into simple binary hash codes via deep neural networks and thus enable effective large-scale image retrieval. Recently, hybrid networks that combine convolution and Transformer have achieved superior performance on various computer tasks and have attracted extensive attention from researchers. Nevertheless, the potential benefits of such hybrid networks in image retrieval still need to be verified. To this end, we propose a hybrid convolutional and self-attention deep hashing method known as HybridHash. Specifically, we propose a backbone network with stage-wise architecture in which the block aggregation function is introduced to achieve the effect of local self-attention and reduce the computational complexity. The interaction module has been elaborately designed to promote the communication of information between image blocks and to enhance the visual representations. We have conducted comprehensive experiments on three widely used datasets: CIFAR-10, NUS-WIDE and IMAGENET. The experimental results demonstrate that the method proposed in this paper has superior performance with respect to state-of-the-art deep hashing methods. Source code is available https://github.com/shuaichaochao/HybridHash. Hongxi Wei |
ICMR | 2 |
| 2024 | A Multi-feature Fusion Approach for Words Recognition of Ancient Mongolian Documents
Shiwen Sun, Hongxi Wei |
PRCV (7) | 2 |
| 2023 | Transformer-Based Deep Hashing Method for Multi-Scale Feature FusionabstractThe deep image hashing aims to map the input image into simply binary hash codes via deep neural networks. Motivated by the recent advancements of Vision Transformers (ViT), many deep hashing methods based on ViT have been proposed. Nevertheless, the ViT has enormous number of model parameters and high computational complexity. Moreover, the last layer of the ViT outputs only the classification tokens as image feature vectors, while the rest of the vectors are discarded. This results in the inefficiency of the model computation and the neglect of useful image information. Therefore, this paper proposes a Transformer-based deep hashing method for multi-scale feature fusion (TDH). Specifically, we use a hierarchical Transformer backbone to capture both global and local features of images. The hierarchical Transformer utilizes a local self-attention mechanism to process image blocks in parallel, which reduces computational complexity and promotes computational efficiency. Multi-scale feature fusion module captures all image feature vectors of the hierarchical Transformer output to obtain more enriched image feature information. We perform comprehensive experiments on three widely-studied datasets: CIFAR-10, NUS-WIDE and IMAGENET. The experimental results demonstrate that the proposed method in this paper indicates superior results compared to the existing state-of-the-art work. Source code is available https://github.com/shuaichaochao/TDH. Hongxi Wei |
ICASSP | 2 |
| 2023 | AFFGANwriting: A Handwriting Image Generation Method Based on Multi-feature Fusion
Hongxi Wei |
ICDAR (4) | 3 |
| 2023 | A Hybrid Approach Using Convolution and Transformer for Mongolian Ancient Documents Recognition
Shiwen Sun, Hongxi Wei |
ICONIP (13) | 2 |
| 2023 | Multi-feature Fusion-Based Central Similarity Deep Supervised Hashing
Hongxi Wei |
PRCV (7) | 2 |
| 2022 | An Approach Based on Transformer and Deformable Convolution for Realistic Handwriting Samples GenerationabstractIn the field of handwritten recognition, it usually needs to collect a large number of handwriting samples to obtain better results. However, it is so time-consuming and expensive that such a large number of samples cannot be collected manually. There is an effective solution to the above problem through synthesizing handwriting samples. In this study, an approach based on Transformer and deformable convolution has been proposed to generate much more realistic handwriting samples. To be specific, a Transformer has been utilized to capture the relations between writing style and text strings. By this way, global features (e.g. thickness of stroke, slant and so on) of writing style can be obtained. Meanwhile, a feature deformation fusion (FDF) module has been designed for obtaining local features (i.e. personalization) of writing style. Moreover, a focal frequency loss (FFL) is employed to solve the problem of pen-level artifacts. Experimental results demonstrate that the proposed approach can be competent for the task of handwriting samples generation and outperforms various baseline and state-of-the-art methods. Shiwen Sun, Hongxi Wei |
ICPR | 4 |
| 2022 | A Mongolian Handwritten Word Images Generation Approach Based on Generative Adversarial NetworksabstractThe particular word formation manner of the Mongolian language makes its vocabulary reach millions. It is tough to manually collect a Mongolian handwritten word images dataset covering all words. In order to realize the automatic generation of Mongolian handwritten word images, an approach based on generative adversarial networks has been proposed in this paper, which is called Mongolian Generator (MG). The proposed MG generates Mongolian handwritten word images with better writing details and more accurate text content. Also, it can generate handwritten word images based on a specific writing style or text content. Moreover, the perceptual adversarial loss has been integrated into the MG, which guarantees that the generated handwritten word images appear more realistic. The experimental results indicate that the performance of the proposed MG outperforms several baselines and state-of-the-art models. Shiwen Sun, Hongxi Wei |
IJCNN | 2 |
| 2022 | Joint Learning Method Based On Transformer For Image RetrievalabstractIn order to solve the problem that the supervised information is not effectively used and lack of correlation between image features in image retrieval, a joint learning method based on Transformer (JLMT) for image retrieval is proposed in this paper. This method not only combines Transformer with asymmetric learning strategy to extract image features and generate the correlation between image features, but also combines the classification loss to make full use of the supervised information. For training images, we uses the Transformer to generate hash codes, classification loss and the semantic similarity loss are used to learn the hash function, so that the image hash codes are closer to the real hash codes. For the database (retrieval) images, we use the asymmetric learning strategy to learn the hash codes directly through the hash codes of the training images. Finally, the query images generate hash codes through hash function and search similar images in the database image set according to the Hamming distance. Also, a new loss function for the classification of multi-label datasets is proposed in this paper. Experimental results show that the JLMT has the state-of-the-art performance on public datasets of CIFAR-10, NUS-WIDE and MS-COCO. Hongxi Wei |
IJCNN | 1 |
| 2021 | Data Augmentation Based on CycleGAN for Improving Woodblock-Printing Mongolian Words Recognition
Hongxi Wei, Daoerji Fan |
ICDAR (4) | 1 |
| 2020 | A Hybrid Representation of Word Images for Keyword Spotting
Hongxi Wei |
ICONIP (4) | 1 |
| 2020 | Multi-Task Learning Based Traditional Mongolian Words RecognitionabstractIn this paper, a multi-task learning framework has been proposed for solving and improving traditional Mongolian words recognition. To be specific, a sequence-to-sequence model with attention mechanism was utilized to accomplish the task of recognition. Therein, the attention mechanism is designed to fulfill the task of glyph segmentation during the process of recognition. Although the glyph segmentation is an implicit operation, the information of glyph segmentation can be integrated into the process of recognition. After that, the two tasks can be accomplished simultaneously under the framework of multi-task learning. By this way, adjacent image frames can be decoded into a glyph more precisely, which results in improving not only the performance of words recognition but also the accuracy of character segmentation. Experimental results demonstrate that the proposed multi-task learning based scheme outperforms the conventional glyph segmentation-based method and various segmentation-free (i.e. holistic recognition) methods. Hongxi Wei, Hui Zhang 0031 |
ICPR | 1 |
| 2020 | Deep Features Representation of Word Image for Keyword Spotting in Historical Mongolian Document ImagesabstractDue to degradation of historical Mongolian documents, a task for retrieving them is challenging. In the field of document image retrieval, keyword spotting technology is an alternative when optical character recognition is infeasible. Representation of word images plays a very important role in keyword spotting. In this paper, various of convolutional neural networks have been used for representing word images of historical Mongolian documents. To be specific, activations of the fully-connected layer in convolutional neural network are extracted and taken as representation vectors of word images. And then, similarity can be calculated between their representation vectors of word images. Several classic structures of convolutional neural networks have been compared with each other and the best one has been determined. Furthermore, convolutional neural network has been also compared with several baselines and the state-of-the-art method on a dataset of historical Mongolian documents. Experimental results indicates that the performance of convolutional neural network is superior to these baseline and state-of-the-art methods. Hongxi Wei, Hui Zhang 0031 |
ICTAI | 1 |
| 2019 | Woodblock-Printing Mongolian Words Recognition by Bi-LSTM with Attention MechanismabstractWoodblock-printing Mongolian documents are seriously degraded due to aging. Therefore, it is difficult to segment woodblock-printing Mongolian words are into individual glyphs. In this paper, a holistic recognition approach based on sequence to sequence model has been proposed for the woodblock-printing Mongolian words. The input of the proposed model is the sequence of frames of a wood-block printing Mongolian word. In order to generating the corresponding sequence of frames, each word image should be normalized into the same sizes in advance. And then, each word image is segmented into several fragments with equal size along writing direction. The output of the proposed model is a sequence of letters. To be specific, the proposed model contains three parts: an encoder, a decoder and an attention network. The encoder consists of a deep neural network and a bi-directional Long Short-Term Memory (Bi-LSTM). The decoder consists of a Long Short-Term Memory (LSTM) with a softmax layer. The encoder and decoder are connected by an attention network, which can map multiple frames to one letter. Experimental results demonstrate that the proposed approach outperforms the segmentation based method. Yanke Kang, Hongxi Wei, Hui Zhang 0031, Guanglai Gao |
ICDAR | 2 |
| 2019 | A Holistic Recognition Approach for Woodblock-Print Mongolian Words Based on Convolutional Neural NetworkabstractThis paper proposed a holistic recognition approach for woodblock-print Mongolian words using a convolutional neural network (CNN). To be specific, the whole word image is regarded as input of CNN. Hence, all the word images should be normalized into the same size before being inputted into CNN. By comparison, an appropriate normalization size has been determined in our study. Through the above manner, the woodblock-print Mongolian word images do not need to be segmented into glyphs. Thereby, the segmentation errors can be avoided under the circumstance. Furthermore, to solve the problem of imbalance distribution on our dataset, SMOTE technique is adopted to generate samples. In this way, the training procedure can be more efficient and the obtained CNN is more robust. Experimental results demonstrate that the proposed approach outperforms the segmentation based method and other baselines. Hongxi Wei, Guanglai Gao |
ICIP | 1 |
| 2019 | End-to-End Model for Offline Handwritten Mongolian Word Recognition
Hongxi Wei, Hui Zhang 0031, Feilong Bao, Guanglai Gao |
NLPCC (2) | 1 |
| 2018 | Convolutional Neural Network for Machine-Printed Traditional Mongolian Font Recognition
Hongxi Wei, Weiyuan Wang, Guanglai Gao |
ICONIP (5) | 1 |
| 2018 | Word Image Representation Based on Visual Embeddings and Spatial Constraints for Keyword Spotting on Historical DocumentsabstractThis paper proposed a visual embeddings approach to capturing semantic relatedness between visual words. To be specific, visual words are extracted and collected from a word image collection under the Bag-of-Visual-Words framework. And then, a deep learning procedure is used for mapping visual words into embedding vectors in a semantic space. To integrate spatial constraints into the representation of word images, one word image is segmented into several sub-regions with equal size along rows and columns. After that, each sub-region can be represented as an average of embedding vectors, which is the centroid of the embedding vectors of all visual words within the same sub-region. By this way, one word image can be converted into a fixed-length vector by concatenating the corresponding average embedding vectors from its all sub-regions. Euclidean distance can be calculated to measure similarity between word images. Experimental results demonstrate that the proposed representation approach outperforms Bag-of-Visual-Words, visual language model, spatial pyramid matching, latent Dirichlet allocation, average visual word embeddings and recurrent neural network. Hongxi Wei, Hui Zhang 0031, Guanglai Gao |
ICPR | 1 |
| 2017 | Segmentation-Free Printed Traditional Mongolian OCR Using Sequence to Sequence with Attention ModelabstractMongolian Optical Character Recognition (OCR) systems are required for printed document digitization and Mongolian cultural resources utilization. Existing Mongolian OCR systems are based on segmentation. But, the Mongolian segmentation is more difficult than other languages. So, these methods are highly costly and error suffering. In this study, a segmentation-free based traditional Mongolian word recognition method is proposed. Specifically, we formalize the OCR task as a sequence to sequence mapping problem, in which the input Mongolian word image and the output textual string are treated as a sequence of image frames and a sequence of letters, respectively. A sequence to sequence with attention model is adopted to solve this problem. Experimental results on a dataset show the effectiveness of the proposed method. Hui Zhang 0031, Hongxi Wei, Feilong Bao, Guanglai Gao |
ICDAR | 2 |
| 2017 | Representing word image using visual word embeddings and RNN for keyword spotting on historical document imagesabstractVisual words of Bag-of-Visual-Words (BoVW) framework are independent each other, which results in not only discarding spatial orders between visual words but also lacking semantic information. This study is inspired by word embeddings that a similar embedding procedure is applied to a large number of visual words. By this way, the corresponding embedding vectors of the visual words can be formulated. For a word image, the average of embedding vectors of all visual words within the word image is taken as its embedding vector. Moreover, Recurrent Neural Network (RNN) is utilized to encode each word image into embeddings like an auto-encoder. The RNN embeddings and the visual word embeddings are complementary. In this study, all word images are represented by combining visual word embeddings and RNN embeddings. Experimental results show that the proposed representation approach is superior to the traditional BoVW, spatial pyramid matching and latent Dirichlet allocation. Hongxi Wei, Hui Zhang 0031, Guanglai Gao |
ICME | 1 |
| 2017 | Using Word Mover's Distance with Spatial Constraints for Measuring Similarity Between Mongolian Word Images
Hongxi Wei, Hui Zhang 0031, Guanglai Gao, Xiangdong Su |
ICONIP (4) | 1 |
| 2016 | LDA-Based Word Image Representation for Keyword Spotting on Historical Mongolian Documents
Hongxi Wei, Guanglai Gao, Xiangdong Su |
ICONIP (4) | 1 |
| 2016 | A knowledge-based recognition system for historical Mongolian documents
Xiangdong Su, Guanglai Gao, Hongxi Wei, Feilong Bao |
Int. J. Document Anal. Recognit. | 3 |
| 2015 | A multiple instances approach to improving keyword spotting on historical Mongolian document imagesabstractFor keyword spotting of historical Mongolian document images, when user provides different instance image for the same query keyword, the performance will vary a lot. This paper proposed an approach to solving the above problem. Particularly, the whole procedure of keyword spotting is divided into two stages. The main task of the first stage is to generate multiple ranking lists for a query keyword. And the aim of the second stage is to merge the multiple ranking lists to form a final ranking. In the first stage, the ranking list of one query keyword is firstly returned by traditional image matching and then a number of instances for the query keyword are obtained using pseudo relevant feedback. Next, each instance of the query keyword can return the corresponding ranking list separately. In the second stage, the multiple ranking lists from the multiple instances of the query keyword are combined by the data fusion technique. The final ranking will be taken as the retrieval results of the query keyword. The experimental results show that the proposed approach can significantly improve the performance of keyword spotting for the historical Mongolian document images. Hongxi Wei, Guanglai Gao, Xiangdong Su |
ICDAR | 1 |
| 2015 | Enhancing the Mongolian Historical Document Recognition System with Multiple Knowledge-Based Strategies
Xiangdong Su, Guanglai Gao, Hongxi Wei, Feilong Bao |
ICONIP (2) | 3 |
| 2014 | A keyword retrieval system for historical Mongolian document images
Hongxi Wei, Guanglai Gao |
Int. J. Document Anal. Recognit. | 1 |
| 2013 | Word Spotting Application in Historical Mongolian Document Images
Hongxi Wei, Guanglai Gao |
ICIC (1) | 1 |
| 2011 | Classical Mongolian Words Recognition in Historical DocumentabstractThere are many classical Mongolian historical documents which are reserved in image form, and as a result it is difficult for us to explore and retrieve them. In this paper, we investigate the peculiarities of classical Mongolian documents and propose an approach to recognize the words in them. We design an algorithm to segment the Mongolian words into several Glyph Units(Glyph Unit abbr. GU). Each GU is consisted of no more than three characters. Then we used a three-stage method to recognize the GUs. At the first stage, all the GUs are classified into nine groups by decision tree using three features of the GUs. At the second stage, the GUs in each group are classified individually by five independent BP Neutral Networks whose inputs are other five feature vectors of the GUs. At the last stage, the five results of each GU group from the above five classifiers are combined to provide the final recognized result. The recognition rate of the Mongolian words in our experiment achieves 71%, indicating that our method is effective. Guanglai Gao, Xiangdong Su, Hongxi Wei, Yeyun Gong |
ICDAR | 3 |
| 2011 | A Method for Removing Inflectional Suffixes in Word Spotting of Mongolian KanjurabstractAccording to characteristics of Mongolian word-formation, a method for removing inflectional suffixes from word images of the Mongolian Kanjur is proposed in this paper. By removing inflectional suffixes, the amount of clusters equivalent indexing terms might be reduced in word spotting. For the above purpose, we need to determine whether or not one word image contains inflectional suffix. If the word image contains inflectional suffix, the inflectional suffix would be segmented from the word image. The proposed method is as follows: first, many parts are segmented from the bottom of the word image according to the cutting positions of the inflectional suffixes. Then, the segmented parts are represented by a number of profile features and classified by multi-BP neural networks. Finally, the outputs of BP are confirmed by template matching using DTW. Experimental results on our data set prove the feasibility of the proposed method. Hongxi Wei, Guanglai Gao, Yulai Bao |
ICDAR | 1 |