Hongxi Wei

dblp:05/10400 · DBLP profile ↗
← Back
16ranked-venue papers in the field
3as first author
11since 2021 · last 2026
0000-0002-2570-4544ORCID · corroborated

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 14 (3 first)Information Retrieval & Web Search · 2
YearPublicationVenuePosition
2026 A MambaVision-Based Cross-Modal Feature Enhancement Network for Scene Text Super-Resolution
Ruichang Zhu, Hongxi Wei
ICDAR (2)2
2025 DCC: Plug-and-Play Dynamic Category Compression for Enhanced Handwritten Text Generation
Hongxi Wei, Shiwen Sun
ICDAR (1)2
2025 VMF-Net: Visual-Aware Multi-representation Fusion Network for Artifact-Free Handwritten Mathematical Expressions Generation
Hongxi Wei
ICDAR (3)2
2025 SFRD: Handwritten Mathematical Expressions Generation by Spatial-Aware Feature Refinement Diffusion
Hongxi Wei, Shiwen Sun
ICDAR (4)2
2025 MambaHash: Visual State Space Deep Hashing Model for Large-Scale Image Retrieval
abstract
Deep image hashing aims to enable effective large-scale image retrieval by mapping the input images into simple binary hash codes through deep neural networks. More recently, Vision Mamba with linear time complexity has attracted extensive attention from researchers by achieving outstanding performance on various computer tasks. Nevertheless, the suitability of Mamba for large-scale image retrieval tasks still needs to be explored. Towards this end, we propose a visual state space hashing model, called MambaHash. Concretely, we propose a backbone network with stage-wise architecture, in which grouped Mamba operation is introduced to model local and global information by utilizing Mamba to perform multi-directional scanning along different groups of the channel. Subsequently, the proposed channel interaction attention module is used to enhance information communication across channels. Finally, we meticulously design an adaptive feature enhancement module to increase feature diversity and enhance the visual representation capability of the model. We have conducted comprehensive experiments on three widely used datasets: CIFAR-10, NUS-WIDE and IMAGENET. The experimental results demonstrate that compared with the state-of-the-art deep hashing methods, our proposed MambaHash has well efficiency and superior performance to effectively accomplish large-scale image retrieval tasks. Source code is available https://github.com/shuaichaochao/MambaHash.git
Hongxi Wei
ICMR2
2024 LABT: A Sequence-to-Sequence Model for Mongolian Handwritten Text Recognition with Local Aggregation BiLSTM and Transformer
Hongxi Wei, Shiwen Sun
ICDAR (2)2
2024 Recognition and Link Prediction of Onomatopoeia Texts with Arbitrary Shapes
Hongxi Wei
ICDAR (3)2
2024 Deepfake In-Air Signature Verification via Two-Channel Model
Hongxi Wei
ICDAR (2)2
2024 HybridHash: Hybrid Convolutional and Self-Attention Deep Hashing for Image Retrieval
abstract
Deep image hashing aims to map input images into simple binary hash codes via deep neural networks and thus enable effective large-scale image retrieval. Recently, hybrid networks that combine convolution and Transformer have achieved superior performance on various computer tasks and have attracted extensive attention from researchers. Nevertheless, the potential benefits of such hybrid networks in image retrieval still need to be verified. To this end, we propose a hybrid convolutional and self-attention deep hashing method known as HybridHash. Specifically, we propose a backbone network with stage-wise architecture in which the block aggregation function is introduced to achieve the effect of local self-attention and reduce the computational complexity. The interaction module has been elaborately designed to promote the communication of information between image blocks and to enhance the visual representations. We have conducted comprehensive experiments on three widely used datasets: CIFAR-10, NUS-WIDE and IMAGENET. The experimental results demonstrate that the method proposed in this paper has superior performance with respect to state-of-the-art deep hashing methods. Source code is available https://github.com/shuaichaochao/HybridHash.
Hongxi Wei
ICMR2
2023 AFFGANwriting: A Handwriting Image Generation Method Based on Multi-feature Fusion
Hongxi Wei
ICDAR (4)3
2021 Data Augmentation Based on CycleGAN for Improving Woodblock-Printing Mongolian Words Recognition
Hongxi Wei, Daoerji Fan
ICDAR (4)1
2019 Woodblock-Printing Mongolian Words Recognition by Bi-LSTM with Attention Mechanism
abstract
Woodblock-printing Mongolian documents are seriously degraded due to aging. Therefore, it is difficult to segment woodblock-printing Mongolian words are into individual glyphs. In this paper, a holistic recognition approach based on sequence to sequence model has been proposed for the woodblock-printing Mongolian words. The input of the proposed model is the sequence of frames of a wood-block printing Mongolian word. In order to generating the corresponding sequence of frames, each word image should be normalized into the same sizes in advance. And then, each word image is segmented into several fragments with equal size along writing direction. The output of the proposed model is a sequence of letters. To be specific, the proposed model contains three parts: an encoder, a decoder and an attention network. The encoder consists of a deep neural network and a bi-directional Long Short-Term Memory (Bi-LSTM). The decoder consists of a Long Short-Term Memory (LSTM) with a softmax layer. The encoder and decoder are connected by an attention network, which can map multiple frames to one letter. Experimental results demonstrate that the proposed approach outperforms the segmentation based method.
Yanke Kang, Hongxi Wei, Hui Zhang 0031, Guanglai Gao
ICDAR2
2017 Segmentation-Free Printed Traditional Mongolian OCR Using Sequence to Sequence with Attention Model
abstract
Mongolian Optical Character Recognition (OCR) systems are required for printed document digitization and Mongolian cultural resources utilization. Existing Mongolian OCR systems are based on segmentation. But, the Mongolian segmentation is more difficult than other languages. So, these methods are highly costly and error suffering. In this study, a segmentation-free based traditional Mongolian word recognition method is proposed. Specifically, we formalize the OCR task as a sequence to sequence mapping problem, in which the input Mongolian word image and the output textual string are treated as a sequence of image frames and a sequence of letters, respectively. A sequence to sequence with attention model is adopted to solve this problem. Experimental results on a dataset show the effectiveness of the proposed method.
Hui Zhang 0031, Hongxi Wei, Feilong Bao, Guanglai Gao
ICDAR2
2015 A multiple instances approach to improving keyword spotting on historical Mongolian document images
abstract
For keyword spotting of historical Mongolian document images, when user provides different instance image for the same query keyword, the performance will vary a lot. This paper proposed an approach to solving the above problem. Particularly, the whole procedure of keyword spotting is divided into two stages. The main task of the first stage is to generate multiple ranking lists for a query keyword. And the aim of the second stage is to merge the multiple ranking lists to form a final ranking. In the first stage, the ranking list of one query keyword is firstly returned by traditional image matching and then a number of instances for the query keyword are obtained using pseudo relevant feedback. Next, each instance of the query keyword can return the corresponding ranking list separately. In the second stage, the multiple ranking lists from the multiple instances of the query keyword are combined by the data fusion technique. The final ranking will be taken as the retrieval results of the query keyword. The experimental results show that the proposed approach can significantly improve the performance of keyword spotting for the historical Mongolian document images.
Hongxi Wei, Guanglai Gao, Xiangdong Su
ICDAR1
2011 Classical Mongolian Words Recognition in Historical Document
abstract
There are many classical Mongolian historical documents which are reserved in image form, and as a result it is difficult for us to explore and retrieve them. In this paper, we investigate the peculiarities of classical Mongolian documents and propose an approach to recognize the words in them. We design an algorithm to segment the Mongolian words into several Glyph Units(Glyph Unit abbr. GU). Each GU is consisted of no more than three characters. Then we used a three-stage method to recognize the GUs. At the first stage, all the GUs are classified into nine groups by decision tree using three features of the GUs. At the second stage, the GUs in each group are classified individually by five independent BP Neutral Networks whose inputs are other five feature vectors of the GUs. At the last stage, the five results of each GU group from the above five classifiers are combined to provide the final recognized result. The recognition rate of the Mongolian words in our experiment achieves 71%, indicating that our method is effective.
Guanglai Gao, Xiangdong Su, Hongxi Wei, Yeyun Gong
ICDAR3
2011 A Method for Removing Inflectional Suffixes in Word Spotting of Mongolian Kanjur
abstract
According to characteristics of Mongolian word-formation, a method for removing inflectional suffixes from word images of the Mongolian Kanjur is proposed in this paper. By removing inflectional suffixes, the amount of clusters equivalent indexing terms might be reduced in word spotting. For the above purpose, we need to determine whether or not one word image contains inflectional suffix. If the word image contains inflectional suffix, the inflectional suffix would be segmented from the word image. The proposed method is as follows: first, many parts are segmented from the bottom of the word image according to the cutting positions of the inflectional suffixes. Then, the segmented parts are represented by a number of profile features and classified by multi-BP neural networks. Finally, the outputs of BP are confirmed by template matching using DTW. Experimental results on our data set prove the feasibility of the proposed method.
Hongxi Wei, Guanglai Gao, Yulai Bao
ICDAR1