Kurban Ubul

dblp:87/7435 · DBLP profile ↗
← Back
59ranked-venue papers
0as first author
56since 2021 · last 2026
0000-0002-7566-6494ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 30 since 2021Databases, data management, data science and information retrieval · 10 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CalibLoop: Self-Calibrating Pseudo-Label Learning for Noise-Robust Source-Free QA Domain Adaptation
Chaorui Shi, Tianqing Zhang, Zhongyuan Yang, Alimjan Aysa, Kurban Ubul
ICIC (24)5
2026 S2C-2DGS: Scale-Robust and Structure-Consistent 2D Gaussian Splatting for Sparse-View Reconstruction
Ghalip lbrahim, Guzalnur Amrurla, Kurban Ubul
ICIC (12)5
2026 MFNet: A Multimodal Fingerprint-Vein Recognition Network with Frequency-Domain Enhancement and Cross-Modal Fusion
Jiachang Wang, Tingting Xian, Alimjan Aysa, Kurban Ubul
ICPR (10)5
2026 HCT-Net: A Hybrid CNN-Transformer Network for Robust Oracle Bone Script Recognition
Yanni Zuo, Lanxin Jiang, Chong Qian, Kurban Ubul
ICPR (3)5
2026 Cross-Granularity Fusion Vision Mamba UNet for medical image segmentation
Tuersunjiang Baidi, Zitong Ren, Kurban Ubul, Alimjan Aysa
Eng. Appl. Artif. Intell.3
2026 A lightweight multi-domain collaborative representation network for infrared small target detection
Zitong Ren, Tuersunjiang Baidi, Saisai Ma, Kurban Ubul
Expert Syst. Appl.5
2026 MTTrack: A joint mamba-transformer framework with memory enhancement for real-time satellite remote sensing video object tracking
Guocai Du, Peiyong Zhou, Nurbiya Yadikar, Alimjan Aysa, Kurban Ubul
Knowl. Based Syst.5
2025 KCGAFormer: When Large-Kernel ConvFormer Meets KAN in Semantic Segmentation
abstract
Vision Transformer, with the distinctive architecture and self-attention mechanisms, had profoundly influenced the field of computer vision, establishing Transformer-based models as benchmarks for semantic segmentation. In this study, we propose a pioneering hybrid model that fuses Kolmogorov-Arnold convolutions with ViT architecture to tackle the intrinsic challenges of semantic segmentation. By leveraging the unique attributes of Kolmogorov-Arnold convolutions, our approach introduces a convolutional attention mechanism within the Vision Transformer framework, effectively alleviating the quadratic complexity associated with self-attention. Furthermore, we integrate large-kernel convolutions and an upsampling module into the decoder, which is designed to enhance feature resolution, capture fine details, and maintain robust performance in complex scenarios for dense prediction tasks. Comprehensive experiments conducted on the ADE20K, Cityscapes, and COCO-Stuff datasets reveal that our method achieves mean Intersection over Union (mIoU) scores of 55.52%, 83.6%, and 51.8%, respectively.
Zhengxing Huang, Enguang Zuo, Alimjan Aysa, Kurban Ubul
ICASSP5
2025 ASANet: Scene Text Recognition With Alternate Self-Attention
abstract
Text recognition in complex scenes is a challenging task in Computer Vision. In this paper, we propose an innovative framework for scene text recognition, ASANet, which features an Alternate Attention Enhancement Encoder and a Masked Dual-modal Decoder. The encoder incorporates a 12-layer Alternating Self-Attention Module (ASAM), consisting of both Channel and Spatial Blocks, which significantly enhance the depth and breadth of feature extraction. The decoder employs a strategy that combines masking and sequence alignment modeling, effectively improving character context relevance and prediction accuracy. Extensive experimental results demonstrate that ASANet achieves state-of-the-art performance across several benchmark datasets, with a notable accuracy of 91.2% on our self-constructed Uyghur text dataset, highlighting its superior performance.
Wenting Xu, Elham Eli, Alimjan Aysa, Xuebin Xu, Kurban Ubul
ICASSP5
2025 Long-tailed Oracle Character Recognition Based on Convolutional Neural Networks and Vision Transformers
abstract
Oracle bone inscriptions, which are among the oldest known hieroglyphics in China, encompass rich historical and cultural information. However, the automatic recognition of oracle characters faces substantial challenges due to issues with data quality and long-tail distribution. This study introduces a hybrid model based on a convolutional neural network (CNN) and a visual transformer (ViT), augmented with a multi-expert learning strategy, to enhance the recognition performance of long-tail oracle bone characters. Initially, the CNN extracts local features, while the ViT captures global dependencies, thereby improving the model’s feature representation capability. Subsequently, through the multi-expert learning mechanism, the prediction results of multiple models are aggregated, effectively mitigating the effects of data imbalance and alleviating the challenges associated with the long-tail distribution of the dataset. Experimental results demonstrate that the proposed model achieves new state-of-the-art performance on two largescale Oracle datasets (Oracle-AYNU and OBC306) in terms of Top-1 accuracy, Top-5 accuracy, F1-Score, and average accuracy, with respective values of 87.97% (93.61%), 96.58% (98.84%), 86.83% (93.49%), and 82.81% (85.23%).
Zhongyuan Yang, Zhiwang Han, Alimjan Aysa, Ghalipjan Ibrahim, Kurban Ubul
ICASSP5
2025 GM-MoE: Low-Light Enhancement with Gated-Mechanism Mixture-of-Experts
abstract
Low-light enhancement has wide applications in autonomous driving, 3D reconstruction, remote sensing, surveillance, and so on, which can significantly improve information utilization. However, most existing methods lack generalization and are limited to specific tasks such as image recovery. To address these issues, we propose Gated-Mechanism Mixture-of-Experts (GM-MoE), the first framework to introduce a mixture-of-experts network for low-light image enhancement. GM-MoE comprises a dynamic gated weight conditioning network and three sub-expert networks, each specializing in a distinct enhancement task. Combining a self-designed gated mechanism that dynamically adjusts the weights of the sub-expert networks for different data domains. Additionally, we integrate local and global feature fusion within sub-expert networks to enhance image quality by capturing multi-scale features. Experimental results demonstrate that the GM-MoE achieves superior generalization with respect to 25 compared approaches, reaching state-of-the-art performance on PSNR on 5 benchmarks and SSIM on 4 benchmarks, respectively.
Minwen Liao, Hao Bo Dong, Kurban Ubul, Yihua Shao, Ziyang Yan
ICCV4
2025 A Novel Multi-modal Dataset and Method for Handwritten Signature Recognition with Image-Audio Fusion
Qixiang Li, Xirali Ablat, Xiaoya Lin, Mahpirat, Kurban Ubul
ICDAR (1)5
2025 Multi-scale Convolution Combined with DTW for Online Signature Verification
Dengshan Yang, Mahpirat, Xuebin Xu, Alimjan Aysa, Kurban Ubul
ICDAR (2)5
2025 Scene Script Identification Using Dense Hierarchical Semantic Fusion
Yaowei Yang, Kaisaier Tuerxun, Kurban Ubul
ICDAR (3)4
2025 Multiscale and Multidimensional Lightweight Network for Small-Target Detection in UAV Images
Hongji Ma, Kurban Ubul
ICIC (11)4
2025 Wavelet-Enhanced Convolution with Multiscale Aggregation Network for Small-Target Detection in UAV Images
Kurban Ubul
ICIC (2)4
2025 MAFF-CrossNet: Multi-scale Attentional Feature Fusion and Cross-Writer Network for Offline Signature Verification
Mahpirat, Xuebin Xu, Alimjan Aysa, Kurban Ubul
ICIC (17)5
2025 EROCNet: Multi-Scale Spectral Synergy for Resource-Efficient Retinal Optical Coherence Tomography Diagnosis
abstract
As a non-invasive diagnostic modality, retinal optical coherence tomography (OCT) imaging has become a cornerstone of clinical ophthalmology. However, existing deep learning-based classification solutions face challenges of data dependence and computational complexity, particularly in resource-constrained scenarios. In this work, we present EROCNet (Efficient Retinal OCT Classification Network), a lightweight architecture that integrates multi-mechanism learning via coordinated technical innovations. The framework utilizes progressively dilated convolutions with exponentially scaled dilation rates in shallow layers to achieve hierarchical context aggregation, capturing local textures to global pathologies. In deeper stages, wavelet-based attention mechanisms are employed through discrete wavelet transforms and temperature-regulated spatial-channel reweighting to enhance discriminative patterns. Additionally, we improve model interpretability via heatmap visualizations, providing more transparent diagnostic insights. Comprehensive evaluations show that EROCNet achieves 98.04% precision on the 8-class Retinal OCT C-8 dataset and 99.90% accuracy on OCT2017 (4-class), with only 1.05M parameters and 0.32 GFLOPs computation. Under severe data scarcity (1/32 of the original training data), it maintains 99.20% accuracy, significantly outperforming existing lightweight pretrained models, while demonstrating dual robustness in computational efficiency and data efficiency with deployment potential in resource-limited clinics.
Tuersunjiang Baidi, Kurban Ubul
IJCNN2
2025 PRNet: Parallel Refinement Network with Selective Feature Enhancement for Infrared Small Target Detection
abstract
Infrared small target detection is a key task in computer vision and plays a crucial role in military defense, aerospace, and maritime surveillance. However, it remains a challenging task due to the small target size, low contrast, and complex background noise. Traditional methods, despite the progress made, still suffer from limited robustness and generalization capabilities. On the other hand, deep learning-based methods may fail to accurately recognize deep small targets due to insufficient background noise suppression. For this reason, we propose a parallel refinement network (PRNet). This network uses MobileNet v3 as the backbone for feature extraction, and then we propose the selective feature enhancement module (SFEM) to selectively enhance the extracted features to effectively improve the representativeness of the target features while suppressing the background interference. Finally, we propose the parallel refinement module (PRM) to optimize and aggregate features at different levels to achieve high accuracy and enhance the robustness of object detection. Experimental results on the public dataset MDFA show that the proposed method outperforms the current state-of-the-art methods in terms of detection performance.
Kurban Ubul
ICMR3
2025 Edge-Aware Network with Confidence Feature Fusion for Infrared Small Target Detection
abstract
Infrared Small Target Detection (IRSTD) plays a critical role in various fields such as military surveillance, autonomous driving, and environmental monitoring. However, traditional IRSTD methods often face high false alarm rates and missed detections when dealing with complex backgrounds, significantly limiting their effectiveness in real-world applications. Although recent advancements in deep learning-based IRSTD techniques have achieved remarkable progress, the detected target edges are often unclear and tend to be overly smooth. Additionally, when detecting extremely small targets, missed detections remain prevalent due to interference from background clutter. To address these issues, we propose a novel Edge-Aware Network (EANet). In the encoder stage, EANet employs a combination of residual blocks and max-pooling layers to extract feature maps containing multi-scale information. To further enhance the edge features of targets, we propose an edge-aware enhancement module (EAEM), which integrates convolutional layers with different receptive fields and central difference convolution to effectively improve boundary segmentation accuracy. In the decoder stage, we propose a confidence feature fusion module (CFFM), which incorporates deep saliency information to guide shallow features. This enables the network to focus on critical features in the target regions while suppressing background noise, resulting in a more complete target shape. Experimental results demonstrate that the proposed EANet significantly improves detection accuracy and segmentation performance for infrared small targets under various complex scenarios. It performs exceptionally well for targets of different scales and under low signal-to-noise ratio conditions.
Zitong Ren, Kurban Ubul
ICMR4
2025 Writer-Independent Signature Verification Method Without Forgery Signature Training
Mahpirat, Xuebin Xu, Alimjan Aysa, Kurban Ubul
PRCV (15)5
2025 DSA-Net: A Hybrid SigNet and Selective Attention Network for Handwritten Signature Verification
Xiaoya Lin, Mahpirat, Kurban Ubul
PRCV (15)3
2025 C3F: A Coarse-to-Fine Feature Fusion Approach for Scene Script Identification
Yaowei Yang, Kaisaier Tuerxun, Alimjan Aysa, Kurban Ubul
PRCV (6)5
2025 MTFM: Multi-target Feature Mining Hashing for Fine-Grained Cave Mural Retrieval
Ismail Tursun, Yanni Zuo, Turhunay Sultan, Elham Eli, Kurban Ubul
PRCV (1)5
2025 Towards Uyghur Lip-Reading: Dataset Development and Attention-Enhanced Recognition with ECA-S Module
Siyuan Wei, Zilong Xing, Mutallip Mamut, Kurban Ubul
PRCV (15)4
2025 BYNet: A Bayesian Approach for Robust Palmprint and Palm Vein Recognition
Tingting Xian, Jiachang Wang, Kurban Ubul
PRCV (15)4
2025 Dynamic token sampling for efficient unmanned aerial vehicles transformer tracking
Guocai Du, Peiyong Zhou, Nurbiya Yadikar, Alimjan Aysa, Kurban Ubul
Eng. Appl. Artif. Intell.5
2025 A comprehensive review of non-Latin natural scene text detection and recognition techniques
Elham Eli, Wenting Xu, Hornisa Mamat, Alimjan Aysa, Kurban Ubul
Eng. Appl. Artif. Intell.6
2025 HyperSegmenter: Reappraising the potential of large kernel CNN architecture in efficient semantic segmentation
Zhengxing Huang, Xirali Ablat, Alimjan Aysa, Kurban Ubul
Expert Syst. Appl.5
2025 Toward a dynamic tree-Mamba encoder for UAV tracking with vision-language
Guocai Du, Peiyong Zhou, Nurbiya Yadikar, Alimjan Aysa, Kurban Ubul
Knowl. Based Syst.5
2025 DPA-MVSNet: Dynamic Context Perception Multi-view Stereo with transformers and data augmentation
Jianjun Ji, Xuebin Xu, Alimjan Aysa, Kurban Ubul
Knowl. Based Syst.6
2025 PGNet: Position guided infrared small target detection
Zitong Ren, Shengbin Hao, Kurban Ubul
Knowl. Based Syst.4
2024 MMHSV: A Multimodal Handwritten Signature Verification Fusing Dynamic and Static Feature
abstract
In recent years, significant progress has been made in the field of handwritten signature verification through methods based on deep learning. However, due to the high intra-class variability and high inter-class similarity of signature samples, achieving high accuracy and security in handwritten signature verification systems remains challenging. In this paper, we propose a deep learning-based multimodal handwritten signature verification framework, MMHSV, to fuse dynamic and static features by utilizing the complementary signal components of signature images and pen stroke sounds, and explore the feasibility of multimodal signature verification. MMHSV consists of an innovative joint-embedded feature representation method based on multi-task learning and a dual-path network model for signature representation extraction. To support our method, we have curated a multimodal signature dataset, serving as a benchmark for the proposed technique. Preliminary findings suggest that our approach offers a pioneering solution in the realm of handwritten signature verification and security metrics that surpass the benchmarks set by the unimodal state-of-the-art methods.
Qixiang Li, Zhaoya Wang, Nurbiya Yadikar, Kurban Ubul
ICASSP5
2024 The Collaboration of 3D Convolutions and CRO-TSM in Lipreading
abstract
Lip reading refers to the recognition of speech solely based on the subtle movements of the lips without audio information. Extracting temporal information in lip reading has always been a challenge in this field. In this work, we propose an effective method for extracting temporal information. Specifically, we make the following contributions: Firstly, We propose a new approach called cro-TSM, which utilizes different channel ratios for temporal shifting based on the existing TSM(Temporal Shift Module). Secondly, we replace the global average pooling of the ResNet with 3D convolutions, which work in collaboration with cro-TSM to extract additional temporal information. Lastly, we apply this method to the state-of-the-art models and achieve a remarkable accuracy of 92.4% on the Lipreading In-The-Wild (LRW) dataset. Our approach surpasses all baseline methods and achieves a new state-of-the-art performance in Lipreading.
Yangzhao Xiang, Mutellip Mamut, Nurbiya Yadikar, Ghalipjan Ibrahim, Kurban Ubul
ICASSP5
2024 Oracle Bone Inscriptions Image Retrieval Based on Metric Learning
Jiaoyan Wang, Alimjan Aysa, Xuebin Xu, Kurban Ubul
ICDAR (3)5
2024 Script Identification in the Wild with FFT-Multi-grained Mix Attention Transformer
Zhi Pan, Yaowei Yang, Kurban Ubul, Alimjan Aysa
ICDAR (2)3
2024 A New Bottom-Up Path Augmentation Attention Network for Script Identification in Scene Images
Zhi Pan, Yaowei Yang, Kurban Ubul, Alimjan Aysa
ICDAR (5)3
2024 Improving Retrieval-Based Dialogue Systems: Fine-Grained Post-training Prompt Adaptation and Pairwise Optimization Fine-Tuning Strategy
Tianqing Zhang, Alimjan Aysa, Kurban Ubul, Enguang Zuo
ICDAR (6)4
2024 Oracle Bone Inscription Image Retrieval Based on Improved ResNet Network
Jiaoyan Wang, Alimjan Aysa, Xuebin Xu, Kurban Ubul
ICPR (21)5
2024 DDCTrack: Dynamic Token Sampling for Efficient UAV Transformer Tracking
Guocai Du, Peiyong Zhou, Nurbiya Yadikar, Alimjan Aysa, Kurban Ubul
ICPR (15)5
2024 Oracle Character Recognition Based on Attention Enhancement and Multi-level Feature Fusion
Zhiwang Han, Nurbiya Yadikar, Xuebin Xu, Alimjan Aysa, Kurban Ubul
ICPR (31)5
2024 One-Shot Classification Is Enough for Automatic Label Mapping
Xiaowen Lin, Alimjan Aysa, Kurban Ubul
ICPR (5)3
2024 Multi-Task Interaction Network Based on a Cross-Attention Fusion Mechanism for Offline Signature Verification
Haotian Meng, Xiaoya Lin, Kurban Ubul, Alimjan Aysa
ICPR (25)3
2024 Scene Uyghur Text Detection Based on Adaptive Feature Fusion
Elham Eli, Alimjan Aysa, Xuebin Xu, Hornisa Mamat, Kurban Ubul
ICPR (20)6
2024 Oracle Bone Script Recognition Based on Multi-scale Feature Fusion and Knowledge Distillation
Jiaoyan Wang, Xuebin Xu, Alimjan Aysa, Kurban Ubul
ICPR (22)5
2024 Online Signature Verification Based on Recurrent Attentional Time-Delay Neural Networks
Xirali Ablat, Qixiang Li, Nurbiya Yadikar, Kurban Ubul
PRCV (15)4
2024 TextViTCNN: Enhancing Natural Scene Text Recognition with Hybrid Transformer and Convolutional Networks
Elham Eli, Wenting Xu, Alimjan Aysa, Hornisa Mamat, Kurban Ubul
PRCV (7)5
2024 Multimodal Finger Recognition Based on Feature Fusion Attention for Fingerprints, Finger-Veins, and Finger-Knuckle-Prints
Xinbo Lai, Yimin Xue, Tayir Tursun, Nurbiya Yadikar, Kurban Ubul
PRCV (15)5
2024 A Stochastic Model for Video Object Tracking
Mohammed Leo, Kurban Ubul, Alimjan Aysa, Shengjie Cheng, Elham Eli, Xuebin Xu
PRCV (12)2
2024 Semantic Consistency-Enhanced Refined Hashing for Fine-Grained Image Retrieval
Shuoshuo Li, Kurban Ubul
PRCV (3)2
2024 SCMA Codebooks Design for Three Optimization Algorithms Based on Eisenstein Integer Unit Circle
abstract
Codebook design plays a crucial role in non-orthogonal sparse code multiple access technology. In this paper, a mother constellation constructed by the Eisenstein integer unit circle in the complex plane is proposed, and the power imbalance and dimensionality reduction are introduced into the codebook design. Three optimization algorithms are used to maximize the minimum Euclidean distance (MED) of superimposed codewords on the resource element as the objective function, and the optimal solution of the rotation angle is finally obtained, so as to obtain the three expected codebooks. Simulation results demonstrate that the proposed codebooks have smaller bit error rate (BER) and better performance than the four benchmark codebooks provided under the condition of a Gaussian channel.
Teng Wan, Wenping Ge, Kurban Ubul
SMC3
2024 Research on knowledge distillation algorithm based on Yolov5 attention mechanism
Shengjie Cheng, Peiyong Zhou, Yuliu, Hongjima, Alimjan Aysa, Kurban Ubul
Expert Syst. Appl.6
2023 SUCOLA: Self-adaptive structure refinement unsupervised contrastive learning framework for food safety risk early warning
Enguang Zuo, Junyi Yan, Alimjan Aysa, Chen Chen 0078, Hongbing Ma, Xiaoyi Lv, Kurban Ubul
Eng. Appl. Artif. Intell.8
2023 A survey: object detection methods from CNN to transformer
abstract
Abstract Object detection is the most important problem in computer vision tasks. After AlexNet proposed, based on Convolutional Neural Network (CNN) methods have become mainstream in the computer vision field, many researches on neural networks and different transformations of algorithm structures have appeared. In order to achieve fast and accurate detection effects, it is necessary to jump out of the existing CNN framework and has great challenges. Transformer’s relatively mature theoretical support and technological development in the field of Natural Language Processing have brought it into the researcher’s sight, and it has been proved that Transformer’s method can be used for computer vision tasks, and proved that it exceeds the existing CNN method in some tasks. In order to enable more researchers to better understand the development process of object detection methods, existing methods, different frameworks, challenging problems and development trends, paper introduced historical classic methods of object detection used CNN, discusses the highlights, advantages and disadvantages of these algorithms. By consulting a large amount of paper, the paper compared different CNN detection methods and Transformer detection methods. Vertically under fair conditions, 13 different detection methods that have a broad impact on the field and are the most mainstream and promising are selected for comparison. The comparative data gives us confidence in the development of Transformer and the convergence between different methods. It also presents the recent innovative approaches to using Transformer in computer vision tasks. In the end, the challenges, opportunities and future prospects of this field are summarized.
Ershat Arkin, Nurbiya Yadikar, Xuebin Xu, Alimjan Aysa, Kurban Ubul
Multim. Tools Appl.5
2023 Multimodal sentiment system and method based on CRNN-SVM
abstract
Abstract Traditional sentiment analysis focuses on text-level sentiment mining, transforming sentiment mining into classification or regression problems, resulting in a sentiment analysis low accuracy rate. Sentiment analysis refers to the use of natural language processing, text analysis, and computational linguistics to systematically identify, extract, quantify, and study sentimental states. Therefore, more scholars have begun to focus on speech recognition and facial expression recognition research, and extracting and analysing people’s sentiment tendencies can improve sentiment recognition accuracy. Traditional single-modal sentiment analysis can no longer meet people’s needs. Therefore, this paper proposes a multimodal sentiment analysis method based on the multimodal sentiment analysis method that can obtain more sentimental information sources and help people make better decisions. The experimental results in this paper show that the highest recognition rates of CNN-SVM, RNN-SVM, and CRNN-SVM were 76.8%, 71.2%, and 93.5%, respectively. It can be seen that CRNN-SVM has the highest sentiment tendency recognition rate in deep learning, so it is suitable to apply CRNN-SVM to sentiment tendency analysis system design in this paper. The average accuracy rate of the system designed in this paper was 91%, and the stability was also very strong, which shows that the system designed in this paper is meaningful. The main contribution of this paper is based on the limitations of single-mode emotion analysis. It proposes a multimode emotion analysis method and introduces a convolutional neural network to help people obtain more emotional information sources to meet their needs.
Yuxia Zhao, Mahpirat Mamat, Alimjan Aysa, Kurban Ubul
Neural Comput. Appl.4
2021 How to Use Time Information Effectively? Combining with Time Shift Module for Lipreading
abstract
Lipreading refers to recognizing the speaker's speech content through the image sequence of lip movement without the speech signal. Currently, most models use a spatiotemporal (3D) convolutional layer combined with 2D CNN to extract spatial and temporal features from image sequences. However, compared with 2D convolutional layers, which can extract fine-grained spatial features from the spatial domain, the single-layer 3D convolutional layer used in the model cannot extract temporal information well. This point is improved in this paper. Firstly, the Time Shift Module (TSM) is applied to two different front-ends (full 2D CNN based and mixture of 2D and 3D convolution) to enhance the ability of time information extraction. Secondly, the influence of different shift proportion of TSM and different sampling interval input on extracting time information is verified. Thirdly, the influence of different time shifts on the ability of spatiotemporal feature extraction is compared. The proposed method verified on two challenging word-level lipreading datasets LRW and LRW-1000 and achieved new state-of-the-art performance.
Mingfeng Hao, Mutallip Mamut, Nurbiya Yadikar, Alimjan Aysa, Kurban Ubul
ICASSP5
2018 Script Identification of Central Asia Based on Fused Texture Features
abstract
Script identification is an important step in multi-script recognition. Despite the achieved results in this field, the identification of Central Asian scripts has not been considered in-depth. In the Central Asian region, there are many similar scripts, and the traditional texture features can not discriminate them accurately. This paper proposes a script identification method based on fused texture features for Central Asian document images. On preprocessed multilingual document images, the method first performs Non-subsampled Contourlet Transform (NSCT), and then extracts Tamura texture features of the generated sub-bands. A Support Vector Machine (SVM) classifier is trained for classification. For experimental evaluation, it is collected a dataset of 30, 000 document images for 10 scripts, such as Arabic, Chinese, English, Russian, Kazakhstan, Turkish, Uyghur, Kyrgyzstan, Mongolian and Tibetan. The experimental results show that the proposed method can extract multi-scale and multi-directional texture features, and the fusion of texture features leads to superior performance of script identification.
Xing-kun Han, Alimjan Aysa, Hornisa Mamat, Nurbiya Yadikar, Kurban Ubul
ICPR5
2018 Complex Printed Uyghur Document Image Retrieval Based on Modified SURF Features
Aliya Batur, Patigul Mamat, Ya-li Zhu, Kurban Ubul
PRCV (3)5
2017 Script Identification Based on Nonsubsampled Contourlet Transform
Xing-kun Han, Alimjan Aysa, Nurbiya Yadikar, Hornisa Mamat, Kurban Ubul
ICDAR5