Nurbiya Yadikar

dblp:158/1015 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
11since 2021 · last 2026
0009-0007-7955-7846ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 8 since 2021Artificial intelligence and machine learning · 7 · 5 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 MTTrack: A joint mamba-transformer framework with memory enhancement for real-time satellite remote sensing video object tracking
Guocai Du, Peiyong Zhou, Nurbiya Yadikar, Alimjan Aysa, Kurban Ubul
Knowl. Based Syst.3
2025 Dynamic token sampling for efficient unmanned aerial vehicles transformer tracking
Guocai Du, Peiyong Zhou, Nurbiya Yadikar, Alimjan Aysa, Kurban Ubul
Eng. Appl. Artif. Intell.3
2025 Toward a dynamic tree-Mamba encoder for UAV tracking with vision-language
Guocai Du, Peiyong Zhou, Nurbiya Yadikar, Alimjan Aysa, Kurban Ubul
Knowl. Based Syst.3
2024 MMHSV: A Multimodal Handwritten Signature Verification Fusing Dynamic and Static Feature
abstract
In recent years, significant progress has been made in the field of handwritten signature verification through methods based on deep learning. However, due to the high intra-class variability and high inter-class similarity of signature samples, achieving high accuracy and security in handwritten signature verification systems remains challenging. In this paper, we propose a deep learning-based multimodal handwritten signature verification framework, MMHSV, to fuse dynamic and static features by utilizing the complementary signal components of signature images and pen stroke sounds, and explore the feasibility of multimodal signature verification. MMHSV consists of an innovative joint-embedded feature representation method based on multi-task learning and a dual-path network model for signature representation extraction. To support our method, we have curated a multimodal signature dataset, serving as a benchmark for the proposed technique. Preliminary findings suggest that our approach offers a pioneering solution in the realm of handwritten signature verification and security metrics that surpass the benchmarks set by the unimodal state-of-the-art methods.
Qixiang Li, Zhaoya Wang, Nurbiya Yadikar, Kurban Ubul
ICASSP4
2024 The Collaboration of 3D Convolutions and CRO-TSM in Lipreading
abstract
Lip reading refers to the recognition of speech solely based on the subtle movements of the lips without audio information. Extracting temporal information in lip reading has always been a challenge in this field. In this work, we propose an effective method for extracting temporal information. Specifically, we make the following contributions: Firstly, We propose a new approach called cro-TSM, which utilizes different channel ratios for temporal shifting based on the existing TSM(Temporal Shift Module). Secondly, we replace the global average pooling of the ResNet with 3D convolutions, which work in collaboration with cro-TSM to extract additional temporal information. Lastly, we apply this method to the state-of-the-art models and achieve a remarkable accuracy of 92.4% on the Lipreading In-The-Wild (LRW) dataset. Our approach surpasses all baseline methods and achieves a new state-of-the-art performance in Lipreading.
Yangzhao Xiang, Mutellip Mamut, Nurbiya Yadikar, Ghalipjan Ibrahim, Kurban Ubul
ICASSP3
2024 DDCTrack: Dynamic Token Sampling for Efficient UAV Transformer Tracking
Guocai Du, Peiyong Zhou, Nurbiya Yadikar, Alimjan Aysa, Kurban Ubul
ICPR (15)3
2024 Oracle Character Recognition Based on Attention Enhancement and Multi-level Feature Fusion
Zhiwang Han, Nurbiya Yadikar, Xuebin Xu, Alimjan Aysa, Kurban Ubul
ICPR (31)2
2024 Online Signature Verification Based on Recurrent Attentional Time-Delay Neural Networks
Xirali Ablat, Qixiang Li, Nurbiya Yadikar, Kurban Ubul
PRCV (15)3
2024 Multimodal Finger Recognition Based on Feature Fusion Attention for Fingerprints, Finger-Veins, and Finger-Knuckle-Prints
Xinbo Lai, Yimin Xue, Tayir Tursun, Nurbiya Yadikar, Kurban Ubul
PRCV (15)4
2023 A survey: object detection methods from CNN to transformer
abstract
Abstract Object detection is the most important problem in computer vision tasks. After AlexNet proposed, based on Convolutional Neural Network (CNN) methods have become mainstream in the computer vision field, many researches on neural networks and different transformations of algorithm structures have appeared. In order to achieve fast and accurate detection effects, it is necessary to jump out of the existing CNN framework and has great challenges. Transformer’s relatively mature theoretical support and technological development in the field of Natural Language Processing have brought it into the researcher’s sight, and it has been proved that Transformer’s method can be used for computer vision tasks, and proved that it exceeds the existing CNN method in some tasks. In order to enable more researchers to better understand the development process of object detection methods, existing methods, different frameworks, challenging problems and development trends, paper introduced historical classic methods of object detection used CNN, discusses the highlights, advantages and disadvantages of these algorithms. By consulting a large amount of paper, the paper compared different CNN detection methods and Transformer detection methods. Vertically under fair conditions, 13 different detection methods that have a broad impact on the field and are the most mainstream and promising are selected for comparison. The comparative data gives us confidence in the development of Transformer and the convergence between different methods. It also presents the recent innovative approaches to using Transformer in computer vision tasks. In the end, the challenges, opportunities and future prospects of this field are summarized.
Ershat Arkin, Nurbiya Yadikar, Xuebin Xu, Alimjan Aysa, Kurban Ubul
Multim. Tools Appl.2
2021 How to Use Time Information Effectively? Combining with Time Shift Module for Lipreading
abstract
Lipreading refers to recognizing the speaker's speech content through the image sequence of lip movement without the speech signal. Currently, most models use a spatiotemporal (3D) convolutional layer combined with 2D CNN to extract spatial and temporal features from image sequences. However, compared with 2D convolutional layers, which can extract fine-grained spatial features from the spatial domain, the single-layer 3D convolutional layer used in the model cannot extract temporal information well. This point is improved in this paper. Firstly, the Time Shift Module (TSM) is applied to two different front-ends (full 2D CNN based and mixture of 2D and 3D convolution) to enhance the ability of time information extraction. Secondly, the influence of different shift proportion of TSM and different sampling interval input on extracting time information is verified. Thirdly, the influence of different time shifts on the ability of spatiotemporal feature extraction is compared. The proposed method verified on two challenging word-level lipreading datasets LRW and LRW-1000 and achieved new state-of-the-art performance.
Mingfeng Hao, Mutallip Mamut, Nurbiya Yadikar, Alimjan Aysa, Kurban Ubul
ICASSP3
2018 Script Identification of Central Asia Based on Fused Texture Features
abstract
Script identification is an important step in multi-script recognition. Despite the achieved results in this field, the identification of Central Asian scripts has not been considered in-depth. In the Central Asian region, there are many similar scripts, and the traditional texture features can not discriminate them accurately. This paper proposes a script identification method based on fused texture features for Central Asian document images. On preprocessed multilingual document images, the method first performs Non-subsampled Contourlet Transform (NSCT), and then extracts Tamura texture features of the generated sub-bands. A Support Vector Machine (SVM) classifier is trained for classification. For experimental evaluation, it is collected a dataset of 30, 000 document images for 10 scripts, such as Arabic, Chinese, English, Russian, Kazakhstan, Turkish, Uyghur, Kyrgyzstan, Mongolian and Tibetan. The experimental results show that the proposed method can extract multi-scale and multi-directional texture features, and the fusion of texture features leads to superior performance of script identification.
Xing-kun Han, Alimjan Aysa, Hornisa Mamat, Nurbiya Yadikar, Kurban Ubul
ICPR4
2017 Script Identification Based on Nonsubsampled Contourlet Transform
Xing-kun Han, Alimjan Aysa, Nurbiya Yadikar, Hornisa Mamat, Kurban Ubul
ICDAR3