Askar Hamdulla

dblp:13/3035 · DBLP profile ↗
← Back
55ranked-venue papers
0as first author
48since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 12 since 2021Databases, data management, data science and information retrieval · 10 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Systems, architecture and hardware · 5 · 4 since 2021Computer networks · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2027 LiteNER: A novel lightweight method for long text named entity recognition
Yelin Chen, Huaping Zhang, Ruohao Yan, Askar Hamdulla
Inf. Process. Manag.5
2026 MoTE-Detox: Multi-dimensional Detoxification of Large Language Models via Mixture-of-Experts-Injected Dataset
Qiwei Dai, Ailiyaer Abudukelimu, Hankiz Yilahun, Abdusalam Dawut, Askar Hamdulla
KSEM (4)5
2026 ARGUE: Towards LLM-Based Fact Checking via an Argumentation-Guided Evidence-Aware Framework
Qiuhua Wu, Ailiyaer Abudukelimu, Hankiz Yilahun, Askar Hamdulla
KSEM (2)4
2026 LLaMA-MoT: A cost-effective framework for visual-linguistic instruction tuning based on multi-head adapters and chain-of-thought
Turdi Tohti, Wenpeng Hu, Tianwei Yan 0001, Shaohuang Wang, Askar Hamdulla
Expert Syst. Appl.6
2026 Cattle herd viewpoint detection based on lightweight convolution and cross-view fusion
Xinxin Luo, Fuzeng Zhang, Eksan Firkat, Askar Hamdulla, Aizimaiti Xiaokaiti, Abdusalam Dawut
Expert Syst. Appl.4
2025 MSTDD: A Multi-scale Transformer Framework for Automatic Depression Detection
Dongfang Han, Yuanyuan Liao, Askar Hamdulla, Turdi Tohti
ADMA (2)5
2025 MSACC: A Unified Multimodal Sentiment Analysis Framework for High Interpretability and Zero-shot Performance
abstract
Compared to large language models, traditional multimodal sentiment analysis frameworks are constrained by their classification heads, resulting in poor performance on zero-shot tasks. Moreover, due to limitations in visual encoders and multimodal fusion modules, most existing frameworks can only process a small number of images, leading to a loss of visual information. In light of these issues, this paper proposes a new framework, MSACC. This framework enhances the model’s zero-shot performance by adopting a contrastive classification method and reduces the loss of visual information through visual relation extraction and three-dimensional sentiment analysis. We conducted extensive experiments on the Yelp dataset. The experimental results show that MSACC outperforms models of the same category in zero-shot MSA tasks, achieving a 48% performance improvement. Furthermore, compared to the large language model ChatGLM2-6B, MSACC still achieved a 7% performance increase while saving 90% of the model size. In addition, in supervised tasks, MSACC also achieved a 3.27% performance improvement compared to the baseline model.
Turdi Tohti, Bo Kong 0002, Dongfang Han, Tianwei Yan 0001, Askar Hamdulla
ICASSP6
2025 SiamMFT: Siamese MultiFrame Network in Infrared Small Target Tracking
Zhuoxu Jiang, Abdusalam Dawut, Askar Hamdulla
ICIC (6)3
2025 ISL-MED: A General Iterative Self-Learning Framework for Speech Complex Emotion Detection
abstract
Speech complex emotion detection (SCED) aims to identify all emotion categories and their intensities in speech, which is crucial for understanding the speaker’s genuine intentions. A significant challenge inherent to SCED is the coarse nature of manual annotations, such as one-hot labels. These labels merely indicate the presence or absence of an emotion, thereby lacking the granularity required to guide models in capturing crucial intensity information. To overcome this limitation, we propose a novel framework: the Iterative Self-Learning-based Multiple Emotion Detector (ISL-MED). This framework utilizes an iterative self-learning approach to infer the latent complex emotion distribution directly from coarse one-hot labels. Specifically, ISL-MED employs multiple dedicated emotion detectors, each responsible for estimating the intensity component for a distinct emotion category. Notably, these detectors can be instantiated using various existing speech emotion recognition (SER) models, and their quantity can be flexibly configured based on task requirements. Furthermore, this paper proposes a data selection strategy based on Curriculum Learning and Human-Machine Consensus (HMCC). This strategy enhances model performance and accelerates convergence by systematically identifying and excluding highly ambiguous samples from the training set. We validated the effectiveness of ISL-MED on the complex emotions dataset CNSCED for speech complex emotion detection tasks and further evaluated its generalization capability on the IEMOCAP dataset for single emotion recognition tasks.
Xinxin Luo, Chang Feng, Hankiz Yilahun, Mingxing Xu, Askar Hamdulla, Thomas Fang Zheng
IJCNN6
2025 Speech Mutil-label Emotion Recognition Using Asymmetric Class Loss Function Based on Effective Samples
Shanshan Xiang, Hankiz Yilahun, Askar Hamdulla
INTERSPEECH3
2025 CalibMutiL: Online Calibration Of LiDAR-Camera Based On Multi-level Visual Feature Fusion
abstract
Multi-sensor fusion is a key technology in the field of autonomous driving and robotics. Traditional offline multi-sensor fusion calibration methods rely on manual operations and fail to meet real-time requirements, while recent online calibration technologies have limited generalization capabilities. This paper proposes CalibMutiL, an end-to-end calibration network that departs from conventional deep feature fusion by leveraging multi-level RGB image features to guide point cloud alignment. CalibMutiL introduces a Multi-level Fusion module (MLF) that effectively utilizes the rich visual features of the image. In addition, we regard the alignment process as a sequence prediction problem and further improve the performance through an Iterative Refinement module (IRM). Evaluation of the KITTI odometry and raw dataset demonstrates the average calibration error reaches 0.81cm and 0.09°. The generalization tests resulted in errors of 4.24cm and 0.13°, outperforming existing methods. Our implementation will be publicly available at https://github.com/VIP-G/CalibMutiL.
Eksan Firkat, Eliyas Suleyman, Bangquan Xie, Fengze Li, Askar Hamdulla
IROS6
2025 HI-SLAM: Hierarchical implicit neural representation for SLAM
Eksan Firkat, Jihong Zhu 0001, Askar Hamdulla
Expert Syst. Appl.6
2025 Enhanced spatial and interaction channel feature network for skin lesion segmentation
Rumeng Wang, Mayire Ibrayim, Askar Hamdulla
Multim. Tools Appl.3
2025 Domain-adaptive transfer network for visual-textual cross-domain sentiment classification
Turdi Tohti, Dongfang Han, Zicheng Zuo, Yuanyuan Liao, Qingwen Yang, Askar Hamdulla
J. Supercomput.8
2025 DARI: Transformer-Based Data Augmentation and Rotation Invariance for UAV Person Re-Identification
abstract
The rapid development of Uncrewed Aerial Vehicles (UAVs) and their unique vantage points present both new opportunities and challenges for person Re-Identification (ReID). Uncertain rotations and scale variations of targets in UAV images, coupled with complex environmental factors, hinder existing methods from extracting robust feature representations. Some methods either make minor modifications to the traditional model architecture or apply simple image rotations but still fail to effectively address the challenges of UAV person ReID. To overcome these limitations, we propose a novel Data Augmentation and Rotation Invariance (DARI) algorithm. First, rotation-invariant convolution is introduced to adaptively extract features, mitigating the uncertainty caused by target rotation. Second, a refined data augmentation correction strategy is employed to reduce noise interference by increasing the richness of global features at different stages. Additionally, considering that multiple features of the same identity should yield consistent recognition result, invariant constraints are designed to enhance the clustering effect. We conducted extensive experiments on both UAV and fixed-camera datasets. The results on PRAI-1581 demonstrate a 5.6% and 6.1% improvement in mAP and Rank-1, respectively, compared to baseline. These findings highlight the model’s effectiveness in addressing the challenges of UAV ReID, demonstrating its robustness and superiority.
Fuzeng Zhang, Eksan Firkat, Hongbing Ma, Jihong Zhu 0001, Askar Hamdulla
IEEE Trans. Multim.6
2024 Recommendation model based on knowledge graphs and semantic alignment
abstract
Most current recommendation models based on knowledge graphs use graph attention mechanisms to perform semantic aggregation of nodes and their neighbors, but this approach can lead to the loss of consistency information among neighboring nodes. Moreover, many recommendation models incorporate contrastive learning to suppress noise in knowledge graphs, often neglecting to enhance model robustness through parameter updates. Therefore, we proposed a recommendation model called FITAALIGN that integrates knowledge graphs with semantic alignment. Firstly, it is argued that the embedded representation of a node aggregated with its neighboring nodes in the knowledge graphs should possess similar semantic representations, which is achieved through similarity calculations to ensure embedded similarity of node representations. Secondly, conditional alignment is used to overcome the loss of node consistency information caused by encoding user-item graphs, thereby obtaining high-quality node semantic representations. Finally, to mitigate the impact of noise interference in the knowledge graphs, adversarial noise is firstly introduced into the model's network parameters, followed by adversarial training to strengthen the robustness of the model. Experiment results on two public datasets validate that our model outperforms some advanced methods. Respectively, on the Recall@20, there was an average increase of 54.99% and 14.57%; on the NDCG@20, the increases were 60.22% and 23.76%; and on the HR@20, the increases were 42.94% and 10.70%.
Zhiyue Xiong, Hankiz Yilahun, Askar Hamdulla
IEEE Big Data3
2024 Source Free Domain Adaptation via Adapting to the Enhanced Style
abstract
Unsupervised domain adaptation (UDA) assumes a labeled source domain and an unlabeled target domain. It aims to train a target domain model by transferring the knowledge from source domain to target domain. With the same goal, source-free domain adaptation (SFDA) only uses the trained source model and target domain data for target model training, which is beneficial for data transmission and data privacy protection etc. Existing SFDA methods suffer from domain shift, resulting in unsatisfactory results. Different from existing methods, we eliminate the domain shift from the domain style perspective. Specifically, we propose a novel method named Adapting to the Enhanced Style (AES). We first increase the diversity of target domain style by style enhancement, then a contrastive loss is used to adapt to the diverse generated styles by requiring consistent feature representations. This process forces our classification model adapting to target domain style, thus increasing the robustness of the model. We conduct extensive experiments on standard benchmarks, and the results show the superiority of our method.
Chaofeng Yang, Zhiheng Zhao, Hankiz Yilahun, Askar Hamdulla
CSCWD4
2024 Cross-Modal Alignment for End-to-End Spoken Language Understanding Based on Momentum Contrastive Learning
abstract
The end-to-end spoken language understanding system extracts the semantic intent directly from an input speech. It effectively avoids problems such as semantic drift in traditional cascade models. However, the lack of semantically labeled speech data makes the model training process diffi-cult. Several recent multi-modal research perspectives have demonstrated that aligning speech and text embeddings based on space distance can improve the model’s performance. In this study, inspired by the work related to contrastive learning, a speech and text aligning method using momentum contrast learning is proposed, and a momentum distillation method is also used in the model to learn from imperfectly matched speech and text data. The proposed method has improved intent detection accuracy by 2.14% and 5.98% on Fluent Speech Command and SmartLights datasets.
Beida Zheng, Mijit Ablimit, Askar Hamdulla
ICASSP3
2024 Attention-based Dual-Branch Network for Micro-Expression Recognition with Global-Local Feature Fusion
abstract
Micro-expression(ME) is an uncontrollable muscle movement that appears on the face when people try to hide or inhibit their real emotions, which has the problems of short duration, small movement amplitude and uneven distribution. In order to solve the problem of localization and asymmetry in the distribution of ME features, this paper proposes a new two-branch attention network to recognize MEs, which utilizes the attention mechanism to capture global and local ME features. The network is mainly divided into three parts: data preprocessing, ME feature learning and feature fusion classification. First, the data preprocessing first extracts the ME optical flow features and then divides them into four regions, which are used as inputs to the two-branch network respectively. Second, the two-branch network uses the Inception-MSFE global multi-scale feature extraction network incorporating the attention mechanism (CBAM) and the Swin Transformer-based local feature extraction network for feature learning, respectively. Finally, MEs are predicted by fusing global and local features of MEs. Experimental validation is carried out on three datasets, CASME II, SAMM, and SMIC, which proves that Acc, UAR, and UF1 are 0.797, 0.702, and 0.698 on the SAMM dataset, respectively; Acc, UAR, and UF1 are 0.734, 0.723, and 0.729 on SMIC, respectively; and on the CASME II dataset, Acc, UAR and UF1 are 0.865, 0.872, and 0.889, respectively, which are competitive with other state-of-the-art methods.
Yupeng Qi, Mayire Ibrayim, Askar Hamdulla
IJCB3
2024 Doc-DINO: A Transformer Model for Complex Logical Document Layout Analysis
Mayire Ibrayim, Askar Hamdulla, Hailong Luo, Chunhu Zhang
ICDAR (4)3
2024 A Real-Time Scene Uyghur Text Detection Network Based on Feature Complementation
Mayire Ibrayim, Askar Hamdulla, Jianjun Kang, Chunhu Zhang
ICDAR (5)3
2024 More and Less: Enhancing Abundance and Refining Redundancy for Text-Prior-Guided Scene Text Image Super-Resolution
Yihong Luo, Mayire Ibrayim, Askar Hamdulla
ICDAR (5)4
2024 Emotional Atmosphere Soft Label for Emotion Recognition in Conversations
Chang Feng, Hankiz Yilahun, Mingxing Xu, Askar Hamdulla, Thomas Fang Zheng
ICONIP (3)5
2024 Multi-Modal Fake News Detection Based on Image Captions
abstract
Faked news could cause hazardous results or even significant social problems, and it is difficult for people to identify fallacies or facts. Hence, the prompt identifying of misleading multi-modal faked news are pressing issues. Multi-modal fake news detection has witnessed rapid advancements in recent years. However, the majority of fake news detection models have primarily focused on using raw image and text features extracted by models. This paper introduces an innovative approach that incorporates external image information, using a multi-modal pre-trained model (OFA) to extract image captions for recognizing image content serves as external information for the image. The proposed method employ multi-modal Coordinate Attention to fuse text and image features. Extensive experiments on two real-world datasets demonstrate the superiority of our model in detecting fake news.
Yadong Gu, Askar Hamdulla, Mijit Ablimit
IJCNN2
2024 A Novel Loss Incorporating Residual Signal Information for Target Speaker Extraction Under Low-SNR Conditions
abstract
Most target speaker extraction (TSE) methods primarily focus on scenarios characterized by conventional signal-to-noise ratio (SNR), overlooking the substantial decline in the quality of extracted speech under low-SNR conditions. To address this issue, we propose to add an item related to the residual signal to the original SI-SDR-based loss function, with no need of additional training labels or network modules. Specifically, we investigate the impact of residual signal information by applying three types of constraints, which are direction, Euclidean distance, and projection distance + Euclidean distance. The proposed methods are experimentally evaluated using the WSJ0-2mix-extr datasets, with SNR range selected as -10 to -5 dB, -15 to -10 dB and -20 to -15 dB, and the SpEx series models are selected as the validation models. Experimental results show that the proposed methods are all superior to the baseline, and the performance improvement of the first two methods increases with the decrease of SNR, but the third method shows better performance when SNR is higher.
Askar Hamdulla, Mijit Ablimit
IJCNN2
2024 Adversarial Training for Uncertainty Estimation in Cross-Lingual Text Classification
abstract
Multilingual pre-trained models have achieved remarkable performance in cross-lingual transfer tasks, but their effectiveness heavily depends on the amount of labeled data available for training. Recent research have demonstrated that self-training for semi-supervised learning can effectively improve deep learning models by utilizing unlabeled data in the presence of limited training data. In this paper, we propose a neural network model with heteroscedastic uncertainty estimation based on adversarial training. The task model for cross-lingual learning consists of a multilingual pre-trained encoder and a dual-channel feature extraction layer, enhancing the model’s ability to model contextual features. In the self-training framework with adversarial perturbations, we utilize pseudo-labeled data for dynamic iterative training. By combining uncertainty estimation and adversarial training, we selectively choose representative samples from unlabeled data, mitigating the issue of label noise propagation and benefiting model training. The proposed approach is evaluated on two cross-lingual datasets, MLDoc and PAWS-X, and experimental results demonstrate the effectiveness of our method.
Lina Xia, Askar Hamdulla, Mijit Ablimit
IJCNN2
2024 G-YOLOv5: A Face Mask Detector That Balances Effectiveness and Real-Time
abstract
In crowded scenarios, face mask detection algorithms still suffer from the target misdetection and omission, and the difficulty of reconciling real time and effectiveness. To address these issues, we present an G-YOLOv5 face mask detection algorithm. First, we adopt a weighted bidirectional feature pyramid network as the feature fusion network to enhance multiscale feature fusion; second, we utilize the WIoU loss function to strengthen the impact of good anchor frames while weakening the detrimental impact of poor quality anchor frames. Finally, considering the real-time and effectiveness issues of detector, we designed the Ghostv2-C3 module instead of the C3 module of the backbone to improve the inference speed of the model. The results of the experiment demonstrates that compared with other detectors, our detector plays a positive role in solving the target misdetection and omission and balancing the real-time nature of the model.
Rong Ye, Mayire Ibrayim, Askar Hamdulla
IJCNN3
2024 UY/CH-CHILD - A Public Chinese L2 Speech Database of Uyghur Children
Mewlude Nijat, Askar Hamdulla
INTERSPEECH4
2024 Few-Shot Keyword Spotting from Mixed Speech
abstract
Few-shot keyword spotting (KWS) aims to detect unknown keywords with limited training samples.A commonly used approach is the pre-training and fine-tuning framework.While effective in clean conditions, this approach struggles with mixed keyword spotting -simultaneously detecting multiple keywords blended in an utterance, which is crucial in real-world applications.Previous research has proposed a Mix-Training (MT) approach to solve the problem, however, it has never been tested in the few-shot scenario.In this paper, we investigate the possibility of using MT and other relevant methods to solve the two practical challenges together: few-shot and mixed speech.Experiments conducted on the LibriSpeech and Google Speech Command corpora demonstrate that MT is highly effective on this task when employed in either the pre-training phase or the fine-tuning phase.Moreover, combining SSL-based large-scale pre-training (HuBert) and MT fine-tuning yields very strong results in all the test conditions.
Junming Yuan, Ying Shi 0001, Lantian Li, Dong Wang 0013, Askar Hamdulla
INTERSPEECH5
2024 Lightweight and Multi-scale Adaptive Network for Infrared Small Target Detection
Shuxian Liu, Hankiz Yilahun, Askar Hamdulla
PRCV (13)4
2024 Document-Level Relation Extraction Model Based on Boundary Distance Loss and Long-Tail Relation Enhancement
Hankiz Yilahun, Seyyare Imam, Askar Hamdulla
PRICAI (2)4
2024 Highway Gates Dynamic Adaptation Network For Knowledge Graph Entity Alignment
Nursharbat Yusuf, Hankiz Yilahun, Seyyare Imam, Askar Hamdulla
PRICAI (4)4
2024 Global and item-by-item reasoning fusion-based multi-hop KGQA
Tongzhao Xu, Turdi Tohti, Askar Hamdulla
Data Knowl. Eng.3
2024 Spatio-temporal mix deformable feature extractor in visual tracking
Ziwang Xiao, Eksan Firkat, Jinlai Zhang, Danfeng Wu, Askar Hamdulla
Expert Syst. Appl.6
2024 Contrastive classification: A label-independent generalization model for text classification
Turdi Tohti, Askar Hamdulla
Expert Syst. Appl.3
2024 FRCE: Transformer-based feature reconstruction and cross-enhancement for occluded person re-identification
Fuzeng Zhang, Hongbing Ma, Jihong Zhu 0001, Askar Hamdulla
Expert Syst. Appl.4
2024 Infrared Small Target Detection Based on Background Estimation and Scale Fusion
abstract
Local contrast methods for infrared small target detection techniques have attracted much attention. However, for some small targets with obvious contours and large scales, the multiscale strategy fails to achieve the expected results. To address this problem, this letter proposes an infrared small target detection method based on background estimation and scale fusion. First, a background estimation filter is proposed based on the idea of background estimation, which is combined with a ratio local contrast to design a ratio-estimated local contrast measurement (RELCM) to initially capture the target while suppressing low-frequency clutter. Second, a single-scale and multiscale differential contrast measurement (SMDCM) is proposed by combining single-scale intrafeatures with multiscale interfeatures to further suppress clutter while enhancing the target. Combine RELCM with SMDCM to obtain the ratio estimation scale fusion local contrast measurement (RESFLCM). Finally, an adaptive thresholding operation is used to extract the final target. Extensive experimental results show that the RESFLCM performs better detecting small targets of different scales in complex backgrounds. Different quantitative results also show that the RESFLCM possesses high signal-to-clutter ratio (SCR) gain and background suppression factor (BSF) with robust detection performance.
Ziling Lu, Shuxian Liu, Hankiz Yilahun, Askar Hamdulla
IEEE Geosci. Remote. Sens. Lett.4
2024 Visual and semantic guided scene text retrieval
Hailong Luo, Mayire Ibrayim, Askar Hamdulla
J. Supercomput.3
2023 Unsupervised word Segmentation Based on Word Influence
abstract
Word segmentation task is the cornerstone of text processing. There are 7111 languages worldwide, most of which are low-resource languages. This paper attempts to solve the problem of multilingual unsupervised word segmentation using common points between languages without tagged corpus. We find that words are only a relationship between phrases and non-phrases in each language, and the frequency of their occurrence obeys the normal distribution. Based on the objective law of language and pre-training language model, this paper defines the concept of Word Influence and designs its calculation formula, and loss function. Combined with the fine-tuning word segmentation task, a multilingual unsupervised word segmentation model was proposed. In order to apply to multiple languages, the model’s key parameters can be learned independently. Its validity and advancement have been proved on Chinese, Japanese, and English data sets. Finally, we discuss the challenges of word segmentation in the pre-trained language model environment.
Ruohao Yan, Huaping Zhang, Wushour Slamu, Askar Hamdulla
ICASSP4
2023 Visualizing Data Augmentation in Deep Speaker Recognition
Pengqi Li, Lantian Li, Askar Hamdulla, Dong Wang 0013
INTERSPEECH3
2023 DASTSiam: Spatio-temporal fusion and discriminative enhancement for Siamese visual tracking
abstract
Abstract The use of deep neural networks has revolutionised object tracking tasks, and Siamese trackers have emerged as a prominent technique for this purpose. Existing Siamese trackers use a fixed template or template updating technique, but it is prone to overfitting, lacks the capacity to exploit global temporal sequences, and cannot utilise multi‐layer features. As a result, it is challenging to deal with dramatic appearance changes in complicated scenarios. Siamese trackers also struggle to learn background information, which impairs their discriminative ability. Hence, two transformer‐based modules, the Spatio‐Temporal Fusion (ST) module and the Discriminative Enhancement (DE) module, are proposed to improve the performance of Siamese trackers. The ST module leverages cross‐attention to accumulate global temporal cues and generates an attention matrix with ST similarity to enhance the template's adaptability to changes in target appearance. The DE module associates semantically similar points from the template and search area, thereby generating a learnable discriminative mask to enhance the discriminative ability of the Siamese trackers. In addition, a Multi‐Layer ST module (ST + ML) was constructed, which can be integrated into Siamese trackers based on multi‐layer cross‐correlation for further improvement. The authors evaluate the proposed modules on four public datasets and show comparative performance compared to existing Siamese trackers.
Eksan Firkat, Jinlai Zhang, Lijuan Zhu, Jihong Zhu 0001, Askar Hamdulla
IET Comput. Vis.7
2023 Infrared Small Target Detection Based on Multidirectional Cumulative Measure
abstract
Robustness of small target detection is a researchable hotspot in infrared surveillance system. The residual phenomenon of background clutter is universal in current local comparison methods. Algorithm of sparse low-rank decomposition restoration cannot be applied to the actual situations due to the long time consumption. This letter proposes a multi-directional cumulative measure (MDCM) to enhance saliency and effectiveness of weak-small target detection. Firstly,multi-directional cumulative mean difference is implemented in central layer and background layer to estimate the background, while multi-directional cumulative derivative multiplying is calculated in central-active layer to characterize overall target’s heterogeneity, then technology of image fusion is adopted to eliminate interference of false target. Finally,a simple adjudicative technology is employed toward separated target region from complex scenes. Compared to up to date existing approaches, extensive simulational testing on four public datasets prove that proposed approach is capable of separating small targets efficiently from an irregular background in a single-scale window and achieve a comparable or even better accuracy.
Guofeng Zhang 0016, Askar Hamdulla, Hongbing Ma
IEEE Geosci. Remote. Sens. Lett.2
2023 Infrared Small Target Detection With Patch Tensor Collaborative Sparse and Total Variation Constraint
abstract
Sparse and low-rank modeling has shown the powerful describing abilities to express small targets; however, low-rank model exists the problem of insufficient rank approximation deviation ability and excessive shrinkage, which will lead to inaccurate background estimation. In this letter, a new nonconvex approximation function using the Gaussian model is built toward deeply excavating low-rank information of the background as much as possible. In the sparse collaborative representation, the local prior confidence (LPC) is integrated into the target structure tensor to maximum to distinguish target region and background edge more accurately. In the process of background restoration, total variation constraint (TVC) is employed apropos of describing gray variation of small targets in complex backgrounds more precisely and improving the accuracy of background restoration to further perfectly recover small targets. The low-rank and sparse recovery algorithm engages alternating direction multiplier method (ADMM) for iterative calculation and solution. Compared with advanced optimal algorithms, a great number of experimental results show that the proposed model improves the adaptability and robustness of the detection algorithm to a variety of complex scenes, and has lasting vitality and high-application value.
Guofeng Zhang 0016, Askar Hamdulla, Hongbing Ma
IEEE Geosci. Remote. Sens. Lett.2
2022 Reliable Visualization for Deep Speaker Recognition
abstract
In spite of the impressive success of convolutional neural networks (CNNs) in speaker recognition, our understanding to CNNs' internal functions is still limited. A major obstacle is that some popular visualization tools are difficult to apply, for example those producing saliency maps. The reason is that speaker information does not show clear spatial patterns in the temporal-frequency space, which makes it hard to interpret the visualization results, and hence hard to confirm the reliability of a visualization tool. In this paper, we conduct an extensive analysis on three popular visualization methods based on CAM: Grad-CAM, Score-CAM and Layer-CAM, to investigate their reliability for speaker recognition tasks. Experiments conducted on a state-of-the-art ResNet34SE model show that the Layer-CAM algorithm can produce reliable visualization, and thus can be used as a promising tool to explain CNN-based speaker models. The source code and examples are available in our project page: http://project.cslt.org/.
Pengqi Li, Lantian Li, Askar Hamdulla, Dong Wang 0013
INTERSPEECH3
2022 Recurrent Deformable Fusion for Compressed Video Artifact Reduction
abstract
The compressed video inevitably appears in compression artifacts, which seriously affect the Quality of Experience. The state-of-the-art methods employ deformable alignment to gather similar information from multiple neighborhood frames to enhance target frame quality. However, they always align multiple frames to the target frame simultaneously, which brings repetitive and useless information because of multiple and imperfect alignments. In this paper, we propose a recurrent deformable fusion method which considers the alignment quality distortion caused by time distance from the target frame. Specifically, a Deformable Alignment (DA) module aligns each pair of the target frame and an adjacent frame following the time line. At the same time, a Recurrent Fusion (RF) module integrates the current aligned feature with the previous fused feature. After that, the fused features are concatenated along the time line. Then, a Multi-Scale Attention Reconstruction (MSAR) module is proposed to gather useful information from the fused features. Compared with the previous multi-frame alignment approach, our method can avoid obtaining a lot of repetitive and useless information. Experiment results confirm that our method achieves state-of-the-art performance on the standard test sequences.
Liuhan Peng, Askar Hamdulla, Mao Ye 0001, Shuai Li 0005, Hongwei Guo 0001
ISCAS2
2021 Adaptive Morphological Contrast Enhancement Based on Quantum Genetic Algorithm for Point Target Detection
Guofeng Zhang 0016, Askar Hamdulla
Mob. Networks Appl.2
2021 Analysis of phonemes and tones confusion rules obtained by ASR
Gulnur Arkin, Askar Hamdulla, Mijit Ablimit
Wirel. Networks2
2021 An adaptive threshold algorithm for offline Uyghur handwritten text line segmentation
Eliyas Suleyman, Askar Hamdulla, Palidan Tuerxun, Kamil Moydin
Wirel. Networks2
2020 A novel deep learning method for query task execution time prediction in graph database
Zheng Chu 0002, Askar Hamdulla
Future Gener. Comput. Syst.3
2020 A benchmark for unconstrained online handwritten Uyghur word recognition
Wujiahemaiti Simayi, Mayire Ibrayim, Xu-Yao Zhang, Cheng-Lin Liu 0001, Askar Hamdulla
Int. J. Document Anal. Recognit.5
2020 LPG-model: A novel model for throughput prediction in stream processing, using a light gradient boosting machine, incremental principal component analysis, and deep gated recurrent unit network
Zheng Chu 0002, Askar Hamdulla
Inf. Sci.3
2020 Infrared Small Target Detection Using Homogeneity-Weighted Local Contrast Measure
abstract
Detecting small targets in infrared (IR) image sequences is an important task in IR guidance systems. The clutter of complex backgrounds often submerges small targets, making detection difficult. Achieving high detection and low false alarm rates with complex backgrounds is a primary problem. We propose an IR small target detection method using our new homogeneity-weighted local contrast measure (HWLCM). Inspired by the ability of the human visual system (HVS) to determine saliency characteristics, we implement our method to use the local contrast features of the central and surrounding regions and the weighted homogeneity characteristics of the surrounding regions to enhance the target while suppressing the complex background. Our method divides each image into blocks with a sliding window for which the HWLCM is calculated. The HWLCM enhances the actual target and suppresses interference simultaneously. We apply an adaptive threshold to target region extraction to further refine the results. Our experimental results show that our proposed method is more effective than six comparable methods, especially in terms of the signal-to-clutter gain (SCRG) and background suppression factor (BSF) indicators.
Peng Du 0009, Askar Hamdulla
IEEE Geosci. Remote. Sens. Lett.2
2020 Infrared Moving Small-Target Detection Using Spatial-Temporal Local Difference Measure
abstract
Effectiveness and false alarm (FA) suppression are key issues in infrared (IR) moving small-target detection. In this letter, we propose a novel spatial-temporal local difference measure (STLDM) algorithm to detect a moving IR small target. The method we propose involves three steps. First, we block off three frames in a certain range of time (temporal) domain. Next, in a 3-D spatial-temporal domain, we analyze the local grayscale intensity difference of the small targets moving between the three frames and calculate the difference between the grayscale intensity in the center area and the grayscale intensity in the eight directions in the area surrounding the target. Finally, we segment the small target in a detection result map. The results of our experiment demonstrate that our proposed STLDM method has a higher rate of detection of IR moving small targets, as well as fewer FAs than other existing methods.
Peng Du 0009, Askar Hamdulla
IEEE Geosci. Remote. Sens. Lett.2
2014 Lexicon optimization based on discriminative learning for automatic speech recognition of agglutinative language
Mijit Ablimit, Tatsuya Kawahara, Askar Hamdulla
Speech Commun.3
2012 Discriminative approach to lexical entry selection for automatic speech recognition of agglutinative language
abstract
In agglutinative languages, selection of lexical unit is not obvious. Morpheme unit is usually adopted to ensure the sufficient coverage, but many morphemes are short, resulting in weak constraints and possible confusions. In this paper, we propose a discriminative approach to select lexical entries which will directly contribute to ASR error reduction. We define an evaluation function for each word by a set of features and their weights, and the measure for optimization by the difference of WERs by the morpheme-based model and by the word-based model. Then, the weights of the features are learned by a perceptron algorithm. Finally, word (or sub-word) entries with higher evaluation scores are selected to be added to the lexicon. This method is successfully applied to an Uyghur large-vocabulary continuous speech recognition system, resulting in a significant reduction of WER and the lexicon size. Further improvement is achieved by combining with a statistical method based on mutual information criterion.
Mijit Ablimit, Tatsuya Kawahara, Askar Hamdulla
ICASSP3