EDBT 2026 Demo / reviewers in the wild / expert
Wushour Slamu
dblp:129/9382 · also Wushouer Silamu, Wushouer Slamu, Wushour Silamu
· DBLP profile ↗
38ranked-venue papers
1as first author
32since 2021 · last 2026
0000-0003-4592-7806ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 15 since 2021Artificial intelligence and machine learning · 15 · 1 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KG-BiLM: Knowledge Graph Embedding via Bidirectional Language Models
Xin Wang 0030, Zhao Li 0009, Dongxiao He, Yanbing Li, Wushour Slamu |
WWW | 7 |
| 2026 | Dimensional feature-enhanced transformer for low-resource Uyghur scene text recognition
Miaomiao Xu, Yanbing Li, Wushour Slamu |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Dependency-aware graph transformer with hierarchical syntax encoding for aspect-level sentiment classification
Shaokun Liu, Mieradilijiang Maimaiti, Nurmemet Yolwas, Nilufar Abdurakhmonova, Wu Le, Zhuofei Xie, Wushour Slamu |
Expert Syst. Appl. | 8 |
| 2026 | Hierarchical analysis and efficient fine-tuning of weakly supervised speech models
Lixu Sun, Yineng Cai, Nurmemet Yolwas, Wushour Slamu |
Expert Syst. Appl. | 6 |
| 2026 | Text-to-graph query using semantic subgraph retrieval
Yongzhe Jia, Xin Wang 0030, Jianguo Wei, Yurong Qian, Wushour Slamu |
Knowl. Based Syst. | 10 |
| 2025 | Improved Cross-Lingual Speaker Verification Using Speaker Sensitive Feature Guidance and Fine-grained Phonetic InformationabstractSpeaker verification performance significantly degrades when there exists a language mismatch between training and evaluation. Domain Adversarial Training (DAT) has shown to be effective in mitigating this gap by incorporating adversarial training with domain information (language id). Inspired by recent research of DAT, we propose to decompose and locate features sensitive to speaker identity, so that domain adaptation can better serve the speaker classification task. Additionally, we further promote DAT by replacing alignment based on language identification with alignment of fine-grained phonetic information, where pre-trained speech recognition models are utilized to provide frame-level phonetic labels. The speaker verification model is trained on VoxCeleb, while CnCeleb is used for adversarial training and evaluation. Results show the methods effectively mitigate the performance degradation caused by language mismatch. Yongtai Ji, Guangxing Li, Hao Huang 0009, Yanbing Li, Wushour Slamu |
ICASSP | 5 |
| 2025 | Robust and Efficient Text-based Speech Editing using Noise Conditioning and Rectified FlowabstractSignificant advancements have been made in text-based speech editing (TSE) for clear speech, but effectively editing the noise-contaminated speech remains a challenge. Background noise degrades the quality of generated speech, and edited speech that fails to maintain noise context consistency often sounds unnatural. We propose Reflow-TSE, a robust and efficient TSE model for noise-consistent speech editing. For noise robustness, 1) we design a noise condition module to extract frame-level noise sequences as the conditioning information, and 2) we introduce an enhanced context-conditioned prediction module that predicts masked noise sequences along with conventional duration and pitch, using these conditions to guide generation. For boosting efficiency, 3) we introduce the rectified flow model, leveraging speech context and predicted conditions to achieve high-quality editing with limited sampling steps. Experimental results show that with just two steps of sampling, Reflow-TSE achieves context-consistent noisy speech editing, a capability absent in other TSE models. Additionally, for clean speech, Reflow-TSE also matches or surpasses baseline models. Audio samples are available at https://hw-su.github.io. Haowen Yin, Hao Huang 0009, Wushour Slamu |
ICASSP | 5 |
| 2025 | RAST: Residual-Attentive and Scale-Aware Transformer for Robust Scene Text Recognition
Yongbin Mu, Miaomiao Xu, Mieradilijiang Maimaiti, Yanbing Li, Wushour Slamu |
PRCV (7) | 6 |
| 2025 | Visual-Semantic Dual-Decoder Collaboration for Scene Text Recognition
Chuanlong Liu, Shuangying Li, Miaomiao Xu, Mieradilijiang Maimaiti, Wushour Slamu |
PRCV (7) | 5 |
| 2025 | MixFormer: A Cross-Modal Transformer for Arbitrary-Shaped Scene Text Detection
Yaolin Weng, Chuanlong Liu, Miaomiao Xu, Mieradilijiang Maimaiti, Wushour Slamu |
PRCV (7) | 5 |
| 2025 | DSDGT: Dual-stage Dependency Enhanced Graph Transformer for Aspect-Based Sentiment AnalysisabstractAspect-Based Sentiment Analysis (ABSA) seeks to determine the sentiment polarity of specific aspects within a text. Despite the strong performance of Graph Neural Networks (GNNs) based on dependency syntax trees in ABSA, existing methods often fail to differentiate the importance of dependency relations and inadequately capture semantic interactions, limiting their ability to detect implicit sentiments. To address these limitations, we propose Dual-stage Dependency Enhanced Graph Transformer (DSDGT), a serial architecture that combines Transformer and Graph Convolutional Network (GCN) modules with an enhanced dependency mechanism. The Transformer module captures global semantic information, while the GCN models local dependencies. A dual-stage feature processing strategy is employed to integrate both original and enhanced dependency features effectively. Experiments on multiple benchmark datasets demonstrate that DSDGT achieves superior performance in terms of accuracy and F1 scores compared to state-of-the-art methods. Shaokun Liu, Wushour Slamu, Yanbing Li |
SMC | 2 |
| 2025 | Feature enhanced attention decoder for scene text recognition
Miaomiao Xu, Lianghui Xu, Wushour Slamu, Yanbing Li |
Multim. Tools Appl. | 4 |
| 2025 | A classifier expansion framework with dual knowledge distillation and dynamic weighting for continual relation extraction
Aonan Mao, Di Wu 0088, Liting Jiang, Shuangyong Song, Yanbing Li, Hao Huang 0009, Wushour Slamu |
J. Supercomput. | 7 |
| 2024 | Fact-Aware Summarization with Contrastive Learning for Few-Shot Dialogue State TrackingabstractDialogue state tracking (DST) is a crucial component of task-oriented dialogue systems, as it aims to accurately track the user’s goals throughout the dialogue history. However, DST models struggle with new domains due to limited annotated data, leading to poor performance. To solve this key challenge in DST, we propose a model called Fact-aware Summarization model for few-shot DST (FaS-DST), which introduces a "Summarize, Extract, and Select" pattern. Specifically, we decompose DST into three sub-tasks: generating candidate summaries, extracting dialogue states, and scoring the candidates to select the most accurate one. Contrastive learning is incorporated to train a candidate scorer, which improves faithfulness and factuality in dialogue summarization. Additionally, we employ two strategies namely data augmentation and summary & state concatenation to improve the model’s training effectiveness. Experimental results demonstrate that FaS-DST outperforms state-of-the-art models on both MultiWOZ 2.0 and MultiWOZ 2.1 datasets in few-shot settings. Sijie Feng, Haoxiang Su, Hongyan Xie, Di Wu 0088, Hao Huang 0009, Wushour Slamu |
ICASSP | 6 |
| 2024 | Domain-Slot Aware Contrastive Learning for Improved Dialogue State TrackingabstractLarge-scale pre-trained neural language model has facilitated to achieve the state-of-the-art performance on Dialogue State Tracking (DST) tasks. One of the existing works models the semantic correlation between the dialogue context and (domain, slot) pair encoded by BERT and make the prediction. Despite the effectiveness, they ignore the fact that there is no perfect semantic correspondence between (domain, slot) pair and the dialogue context. In this paper, we propose a domain-slot aware contrastive learning framework to solve this problem, which proposes three methods to bridge the semantic gap between the dialogue context and the (domain, slot) by constructing training sample pairs to fine-tune the BERT model and use it for base DST model. The experiments demonstrate that our proposed method has improved the performance of the baseline model on the MultiWOZ2.1 and MultiWOZ2.4 datasets, yielding competitive results. Haoxiang Su, Sijie Feng, Hongyan Xie, Di Wu 0088, Hao Huang 0009, Zhongjiang He, Shuangyong Song, Ruiyu Fang, Xiaomeng Huang, Wushour Slamu |
ICASSP | 10 |
| 2024 | Context-Sensitive Adapter: Contextual Biasing for Personalized End-to-End Speech Recognition with Attention Fusion and Bias Filtering
Yineng Cai, Lixu Sun, Nurmemet Yolwas, Wushour Slamu |
ICIC (5) | 5 |
| 2024 | An Uyghur Extension to the MASSIVE Multi-lingual Spoken Language Understanding Corpus with Comprehensive Evaluations
Ainikaerjiang Aimaiti, Di Wu 0088, Liting Jiang, Gulinigeer Abudouwaili, Hao Huang 0009, Wushour Slamu |
INTERSPEECH | 6 |
| 2024 | Research on Two-Stage Text Language Identification Algorithm for Chinese, Japanese, and Korean
Mamtimin Qasim, Wushour Slamu, Minghui Qiu |
PDCAT | 2 |
| 2024 | Dual Feature Enhanced Scene Text Recognition Method for Low-Resource Uyghur
Miaomiao Xu, Lianghui Xu, Yanbing Li, Wushour Slamu |
PRCV (7) | 5 |
| 2024 | Hybrid Encoding Method for Scene Text Recognition in Low-Resource Uyghur
Miaomiao Xu, Lianghui Xu, Yanbing Li, Wushour Slamu |
PRCV (7) | 5 |
| 2024 | Pedestrian Trajectory Prediction Using Spatio-Temporal VAE
Qing Yu 0010, Zhenwei Xu, Yaoyong Zhou, Zhida Liu, Wushour Slamu |
PRCV (4) | 5 |
| 2024 | Partial channel pooling attention beats convolutional attention
Jun Zhang 0113, Wushour Slamu |
Expert Syst. Appl. | 2 |
| 2024 | Knowledge-aware image understanding with multi-level visual representation enhancement for visual question answering
Zhe Li 0030, Wushour Slamu, Yanbing Li |
Mach. Learn. | 3 |
| 2024 | OECA-Net: A co-attention network for visual question answering based on OCR scene text feature enhancement
Wushour Slamu, Yachuang Chai, Yanbing Li |
Multim. Tools Appl. | 2 |
| 2024 | TADS: a novel dataset for road traffic accident detection from a surveillance perspective
Yachuang Chai, Jianwu Fang, Haoquan Liang, Wushour Slamu |
J. Supercomput. | 4 |
| 2023 | Empathy Intent Drives Empathy DetectionabstractEmpathy plays an important role in the human dialogue.Detecting the empathetic direction expressed by the user is necessary for empathetic dialogue systems because it is highly relevant to understanding the user's needs.Several studies have shown that empathy intent information improves the ability to response capacity of empathetic dialogue.However, the interaction between empathy detection and empathy intent recognition has not been explored.To this end, we invite 3 experts to manually annotate the healthy empathy detection datasets IEMPATHIZE and TwittEmp with 8 empathy intent labels, and perform joint training for the two tasks.Empirical study has shown that the introduction of empathy intent recognition task can improve the accuracy of empathy detection task, and we analyze possible reasons for this improvement.To make joint training of the two tasks more challenging, we propose a novel framework, Cascaded Label Signal Network, which uses the cascaded interactive attention module and the label signal enhancement module to capture feature exchange information between empathy and empathy intent representations.Experimental results show that our framework outperforms all baselines under both settings on the two datasets.1 Liting Jiang, Di Wu 0088, Bohui Mao, Yanbing Li, Wushour Slamu |
EMNLP | 5 |
| 2023 | Unsupervised word Segmentation Based on Word InfluenceabstractWord segmentation task is the cornerstone of text processing. There are 7111 languages worldwide, most of which are low-resource languages. This paper attempts to solve the problem of multilingual unsupervised word segmentation using common points between languages without tagged corpus. We find that words are only a relationship between phrases and non-phrases in each language, and the frequency of their occurrence obeys the normal distribution. Based on the objective law of language and pre-training language model, this paper defines the concept of Word Influence and designs its calculation formula, and loss function. Combined with the fine-tuning word segmentation task, a multilingual unsupervised word segmentation model was proposed. In order to apply to multiple languages, the model’s key parameters can be learned independently. Its validity and advancement have been proved on Chinese, Japanese, and English data sets. Finally, we discuss the challenges of word segmentation in the pre-trained language model environment. Ruohao Yan, Huaping Zhang, Wushour Slamu, Askar Hamdulla |
ICASSP | 3 |
| 2023 | S-CGRU: An Efficient Model for Pedestrian Trajectory Prediction
Zhenwei Xu, Qing Yu 0010, Wushour Slamu, Yaoyong Zhou, Zhida Liu |
ICONIP (10) | 3 |
| 2023 | Emp-USIR: A Unidirectional Synchronous Interactive Reasoning Model for Empathetic DialogueabstractThe goal of an empathetic dialogue generative sys-tem is to generate coherently and relate emotional responses after finite turns of dialogue. Most of the previous work guided the dialogue system to generate empathetic responses from the dialogue history or emotional reasons. However, in people's daily communication, the empathetic response is often generated by deep reasoning through dialogue history and existing commonsense knowledge. To fill this gap, we propose an empathetic dialogue model of unidirectional synchronous interactive reasoning, Emp- Usir.Specifically, according to the existing commonsense knowledge base, we extract commonsense knowledge with emotional representation from it, and use it as the connection feature of dialogue history and remaining common-sense knowledge to guide the interactive reasoning between multi-turns dialogue history and commonsense knowledge, to imitate human reasoning based on commonsense to generate responses. Then, we propose a cross-token level attention mechanism to learn emotional dependence from reasoning features, to produce more smooth and diverse empathy responses. The experimental results and manual evaluation prove the effectiveness of the Emp-USIR model proposed in this paper on the widely-used benchmark dataset. Liting Jiang, Di Wu 0088, Yanbing Li, Wushour Slamu |
IJCNN | 4 |
| 2022 | Automatic Academic Paper Rating Based on Modularized Hierarchical Attention Network
Huaping Zhang, Yugang Li, Wushour Slamu |
NLPCC (1) | 5 |
| 2022 | An Enhanced New Word Identification Approach Using Bilingual Alignment
Huaping Zhang, Jianyun Shang, Wushour Slamu |
NLPCC (1) | 4 |
| 2022 | SPCA-Net: a based on spatial position relationship co-attention network for visual question answering
Wushour Slamu, Yachuang Chai |
Vis. Comput. | 2 |
| 2020 | A Lightweight Model Based on Separable Convolution for Speech Emotion Recognition
Ying Hu 0005, Hao Huang 0009, Wushour Slamu |
INTERSPEECH | 4 |
| 2020 | A Case Representation and Similarity Measurement Model with Experience-Grounded SemanticsabstractCase-based reasoning heavily depends on the structure and content of the cases, and semantics is essential to effectively represent cases. In the field of structured case representation, most of the works regarding case representation and measurement of semantic similarity between cases are based on model-theoretic semantics and their extensions. The purpose of this study is to explore the potential of experienced-grounded semantics in case representation and semantic similarity measurement. The main contents in this study are as follows: (i) a case representation model based on experience-grounded semantic is proposed, (ii) a novel semantic similarity measurement method with multi-strategy reasoning is introduced, and (iii) a case-based reasoning software for urban firefighting field based on the proposed model is designed and implemented. Theoretically, compared with traditional structured case representation methods, the proposed model not only represents case in a fully formalized way, but also provides a novel metric for computing the strength of the semantic relationship between cases. The proposed model has been applied in an intelligent decision-support software for urban firefighting. Nady Slam, Wushour Slamu |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2019 | A keyword-based combination approach for detecting phishing webpages
Yan Ding 0004, Nurbol Luktarhan, Keqin Li 0001, Wushour Slamu |
Comput. Secur. | 4 |
| 2015 | Maximum F1-Score Discriminative Training Criterion for Automatic Mispronunciation DetectionabstractWe carry out an in-depth investigation on a newly proposed Maximum F1-score Criterion (MFC) discriminative training objective function for Goodness of Pronunciation (GOP) based automatic mispronunciation detection that makes use of Gaussian Mixture Model-hidden Markov model (GMM-HMM) as acoustic models. The formulation of MFC seeks to directly optimize F1-score by converting the non-differentiable F1-score function into a continuous objective function to facilitate optimization. We present model-space training algorithm according to MFC using extended Baum–Welch form like update equations based on the weak-sense auxiliary function method. We then present MFC based feature-space discriminative training. We train a matrix projecting from posteriors of Gaussians to a normal size feature space, and add the projected features to traditional spectral features. Mispronunciation detection experiments show MFC based model-space training and feature-space training are effective in improving F1-score and other commonly used evaluation metrics. It is also shown MFC training in both the feature-space and model-space outperforms either model-space training or feature-space training alone, and is about 11.6% better than the maximum likelihood (ML) trained baseline in terms of F1-score. Further, we review and compare mispronunciation detection results with the use of MFC and some traditional training criteria that minimize word error rate in speech recognition. The experimental analysis and comparison provide useful insight into the correlations between F1-score maximization and optimization of these training criteria. Hao Huang 0009, Haihua Xu 0001, Wushour Slamu |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2009 | An E-learner's Emotion Model of Text Using: I. Fundamental Issues for a DDE ModelabstractAs technology advanced, e-learning becomes increasingly important and popular in our society. Determining how to care about the e-learnerpsilas emotional state is a vital issue. Our research concerns offering a better service to e-learners and making their experience in online learning and communication enjoyable. In this paper, which is based on literature search, we find the main emotional categories; then bring forward a novel discrete-dimensions duality emotion (DDE) model for e-learners and present the 3 characteristics of it. Finally, through statistic emotional words from the 2000 topics of BBS, we prove that the DDE model is reasonable. All those are to serve as the basis for our subsequent studies. Xinyan Jia, Wushour Slamu, Feng Tian 0002, Ruomeng Zhao |
ICALT | 2 |
| 2006 | The research on Uighur speaker-dependent isolated word speech recognition
Wushour Slamu, Nuominghua Caiqin |
PACLIC | 1 |