EDBT 2026 Demo / reviewers in the wild / expert
Qingwei Zhao
dblp:73/5890
· DBLP profile ↗
25ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0001-9272-2614ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 15 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multilevel contextual prompting for conversational ASR: unifying conversation history and hotwords with speech LLM
Gaofeng Cheng, Xuyang Wang 0002, Qingwei Zhao, Yonghong Yan 0002 |
Speech Commun. | 4 |
| 2024 | Snore Sound Features Based on Percussive Enhancing and Positional Encoding Combined with Multi-Task Learning for Osahs DetectionabstractObstructive sleep apnea hypopnea syndrome (OSAHS) is a serious sleep disorder. As the typical symptom of OSAHS, snoring has been proved effective in OSAHS diagnosis and potential to replace the current laborious and expensive polysomnography. However, the lack of analysis on the characteristics of pathological snoring sounds limits the diagnosing performance. In this paper, we propose novel sound features for the classification of OSA, hypopnea and normal snores. The proposed features are based on percussive enhancing and positional encoding as the snores exhibit different percussive properties and temporal traits due to the disease generation mechanisms. To enhance the classification performance, we propose a multi-task learning framework to aid the main classification task by simultaneous learning of two related simple tasks. Experiments on real-recorded snoring sounds show that the proposed methods can greatly improve the classification AUC and ACC and the proposed system performs better than those in other literatures. Aolin Hu, Xueshuai Zhang, Shaoxing Zhang, Pengyuan Zhang, Pengfei Ye, Qingwei Zhao, Yonghong Yan 0002 |
ICASSP | 7 |
| 2024 | One-Epoch Training with Single Test Sample in Test Time for Better Generalization of Cough-Based Covid-19 Detection ModelabstractThe outbreak of COVID-19 has raised researchers’ attention to audio-based rapid disease detection. Most of the previous studies have obtained competitive detection performance. However, these results are usually obtained by testing data from the same source offline. When making cross-dataset testing, the performance may deteriorate dramatically due to inconsistent data distribution between different datasets. In addition, in practical application, the model has to make a prediction for the current test audio without any prior information, which requires good model generalization under limited training data. To address the above issues, we adopt a test-time training framework to achieve a cough-based COVID-19 detection model with better generalizability. In the model development stage, resnet18 serves as the backbone network and a self-supervised learning branch is added as an auxiliary task. In testing stage, the model parameters are first fine-tuned by the self-supervised branch with the single test audio as input, and then the classification head outputs predictions. The proposed method is validated on three open-source datasets using a variety of hyperparameters. In cross-dataset testing, AUC and UAR increase by 3.65% and 3% on average absolutely, respectively. The results show that the proposed framework is applicable to improve the model performance in practical application. Jiakun Shen, Xueshuai Zhang, Pengyuan Zhang, Yonghong Yan 0002, Qingwei Zhao, Ta Li, Yanfen Tang, Shaoxing Zhang |
ICASSP | 5 |
| 2024 | ROUGE-SEM: Better evaluation of summarization using ROUGE combined with semantics
Ming Zhang 0035, Chengzhang Li, Meilin Wan, Xuejun Zhang 0002, Qingwei Zhao |
Expert Syst. Appl. | 5 |
| 2024 | Factorized and progressive knowledge distillation for CTC-based ASR models
Sanli Tian, Zehan Li, Zhaobiao Lyv, Gaofeng Cheng, Ta Li, Qingwei Zhao |
Speech Commun. | 7 |
| 2024 | ASQ: An Ultra-Low Bit Rate ASR-Oriented Speech Quantization MethodabstractFor efficient transmission of speech signals, speech compression methodologies have attracted significant research attention for decades and are widely used in automatic speech recognition (ASR) services. However, most speech codecs are perception-oriented, leaving redundant information and introducing distortion, which harms ASR systems. Recently, the emergence of neural network-based models has significantly advanced the progress of ASR systems and speech coding, laying the foundation for building a speech compression method specially optimized for ASR systems. In this letter, we propose an ASR-oriented Speech Quantization (ASQ) method to reduce communication costs for speech recognition systems. In the proposed method, a speech quantization model first converts the speech into low bit rate tokens. Then the tokens are transmitted to the server and recognized by a quantized speech recognition model. The two models could be jointly trained in the end-to-end (E2E) style. To mitigate the performance degradation introduced by the quantization components, we design an entropy-guided 3-stage training method that encourages the model to fully utilize the token space and promote recognition accuracy. Experiment results on the LibriSpeech corpus show that compared to an existing non-quantized ASR model with a 256 kbps transmission bit rate, the proposed method can achieve a transmission bit rate of 0.6 kbps without any influence on word error rate (WER). It also significantly surpasses the 2-step pipeline that first performs speech codec and then recognizes with a several times lower bit rate. Lingxuan Ye, Changfeng Gao, Gaofeng Cheng, Liuping Luo, Qingwei Zhao |
IEEE Signal Process. Lett. | 5 |
| 2024 | Unsupervised Domain Adaptation on End-to-End Multi-Talker Overlapped Speech RecognitionabstractSerialized Output Training (SOT) has emerged as the mainstream approach for addressing the multi-talker overlapped speech recognition challenge due to its simplicity. However, SOT encounters cross-domain performance degradation which hinders its application. Meanwhile, traditional domain adaption methods may harm the accuracy of speaker change point prediction evaluated by UD-CER, which is an important metric in SOT. To solve these issues, we propose Pseudo-Labeling based SOT (PL-SOT) for domain adaptation by treating speaker change token ($< $sc$>$) specially during training to increase the accuracy of speaker change point prediction. Firstly, we improve CTC loss by proposingWeakening and Enhancing CTC(WE-CTC) loss to weaken the learning of error-prone labels surrounding$<$sc$>$while enhance the emission probability of$< $sc$>$through modifying posteriors of the pseudo-labels. Secondly, we introduceWeighted Confidence Filter(WCF) that assigns higher scores of$<$sc$>$to exclude low-quality pseudo-labels without hurting the$< $sc$>$prediction. Experimental results show that PL-SOT achieves 17.7%/12.8% average relative reduction of CER/UD-CER, with AliMeeting as source domain and AISHELL-4 along with MagicData-RAMC as target domain. Han Zhu 0004, Sanli Tian, Qingwei Zhao, Ta Li |
IEEE Signal Process. Lett. | 4 |
| 2023 | Reminding the incremental language model via data-free self-distillation
Han Wang 0033, Ruiliu Fu, Chengzhang Li, Xuejun Zhang 0002, Jun Zhou 0024, Xing Bai, Yonghong Yan 0002, Qingwei Zhao |
Appl. Intell. | 8 |
| 2023 | So-DAS: A Two-Step Soft-Direction-Aware Speech Separation FrameworkabstractMost existing direction-aware speech separation systems lead to performance degradation when the angle difference between speakers is small due to the low spatial discrimination. To address this issue, we propose a two-step soft-direction-aware speech separation (So-DAS) framework, which consists of a direction of arrival (DOA) estimation module and a speech separation module. First, the two modules are individually optimized, and directional features (DFs) derived from ground-truth DOAs are utilized as spatial information to facilitate the separation module. Next, the two modules are cascaded and optimized with only separation loss, and the DFs are generated using the estimator outputs. By this means, the consistency between the two modules is strengthened, and thus spatial cues that are more beneficial to the separation task can be exploited by the network itself. The experimental results show that compared to the baselines, DFs extracted by our proposed method provides clearer superiority, especially when the angle difference between speakers is small. In addition, our approach yields a state-of-the-art word error rate of 3.4% on the real-recorded utterance-wise LibriCSS dataset. Yi Yang 0057, Qingwei Zhao, Pengyuan Zhang |
IEEE Signal Process. Lett. | 3 |
| 2022 | Ask Question First for Enhancing Lifelong Language LearningabstractLifelong language learning aims to stream learning NLP tasks while retaining knowledge of previous tasks. Previous works based on the language model and following data-free constraint approaches have explored formatting all data as “begin token (B) + context (C) + question (Q) + answer (A)” for different tasks. However, they still suffer from catastrophic forgetting and are exacerbated when the previous task’s pseudo data is insufficient for the following reasons: (1) The model has difficulty generating task-corresponding pseudo data, and (2) A is prone to error when A and C are separated by Q because the information of the C is diminished before generating A. Therefore, we propose the Ask Question First and Replay Question (AQF-RQ), including a novel data format “BQCA” and a new training task to train pseudo questions of previous tasks. Experimental results demonstrate that AQF-RQ makes it easier for the model to generate more pseudo data that match corresponding tasks, and is more robust to both sufficient and insufficient pseudo-data when the task boundary is both clear and unclear. AQF-RQ can achieve only 0.36% lower performance than multi-task learning. Han Wang 0033, Ruiliu Fu, Xuejun Zhang 0002, Jun Zhou 0024, Qingwei Zhao |
COLING | 5 |
| 2011 | Towards precise and robust automatic synchronization of live speech and its transcripts
Jie Gao 0020, Qingwei Zhao, Yonghong Yan 0002 |
Speech Commun. | 2 |
| 2010 | Automatic Synchronization of live speech and its Transcripts based on a frame-synchronous likelihood ratio testabstractIn this paper, we present our initial efforts in the task of Automatically Synchronizing live spoken Utterances with their Transcripts (textual contents) (ASUT) when the texts are known. We treat it as a online speech-text alignment problem. And it is further simplified into the problem of on-the-fly detecting of the end time of a spoken utterance given its textual content. A general framework called frame-synchronous likelihood ratio test (FS-LRT) procedure is proposed for this end time detection task and explored with the hidden Markov models (HMMs). The property of FS-LRT is studied empirically. Extensive experiments indicate that our proposed approach shows satisfying performance. In addition, FS-LRT has been successfully applied in a subtitling system for live broadcast news. Jie Gao 0020, Qingwei Zhao, Yonghong Yan 0002 |
ICASSP | 2 |
| 2009 | Online detecting end times of spoken utterances for synchronization of live speech and its transcripts
Jie Gao 0020, Qingwei Zhao, Yonghong Yan 0002 |
INTERSPEECH | 2 |
| 2009 | Simultaneous Synchronization of Text and Speech for Broadcast News Subtitling
Jie Gao 0020, Qingwei Zhao, Ta Li, Yonghong Yan 0002 |
ISNN (3) | 2 |
| 2009 | A Novel Fuzzy-Based Automatic Speaker Clustering Algorithm
Xiang Zhang 0014, Hongbin Suo, Qingwei Zhao, Yonghong Yan 0002 |
ISNN (2) | 4 |
| 2008 | Mandarin vowel pronunciation quality evaluation by a novel formant classification method and its combination with traditional algorithmsabstractThis paper discusses the vowel pronunciation quality assessment of our computer assisted Mandarin Chinese learning system. Under the speech recognition framework, phonetic pronunciation assessment is usually based on the phonetic posterior probability score, which may be computed by normalizing the frame-based posterior probability or be calculated on the phone segment directly. By the first method, we can achieve a human-machine scoring correlation coefficient (CC) of 0.832 for vowel; and by the second, the CC can be up to 0.847. In order to improve the performance, we suggest employing the formant feature of vowel. This paper proposes a novel method to utilize formant: we plot formant candidates of each frame on the time-frequency plane to form a bitmap, and then extract its Gabor feature for pattern classification. When we use the classification probability score for pronunciation assessment, we get a CC of 0.842. Finally we combine the three scores with various linear or nonlinear methods; the best CC of 0.913 is gotten by using neural network. Fuping Pan, Qingwei Zhao, Yonghong Yan 0002 |
ICASSP | 2 |
| 2008 | Robust speaker change detection using Kernel-Gaussian model
Jie Gao 0020, Xiang Zhang 0014, Qingwei Zhao, Yonghong Yan 0002 |
INTERSPEECH | 3 |
| 2008 | Towards vocabulary-independent speech indexing for large-scale repositories
Jian Shao 0001, Roger Peng Yu, Qingwei Zhao, Yonghong Yan 0002, Frank Seide |
INTERSPEECH | 3 |
| 2007 | Mandarin vowel pronunciation quality evaluation by using formant pattern recognition
Fuping Pan, Qingwei Zhao, Yonghong Yan 0002 |
INTERSPEECH | 2 |
| 2007 | A fast fuzzy keyword spotting algorithm based on syllable confusion network
Jian Shao 0001, Qingwei Zhao, Pengyuan Zhang, Zhaojie Liu, Yonghong Yan 0002 |
INTERSPEECH | 2 |
| 2005 | Fast confidence measure algorithm for continuous speech recognition
Bin Dong 0003, Qingwei Zhao, Yonghong Yan 0002 |
INTERSPEECH | 2 |
| 2000 | Keyword spotting in auto-attendant systemabstractIn this paper, an auto-attendant system using finite state grammar (FSG) based on a continuous speech recognition (CSR) model is introduced. However, by using two virtual garbage models, one is to match the leading extraneous speech before the key name and the other to match the tailing extraneous speech following the key name, we managed to reach a more flexible and robust auto-attendant system. The experiment result show that, in our auto attendant system (about 240 names), to the name only test set and the sentence test set 1 composed of sentences that FSG can recognize, the recognition rate of the keyword spotting system is almost the same as that of FSG. To the sentence test set 2 composed of sentences that undefined in the FSG the keyword spotting system outperforms the FSG system remarkably. Not affecting the recognition accuracy of name only test set and the sentence test set 1, task dependent keyword models cut off additional 20% of error rate comparing with task independent keyword models in the sentence test set 2. Yonghong Yan 0002, Baosheng Yuan, Qingwei Zhao |
INTERSPEECH | 5 |
| 2000 | Optimal maximum likelihood on phonetic decision tree acoustic model for LVCSR
Baosheng Yuan, Qingwei Zhao |
INTERSPEECH | 2 |
| 2000 | Improvements in search algorithm for large vocabulary continuous speech recognition
Qingwei Zhao, Baosheng Yuan, Yonghong Yan 0002 |
INTERSPEECH | 1 |
| 1999 | A study of duration in continuous speech recognition based on DDBHMMabstractAtros is an automatic speech recognition/understanding/translation system whose knowledge sources (acoustic models, lexical models, syntactic language models, semantic models and translation models) can be learnt automatically from training data by using similar techniques. The search process in Atros is performed through a Synchronous Beam Search technique. In this paper, a faster version of Atros is presented and evaluated. This version supports improved acoustic and syntactical models. It also incorporates improved search algorithms to reduce and the computational requirements for decoding: Fast Phoneme Look-Ahead and Histogram Pruning. The system has been tested on a Spanish task of queries to a geographical database (with a vocabulary of 1,264 words). The best result achieved (in real time) was 7.10% of word error rate. Qingwei Zhao, Zuoying Wang, Dajin Lu |
EUROSPEECH | 1 |