Weiqiang Zhang 0001

dblp:72/6674-1 · also Wei-Qiang Zhang 0001 · DBLP profile ↗
← Back
65ranked-venue papers
6as first author
25since 2021 · last 2026
0000-0003-3841-1959ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 52 · 5 first-author · 16 since 2021Artificial intelligence and machine learning · 39 · 3 first-author · 18 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 MetaAug: Task augmentation via entropy increase for robust meta-learning
Chaolong Hao, Hao Zhang 0109, Dan Qu 0003, Weiqiang Zhang 0001
Knowl. Based Syst.5
2026 Gradient-aware knowledge distillation: Tackling gradient insensitivity through teacher guided gradient scaling
Nianwen Si, Hao Zhang 0109, Weiqiang Zhang 0001, Heyu Chang, Dan Qu 0003
Neural Networks4
2025 GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement
abstract
Yifan Yang, Zheshu Song, Jianheng Zhuo, Mingyu Cui, Jinpeng Li, Bo Yang, Yexing Du, Ziyang Ma, Xunying Liu, Ziyuan Wang, Ke Li, Shuai Fan, Kai Yu, Wei-Qiang Zhang, Guoguo Chen, Xie Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yifan Yang 0005, Zheshu Song, Jianheng Zhuo, Bo Yang 0006, Yexing Du, Ziyang Ma 0001, Xunying Liu, Ke Li 0018, Shuai Fan 0005, Kai Yu 0004, Weiqiang Zhang 0001, Guoguo Chen, Xie Chen 0001
ACL (1)14
2025 Data-Efficient Low-Complexity Acoustic Scene Classification via Distilling and Progressive Pruning
abstract
The goal of the acoustic scene classification (ASC) task is to classify recordings into one of the predefined acoustic scene classes. However, in real-world scenarios, ASC systems often encounter challenges such as recording device mismatch, low-complexity constraints, and the limited availability of labeled data. To alleviate these issues, in this paper, a data-efficient and low-complexity ASC system is built with a new model architecture and better training strategies. Specifically, we firstly design a new low-complexity architecture named Rep-Mobile by integrating multi-convolution branches which can be reparameterized at inference. Compared to other models, it achieves better performance and less computational complexity. Then we apply the knowledge distillation strategy and provide a comparison of the data efficiency of the teacher model with different architectures. Finally, we propose a progressive pruning strategy, which involves pruning the model multiple times in small amounts, resulting in better performance compared to a single step pruning. Experiments are conducted on the TAU dataset. With Rep-Mobile and these training strategies, our proposed ASC system achieves the state-of-the-art (SOTA) results so far, while also winning the first place with a significant advantage over others in the DCASE2024 Challenge.
Bing Han 0008, Wen Huang 0004, Zhengyang Chen, Anbai Jiang, Pingyi Fan, Cheng Lu 0007, Zhiqiang Lv, Jia Liu 0001, Weiqiang Zhang 0001, Yanmin Qian
ICASSP9
2025 Empowering Large Language Models for End-to-End Speech Translation Leveraging Synthetic Data
Yu Pu, Weiqiang Zhang 0001, Xie Chen 0001
INTERSPEECH5
2025 MPN: Leveraging Multilingual Patch Neuron for Cross-Lingual Model Editing
Nianwen Si, Heyu Chang, Weiqiang Zhang 0001
KSEM (2)3
2025 SpeechColab leaderboard: An open-source platform for automatic speech recognition evaluation
Jiayu Du, Guoguo Chen, Weiqiang Zhang 0001
Comput. Speech Lang.4
2024 Exploring Large Scale Pre-Trained Models for Robust Machine Anomalous Sound Detection
abstract
Machine anomalous sound detection is a useful technique for various applications, but it often suffers from poor generalization due to the challenges of data collection and complex acoustic environment. To address this issue, we propose a robust machine anomalous sound detection model that leverages self-supervised pre-trained models on large-scale speech data. Specifically, we assign different weights to the features from different layers of the pre-trained model and then use the working condition as the label for self-supervised classification fine-tuning. Moreover, we introduce a data augmentation method that simulates different operating states of the machine to enrich the dataset. Furthermore, we devise a transformer pooling method that fuses the features of different segments. Experiments on the DCASE2023 dataset show that our proposed method outperforms the commonly used reconstruction-based autoencoder and classification-based convolutional network by a large margin, demonstrating the effectiveness of large-scale pre-training for enhancing the generalization and robustness of machine anomalous sound detection. In Task2 of DCASE2023, we achieve 2nd place with these methods.
Bing Han 0008, Zhiqiang Lv, Anbai Jiang, Wen Huang 0004, Zhengyang Chen, Yufeng Deng, Cheng Lu 0007, Weiqiang Zhang 0001, Pingyi Fan, Jia Liu 0001, Yanmin Qian
ICASSP9
2024 Whisper-Based Transfer Learning for Alzheimer Disease Classification: Leveraging Speech Segments with Full Transcripts as Prompts
abstract
Alzheimer’s disease (AD) is a neurodegenerative disorder that can lead to speech impairments. Early diagnosis is crucial for effective treatment, and speech-based diagnosis is currently a hot research topic. In this study, we explore the feasibility of transfer learning for Alzheimer’s disease detection using the state-of-the-art multilingual speech recognition and translation model: Whisper. In order to address the limitation of Whisper’s narrow perspective caused by the restricted audio segment length during fine-tuning, we propose an innovative method to overcome this problem by using the full transcript as a prompt to assist in training speech segments. This approach results in a relative performance improvement of 9%-12% for models with a higher number of parameters. On the ADReSSo test set, the accuracy and F1 score achieved are 84.51% and 84.50% respectively, surpassing both the baseline system and commonly used speech recognition-language model cascade methods, demonstrating its effectiveness.
Weiqiang Zhang 0001
ICASSP2
2024 AnoPatch: Towards Better Consistency in Machine Anomalous Sound Detection
Anbai Jiang, Bing Han 0008, Zhiqiang Lv, Yufeng Deng, Weiqiang Zhang 0001, Xie Chen 0001, Yanmin Qian, Jia Liu 0001, Pingyi Fan
INTERSPEECH5
2024 Improving Anomalous Sound Detection Via Low-Rank Adaptation Fine-Tuning of Pre-Trained Audio Models
abstract
Anomalous Sound Detection (ASD) has gained significant interest through the application of various Artificial Intelligence (AI) technologies in industrial settings. Though possessing great potential, ASD systems can hardly be readily deployed in real production sites due to the generalization problem, which is primarily caused by the difficulty of data collection and the complexity of environmental factors. This paper introduces a robust ASD model that leverages audio pre-trained models. Specifically, we fine-tune these models using machine operation data, employing SpecAug as a data augmentation strategy. Additionally, we investigate the impact of utilizing Low-Rank Adaptation (LoRA) tuning instead of full fine-tuning to address the problem of limited data for fine-tuning. Our experiments on the DCASE2023 Task 2 dataset establish a new benchmark of 77.75% on the evaluation set, with a significant improvement of 6.48% compared with previous state-of-the-art (SOTA) models, including top-tier traditional convolutional networks and speech pre-trained models, which demonstrates the effectiveness of audio pre-trained models with LoRA tuning. Ablation studies are also conducted to showcase the efficacy of the proposed scheme.
Xinhu Zheng, Anbai Jiang, Bing Han 0008, Yanmin Qian, Pingyi Fan, Jia Liu 0001, Weiqiang Zhang 0001
SLT7
2023 Transferring Speech-Generic and Depression-Specific Knowledge for Alzheimer's Disease Detection
abstract
The detection of Alzheimer’s disease (AD) from spontaneous speech has attracted increasing attention while the sparsity of training data remains an important issue. This paper handles the issue by knowledge transfer, specifically from both speech-generic and depression-specific knowledge. The paper first studies sequential knowledge transfer from generic foundation models pretrained on large amounts of speech and text data. A block-wise analysis is performed for AD diagnosis based on the representations extracted from different intermediate blocks of different foundation models. Apart from the knowledge from speech-generic representations, this paper also proposes to simultaneously transfer the knowledge from a speech depression detection task based on the high comorbidity rates of depression and AD. A parallel knowledge transfer framework is studied that jointly learns the information shared between these two tasks. Experimental results show that the proposed method improves AD and depression detection, and produces a state-of-the-art F1 score of 0.928 for AD diagnosis on the commonly used ADReSSo dataset.
Ziyun Cui, Wen Wu 0007, Weiqiang Zhang 0001, Ji Wu 0002, Chao Zhang 0031
ASRU3
2023 Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
abstract
Self-supervised learning (SSL) has achieved great success in speech processing, but always with a large model size to increase the modeling capacity. This may limit its potential applications due to the expensive computation and memory costs introduced by the oversize model. Compression for SSL models has become an important research direction of practical value. To this end, we explore the effective distillation of HuBERT-based SSL models for automatic speech recognition. First, a comprehensive study of different student model structures is conducted. On top of this, as a supplement to the regression loss widely adopted in previous works, a discriminative loss is introduced for HuBERT to enhance the distillation performance, especially in low-resource scenarios. In addition, we design a simple and effective algorithm to distill the front-end input from waveform to Fbank feature, resulting in 17% parameter reduction and doubling inference speed, at marginal performance degradation.
Changli Tang, Ziyang Ma 0001, Zhisheng Zheng, Xie Chen 0001, Weiqiang Zhang 0001
ASRU6
2023 Decoupling Detectors for Scalable Anomaly Detection in AIoT Systems with Multiple Machines
abstract
The fast-developing Artificial Internet of Things (AIoT) technologies enable the consistent monitoring of multiple machines, by which machine failures can be detected in the early phases, and production efficiency and system management can be greatly promoted, bringing huge significance for anomaly detection. However, in most cases, anomalies are not provided for training, and the lack of direct supervision deprecates the anomaly detection performance. For the application viewpoint, the detector is required to generalize well on multiple machines, except for being computationally efficient. The computational cost is strictly limited, which is a great challenge for mobile and embedded devices. In face of these issues, we propose MobileAnoNet, which decouples an end-to-end detector into a front-end feature extractor and a back-end anomaly detector. The front-end extractor, consuming most computation, is unified for all machine types, while the back-end detector is specialized for each machine type, improving the detection capacity. The model is trained by handy labels of machine types and working conditions, in which multiple classification heads are attached behind the feature extractor during training. The performance of the model is evaluated on two DCASE datasets focusing on machine audio anomaly detection. It's shown that MobileAnoNet achieves a general improvement of 6.9% and 8.8% on two datasets, respectively. The ablation study demonstrates that multi-task learning promotes the general representation capacity. The source code is available at: www.github.com/hqj-les30/MobileAnoNet.
Qijun Hou, Anbai Jiang, Weiqiang Zhang 0001, Pingyi Fan, Jia Liu 0001
GLOBECOM3
2023 Unsupervised Anomaly Detection and Localization of Machine Audio: A Gan-Based Approach
abstract
Automatic detection of machine anomaly remains challenging for machine learning. We believe the capability of generative adversarial network (GAN) suits the need of machine audio anomaly detection, yet rarely has this been investigated by previous work. In this paper, we propose AEGAN-AD, a totally unsupervised approach in which the generator (also an autoencoder) is trained to reconstruct input spectrograms. It is pointed out that the denoising nature of reconstruction deprecates its capacity. Thus, the discriminator is redesigned to aid the generator during both training stage and detection stage. The performance of AEGAN-AD on the dataset of DCASE 2022 Challenge TASK 2 demonstrates the state-of-the-art result on five machine types. A novel anomaly localization method is also investigated. Source code available at: www.github.com/jianganbai/AEGAN-AD
Anbai Jiang, Weiqiang Zhang 0001, Yufeng Deng, Pingyi Fan, Jia Liu 0001
ICASSP2
2023 DistilXLSR: A Light Weight Cross-Lingual Speech Representation Model
Haoyu Wang 0014, Siyuan Wang 0002, Weiqiang Zhang 0001, Jinfeng Bai
INTERSPEECH3
2023 Task-Agnostic Structured Pruning of Speech Representation Models
Haoyu Wang 0014, Siyuan Wang 0002, Weiqiang Zhang 0001, Hongbin Suo, Yulong Wan
INTERSPEECH3
2023 Symmetric Saliency-Based Adversarial Attack to Speaker Identification
abstract
Adversarial attack approaches to speaker identification either need high computational cost or are not very effective, to our knowledge. To address this issue, in this letter, we propose a novel generation-network-based approach, called symmetric saliency-based encoder-decoder (SSED), to generate adversarial voice examples to speaker identification. It contains two novel components. First, it uses a novel saliency map decoder to learn the importance of speech samples to the decision of a targeted speaker identification system, so as to make the attacker focus on generating artificial noise to the important samples. It also proposes an angular loss function to push the speaker embedding far away from the source speaker. Our experimental results demonstrate that the proposed SSED yields the state-of-the-art performance, i.e. over 97% targeted attack success rate and a signal-to-noise level of over 39 dB on both the open-set and close-set speaker identification tasks, with a low computational cost.
Jiadi Yao, Xing Chen 0011, Xiao-Lei Zhang 0001, Weiqiang Zhang 0001, Kunde Yang
IEEE Signal Process. Lett.4
2023 LMD: A Learnable Mask Network to Detect Adversarial Examples for Speaker Verification
abstract
Although the security of automatic speaker verification (ASV) is seriously threatened by recently emerged adversarial attacks, there have been some countermeasures to alleviate the threat. However, many defense approaches not only require the prior knowledge of the attackers but also possess weak interpretability. To address this issue, in this paper, we propose anattacker-independentandinterpretablemethod, namedlearnable mask detector(LMD), to separate adversarial examples from the genuine ones. It utilizes score variation as an indicator to detect adversarial examples, where the score variation is the absolute discrepancy between the ASV scores of an original audio recording and its transformed audio synthesized from its masked complex spectrogram. A core component of the score variation detector is to generate the masked spectrogram by a neural network. The neural network needs only genuine examples for training, which makes it an attacker-independent approach. Its interpretability lies that the neural network is trained to minimize the score variation of the targeted ASV, and maximize the number of the masked spectrogram bins of the genuine training examples. Its foundation is based on the observation that, masking out the vast majority of the spectrogram bins with little speaker information will inevitably introduce a large score variation to the adversarial example, and a small score variation to the genuine example. Experimental results with 12 attackers and two representative ASV systems show that our proposed method outperforms five state-of-the-art baselines. The extensive experimental results can also be a benchmark for the detection-based ASV defenses.
Xing Chen 0011, Xiao-Lei Zhang 0001, Weiqiang Zhang 0001, Kunde Yang
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Improving Speech Translation by Cross-Modal Multi-Grained Contrastive Learning
abstract
The end-to-end speech translation (E2E-ST) model has gradually become a mainstream paradigm due to its low latency and less error propagation. However, it is non-trivial to train such a model well due to the task complexity and data scarcity. The speech-and-text modality differences result in the E2E-ST model performance usually inferior to the corresponding machine translation (MT) model. Based on the above observation, existing methods often use sharing mechanisms to carry outimplicit knowledge transferby imposing various constraints. However, the final model often performs worse on the MT task than the MT model trained alone, which means that the knowledge transfer ability of this method is also limited. To deal with these problems, we propose the FCCL (Fine- andCoarse- GranularityContrastiveLearning) approach for E2E-ST, which makesexplicit knowledge transferthrough cross-modal multi-grained contrastive learning. A key ingredient of our approach is applying contrastive learning at both sentence- and frame-level to give the comprehensive guide for extracting speech representations containing rich semantic information. In addition, we adopt a simple whitening method to alleviate the representation degeneration in the MT model, which adversely affects contrast learning. Experiments on the MuST-C benchmark show that our proposed approach significantly outperforms the state-of-the-art E2E-ST baselines on all eight language pairs. Further analysis indicates that FCCL can free up its capacity from learning grammatical structure information and force more layers to learn semantic information.
Hao Zhang 0109, Nianwen Si, Xukui Yang 0001, Dan Qu 0003, Weiqiang Zhang 0001
IEEE ACM Trans. Audio Speech Lang. Process.7
2022 The THUEE System Description for the IARPA OpenASR21 Challenge
abstract
This paper describes the THUEE team's speech recognition system for the IARPA Open Automatic Speech Recognition Challenge (OpenASR21), with further experiment explorations. We achieve outstanding results under both the Constrained and Constrained-plus training conditions. For the Constrained training condition, we construct our basic ASR system based on the standard hybrid architecture. To alleviate the Out-Of-Vocabulary (OOV) problem, we extend the pronunciation lexicon using Grapheme-to-Phoneme (G2P) techniques for both OOV and potential new words. Standard acoustic model structures such as CNN-TDNN-F and CNN-TDNN-F-A are adopted. In addition, multiple data augmentation techniques are applied. For the Constrained-plus training condition, we use the self-supervised learning framework wav2vec2.0. We experiment with various fine-tuning techniques with the Connectionist Temporal Classification (CTC) criterion on top of the publicly available pre-trained model XLSR-53. We find that the frontend feature extractor plays an important role when applying the wav2vec2.0 pre-trained model to the encoder-decoder based CTC/Attention ASR architecture. Extra improvements can be achieved by using the CTC model finetuned in the target language as the frontend feature extractor.
Haoyu Wang 0014, Shuzhou Chai, Guanbo Wang, Guoguo Chen, Weiqiang Zhang 0001
INTERSPEECH7
2021 Automatic Speech Recognition for Low-Resource Languages: The Thuee Systems for the IARPA Openasr20 Evaluation
abstract
The paper introduces our Automatic Speech Recognition (ASR) systems for the IARPA Open Automatic Speech Recognition Challenge (OpenASR20) as well as some post explorations with speech pre-training. We compete in the Constrained training condition for the 10 languages as team THUEE, under which the only speech data permissible for training of each language is a 10-hour corpus. We adopt the hybrid NN-HMM acoustic model and an N-gram Language Model (LM) to construct our basic ASR systems. The acoustic model is proposed as CNN-TDNNF-A, which combines Convolution Neural Network (CNN), Factored Time Delay Neural Network (TDNN-F) and self-attention mechanism. As for low-resource condition, we apply speed and volume perturbation, SpecAugment and reverberation for data enhancement as well as data clean-up to filter interference information. A series of pre-and-post-processing procedures for the evaluation set, such as Speech Activity Detection (SAD), system fusion and results filtering are carried out to obtain the final results. Furthermore, we exploit wav2vec 2.0 pre-trained model to obtain more effective speech representations for the hybrid system with a series of explorations, which brings about evident improvements.
Gui-Xin Shi, Guan-Bo Wang, Weiqiang Zhang 0001
ASRU4
2021 GigaSpeech: An Evolving, Multi-Domain ASR Corpus with 10, 000 Hours of Transcribed Audio
abstract
This paper introduces GigaSpeech, an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training, and 40,000 hours of total audio suitable for semi-supervised and unsupervised training.Around 40,000 hours of transcribed audio is first collected from audiobooks, podcasts and YouTube, covering both read and spontaneous speaking styles, and a variety of topics, such as arts, science, sports, etc.A new forced alignment and segmentation pipeline is proposed to create sentence segments suitable for speech recognition training, and to filter out segments with low-quality transcription.For system training, GigaSpeech provides five subsets of different sizes, 10h, 250h, 1000h, 2500h, and 10000h.For our 10,000-hour XL training subset, we cap the word error rate at 4% during the filtering/validation stage, and for all our other smaller training subsets, we cap it at 0%.The DEV and TEST evaluation sets, on the other hand, are re-processed by professional human transcribers to ensure high transcription quality.Baseline systems are provided for popular speech recognition toolkits, namely Athena, ESPnet, Kaldi and Pika.
Guoguo Chen, Shuzhou Chai, Guan-Bo Wang, Jiayu Du, Weiqiang Zhang 0001, Chao Weng, Dan Su 0002, Daniel Povey, Jan Trmal, Mingjie Jin, Sanjeev Khudanpur, Shinji Watanabe 0001, Shuaijiang Zhao, Xiangang Li, Xuchen Yao, Zhao You, Zhiyong Yan
Interspeech5
2021 The TNT Team System Descriptions of Cantonese and Mongolian for IARPA OpenASR20
Zhiqiang Lv, Ambyer Han, Guan-Bo Wang, Gui-Xin Shi, Jian Kang 0006, Jinghao Yan, Pengfei Hu 0004, Shen Huang, Weiqiang Zhang 0001
Interspeech10
2021 End-to-end keyword search system based on attention mechanism and energy scorer for low resource languages
Zeyu Zhao 0004, Weiqiang Zhang 0001
Neural Networks2
2020 Staged Training Strategy and Multi-Activation for Audio Tagging with Noisy and Sparse Multi-Label Data
Kexin He, Yuhan Shen, Weiqiang Zhang 0001, Jia Liu 0001
ICASSP3
2020 Dynamic Temporal Residual Learning for Speech Recognition
abstract
Long short-term memory (LSTM) networks have been widely used in automatic speech recognition (ASR). This paper proposes a novel dynamic temporal residual learning mechanism for LSTM networks to better explore temporal dependencies in sequential data. The temporal residual learning mechanism is implemented by applying shortcut connections with dynamic weights to temporally adjacent LSTM outputs. Two types of dynamic weight generation methods are proposed: using a secondary network and using a random weight generator. Experimental results on Wall Street Journal (WSJ) speech recognition dataset reveal that our proposed methods have surpassed the baseline LSTM network.
Jiaqi Xie, Ruijie Yan, Shanyu Xiao, Liangrui Peng, Michael T. Johnson, Weiqiang Zhang 0001
ICASSP6
2020 THUEE System for NIST SRE19 CTS Challenge
Ruyun Li, Tianyu Liang, Yi Liu 0049, Yangcheng Wu, Can Xu 0003, Xianhong Chen, Weiqiang Zhang 0001, Shouyi Yin, Liang He 0003
INTERSPEECH10
2020 End-to-End Keyword Search Based on Attention and Energy Scorer for Low Resource Languages
Zeyu Zhao 0004, Weiqiang Zhang 0001
INTERSPEECH2
2019 Hierarchical Pooling Structure for Weakly Labeled Sound Event Detection
abstract
Sound event detection with weakly labeled data is considered as a problem of multi-instance learning.And the choice of pooling function is the key to solving this problem.In this paper, we proposed a hierarchical pooling structure to improve the performance of weakly labeled sound event detection system.Proposed pooling structure has made remarkable improvements on three types of pooling function without adding any parameters.Moreover, our system has achieved competitive performance on Task 4 of Detection and Classification of Acoustic Scenes and Events (DCASE) 2017 Challenge using hierarchical pooling structure.
Kexin He, Yuhan Shen, Weiqiang Zhang 0001
INTERSPEECH3
2019 Learning How to Listen: A Temporal-Frequential Attention Model for Sound Event Detection
abstract
In this paper, we propose a temporal-frequential attention model for sound event detection (SED). Our network learns how to listen with two attention models: a temporal attention model and a frequential attention model. Proposed system learns when to listen using the temporal attention model while it learns where to listen on the frequency axis using the frequential attention model. With these two models, we attempt to make our system pay more attention to important frames or segments and important frequency components for sound event detection. Our proposed method is demonstrated on the task 2 of Detection and Classification of Acoustic Scenes and Events (DCASE) 2017 Challenge and achieves competitive performance.
Yuhan Shen, Kexin He, Weiqiang Zhang 0001
INTERSPEECH3
2019 Music Genre Classification Using Duplicated Convolutional Layers in Neural Networks
Hansi Yang, Weiqiang Zhang 0001
INTERSPEECH2
2018 Argument division based branch-and-bound algorithm for unit-modulus constrained complex quadratic programming
Cheng Lu 0007, Zhibin Deng, Weiqiang Zhang 0001, Shu-Cherng Fang
J. Glob. Optim.3
2018 Semi-supervised minimum redundancy maximum relevance feature selection for audio classification
Xukui Yang 0001, Liang He 0003, Dan Qu 0003, Weiqiang Zhang 0001
Multim. Tools Appl.4
2017 Gated convolutional networks based hybrid acoustic models for low resource speech recognition
abstract
In acoustic modeling for large vocabulary speech recognition, recurrent neural networks (RNN) have shown great abilities to model temporal dependencies. However, the performance of RNN is not prominent in resource limited tasks, even worse than the traditional feedforward neural networks (FNN). Furthermore, training time for RNN is much more than that for FNN. In recent years, some novel models are provided. They use non-recurrent architectures to model long term dependencies. In these architectures, they show that using gate mechanism is an effective method to construct acoustic models. On the other hand, it has been proved that using convolution operation is a good method to learn acoustic features. We hope to take advantages of both these two methods. In this paper we present a gated convolutional approach to low resource speech recognition tasks. The gated convolutional networks use convolutional architectures to learn input features and a gate to control information. Experiments are conducted on the OpenKWS, a series of low resource keyword search evaluations. From the results, the gated convolutional networks relatively decrease the WER about 6% over the baseline LSTM models, 5% over the DNN models and 3% over the BLSTM models. In addition, the new models accelerate the learning speed by more than 1.8 and 3.2 times compared to that of the baseline LSTM and BLSTM models.
Jian Kang 0006, Weiqiang Zhang 0001, Jia Liu 0001
ASRU2
2017 An LSTM-CTC based verification system for proxy-word based OOV keyword search
abstract
Proxy-word based out of vocabulary (OOV) keyword search has been proven to be quite effective in keyword search. In proxy-word based OOV keyword search, each OOV keyword is assigned several proxies and detections of the proxies are regarded as detections of the OOV keywords. However, the confidence scores of these detections are still those of the proxies from lattices. To obtain a better confidence measure, we employ an LSTM-CTC verification method in this work and the confidence scores are regenerated. OOV keyword search results on the evalpart1 dataset of the OpenKWS16 Evaluation have shown consistent improvement and the maximum relative improvement can reach 21.06% for the MWTW metric.
Zhiqiang Lv, Jian Kang 0006, Weiqiang Zhang 0001, Jia Liu 0001
ICASSP3
2017 Deep neural networks based speaker modeling at different levels of phonetic granularity
abstract
Recently, a hybrid deep neural network/i-vector framework has been proved effective for speaker verification, where the DNN trained to predict tied-triphone states (senones) is used to produce frame alignments for sufficient statistics extraction. In this work, in order to better understand the impact of different phonetic precision to speaker verification tasks, three levels of phonetic granularity are evaluated when doing frame alignments, which are tied-triphone state, monophone state and monophone. And the distribution of the features associated to a given phonetic unit is further modeled with multiple Gaussians rather than a single Gaussian. We also propose a fast and efficient way to generate phonetic units of different granularity by tying DNN's outputs according to the clustering results based on DNN derived senone embeddings. Experiments are carried out on the NIST SRE 2008 female tasks. Results show that using DNNs with less precise phonetic units and more Gaussians per phonetic unit for speaker modeling generalize better to different speaker verification tasks.
Liang He 0003, Weiqiang Zhang 0001, Jia Liu 0001
ICASSP4
2016 A Novel Discriminative Score Calibration Method for Keyword Search
Zhiqiang Lv, Weiqiang Zhang 0001, Jia Liu 0001
INTERSPEECH3
2016 Improving Deep Neural Networks Based Speaker Verification Using Unlabeled Data
Liang He 0003, Weiqiang Zhang 0001, Jia Liu 0001
INTERSPEECH4
2016 The NDSC transcription system for the 2016 multi-genre broadcast challenge
abstract
The National Digital Switching System Engineering and Technological R&D Center (NDSC) speech-to-text transcription system for the 2016 multi-genre broadcast challenge is described. Various acoustic models based on deep neural network (DNN), such as hybrid DNN, long short term memory recurrent neural network (LSTM RNN), and time delay neural network (TDNN), are trained. The system also makes use of recurrent neural network language models (RNNLMs) for re-scoring and minimum Bayes risk (MBR) combination. The WER on test dataset of the speech-to-text task is 18.2%. Furthermore, to simulate real applications where manual segmentations were not available an automatic segmentation system based on long-term information is proposed. WERs based on the automatically generated segments were slightly worse than that based on the manual segmentations.
Xukui Yang 0001, Dan Qu 0003, Weiqiang Zhang 0001
SLT4
2015 Improved system fusion for keyword search
abstract
It has been demonstrated that system fusion can significantly improve the performance of keyword search. In this paper, we compare the performance of several widely-used arithmetic-based fusion methods using different normalization pipeline and try to find the best pipeline. A novel arithmetic-based fusion method is proposed in this work. The method supplies a more effective way to incorporate the number of systems which have non-zero scores for a detection. When tested on the development test dataset of the OpenKWS15 Evaluation, the proposed method achieves the highest maximum term-weighted value (MTWV) and actual term-weighted value (ATWV) among all other arithmetic-based fusion methods. Usually, discriminative fusion methods employing classifiers can outperform arithmetic-based fusion methods. A DNN-based fusion method is explored in this work. After word-burst information is added, the DNN-based fusion method outperforms all other methods. In addition, it is notable that our arithmetic-based method achieves the same MTWV as the DNN-based method.
Zhiqiang Lv, Cheng Lu 0007, Jian Kang 0006, Like Hui, Weiqiang Zhang 0001, Jia Liu 0001
ASRU6
2015 The THUEE system for the openKWS14 keyword search evaluation
abstract
The OpenKWS14 keyword search evaluation is one of the most challenging and influential evaluations in the field of speech recognition. Its goal is to build a high-performance keyword search system for a minority language with limited training data in a short period of time. We present the system of the Department of Electronic Engineering, Tsinghua University (THUEE team) for the OpenKWS14 keyword search evaluation. The highlights of the system include the use of convolutional maxout neural networks for acoustic modeling and the use of neural network language models for one-pass lattice generation. The final system is a fusion of 8 sub-systems. The system has achieved an actual term weighted value (ATWV) of 0.5107 for the full language pack (FullLP) condition in the evaluation, ranking third among the participating teams.
Zhiqiang Lv, Beili Song, Yongzhe Shi, Wei-lan Wu, Cheng Lu 0007, Weiqiang Zhang 0001, Jia Liu 0001
ICASSP7
2015 Neuron sparseness versus connection sparseness in deep neural network for large vocabulary speech recognition
abstract
Exploiting sparseness in deep neural networks is an important method for reducing the computational cost. In this paper, we study neuron sparseness in deep neural networks for acoustic modeling. For the feed-forward stage, we only activate neurons whose input values are larger than a given threshold, and set the outputs of inactive nodes to zero. Thus, only a few nonzero outputs are fed to the next layer. Using this method, the output vector of each hidden layer becomes very sparse, so that the computational cost of the feed-forward algorithm can be reduced by adopting sparse matrix operations. The proposed method is evaluated in both small and large vocabulary speech recognition tasks, and results demonstrate that we can reduce the nonzero outputs to fewer than 20% of the total number of hidden nodes, without sacrificing speech recognition performance.
Jian Kang 0006, Cheng Lu 0007, Weiqiang Zhang 0001, Jia Liu 0001
ICASSP4
2015 Multi-resolution time frequency feature and complementary combination for short utterance speaker recognition
Weiqiang Zhang 0001, Jia Liu 0001
Multim. Tools Appl.2
2014 Improved phonotactic language recognition based on RNN feature reconstruction
abstract
Nowadays phone recognition followed by support vector machine (PR-SVM) has been proposed in language recognition tasks and shown encouraging results. However, it still suffers from the problems such as the curse of dimensionality led by the increasing order of the N-gram feature supervector, the fast increasing number of possible parameters because of fast exact match of the phoneme history, etc. These problems hamper the capability of N-gram vector space model (VSM) of handling long-term contexts. In this paper, a recurrent neural networks (RNN) based feature reconstruction (FR) method is presented to compensate for the deficiency of the N-grams feature for phonotactic language recognition in this paper. Experiments are implemented on 2009 National Institute of Standards and Technology language recognition evaluation (NIST LRE) database. The results show that the proposed method gives 8.76%, 3.82%, 11.93% relative error rate reduction for 30s, 10s, 3s respectively comparing with the baseline system.
Weiwei Liu 0001, Weiqiang Zhang 0001, Yongzhe Shi, An Ji, Jia Liu 0001
ICASSP2
2014 Variance regularization of RNNLM for speech recognition
abstract
Recurrent neural network language models (RNNLMs) have been proved superior to many other competitive language modeling techniques in terms of perplexity and word error rate. The remaining problem is the great computational complexity of RNNLMs in the output layer, resulting in long time for evaluation. Typically, a class-based RNNLM with the output layer factorized was proposed for speedup, which was still not fast enough for real-time systems. In this paper, a novel variance regularization algorithm is proposed for RNNLMs to address this problem. All the softmax-normalizing factors in the output layers are penalized to make them converge to one during the training phase, so that the output probability can be estimated efficiently via one dot-product of vectors in the output layer. The computational complexity of the output layer is reduced significantly from O(|V|H) to O(H). We further use this model for rescoring in an advanced CD-HMM-DNN system. Experimental results show that our proposed variance regularization algorithm works quite well, and the word prediction of the model is about 300 times faster than that of RNNLM without any obvious deteriorations in word error rate.
Yongzhe Shi, Weiqiang Zhang 0001, Jia Liu 0001
ICASSP2
2014 Phonotactic language recognition based on time-gap-weighted lattice kernels
Weiwei Liu 0001, Weiqiang Zhang 0001, Jia Liu 0001
INTERSPEECH2
2014 Speaker adaptation based on sparse and low-rank eigenphone matrix estimation
Dan Qu 0003, Weiqiang Zhang 0001, Bi-Cheng Li
INTERSPEECH3
2014 Spoken language recognition based on gap-weighted subsequence kernels
Weiqiang Zhang 0001, Weiwei Liu 0001, Yongzhe Shi, Jia Liu 0001
Speech Commun.1
2014 Efficient One-Pass Decoding with NNLM for Speech Recognition
abstract
Neural network language model (NNLM) has achieved very good results in the field of speech recognition, machine translation, etc. Direct decoding with NNLM is challenging for the overwhelmingly heavy burden in complexity. Most of the previous work focused on rescoring the N-best list and lattice with NNLM in the second pass. In this work, several techniques are explored to directly incorporate the NNLM into the decoder of speech recognition. A novel training algorithm based on variance regularization is proposed to approximate the softmax-normalizing factor as a constant for fast evaluation. Also, the evaluation of NNLM is further speeded up via our advanced storage. Moreover, a simple cache-based strategy is explored to avoid redundant computations during the decoding process. To the authors' knowledge, it is the first time to directly incorporate NNLM into decoding. We evaluate our proposed methods on an English-Switchboard phone-call speech-to-text task. Experimental results show that incorporating the NNLM into the decoder significantly reduces the word error rate (WER) by 1.5% and 1.4% absolutely on the Hub5'00-SWB and RT03S-FSH sets, respectively. Also, the decoding with NNLM is twice as fast as the baseline at the same word error rate.
Yongzhe Shi, Weiqiang Zhang 0001, Jia Liu 0001
IEEE Signal Process. Lett.2
2013 Compact acoustic modeling based on acoustic manifold using a mixture of factor analyzers
abstract
A compact acoustic model for speech recognition is proposed based on nonlinear manifold modeling of the acoustic feature space. Acoustic features of the speech signal is assumed to form a low-dimensional manifold, which is modeled by a mixture of factor analyzers. Each factor analyzer describes a local area of the manifold using a low-dimensional linear model. For an HMM-based speech recognition system, observations of a particular state are constrained to be located on part of the manifold, which may cover several factor analyzers. For each tied-state, a sparse weight vector is obtained through an iteration shrinkage algorithm, in which the sparseness is determined automatically by the training data. For each nonzero component of the weight vector, a low-dimensional factor is estimated for the corresponding factor model according to the maximum a posteriori (MAP) criterion, resulting in a compact state model. Experimental results show that compared with the conventional HMM-GMM system and the SGMM system, the new method not only contains fewer parameters, but also yields better recognition results.
Bi-Cheng Li, Weiqiang Zhang 0001
ASRU3
2013 Temporal kernel neural network language model
abstract
Using neural networks to estimate the probabilities of word sequences has shown significant promise for statistical language modeling. Typical modeling methods include multi-layer neural networks, log-bilinear networks and recurrent neural networks, etc. In this paper, we propose the temporal kernel neural network language model, a variant of models mentioned above. This model explicitly captures long-term dependencies of words with exponential kernel, where the memory of history is decayed exponentially. Additionally, several sentences with variable lengths as a mini-batch are efficiently implemented for speeding up. Experimental results show that the proposed model is very competitive to the recurrent neural network language model and obtains the lower perplexity of 111.6 (more than 10% reduction) than the state-of-the-art results reported in the standard Penn Treebank Corpus. We further apply this model to Wall Street Journal speech recognition task, and observe significant improvements in word error rate.
Yongzhe Shi, Weiqiang Zhang 0001, Jia Liu 0001
ICASSP2
2013 Parallel absolute-relative feature based phonotactic language recognition
Weiwei Liu 0001, Weiqiang Zhang 0001, Jia Liu 0001
INTERSPEECH2
2013 THU-EE system fusion for the NIST 2012 speaker recognition evaluation
Weiqiang Zhang 0001, Weiwei Liu 0001, Jia Liu 0001
INTERSPEECH1
2013 Exploiting articulatory features for pitch accent detection
abstract
Articulatory features describe how articulators are involved in making sounds. Speakers often use a more exaggerated way to pronounce accented phonemes, so articulatory features can be helpful in pitch accent detection. Instead of using the actual articulatory features obtained by direct measurement of articulators, we use the posterior probabilities produced by multi-layer perceptrons (MLPs) as articulatory features. The inputs of MLPs are frame-level acoustic features pre-processed using the split temporal context-2 (STC-2) approach. The outputs are the posterior probabilities of a set of articulatory attributes. These posterior probabilities are averaged piecewise within the range of syllables and eventually act as syllable-level articulatory features. This work is the first to introduce articulatory features into pitch accent detection. Using the articulatory features extracted in this way, together with other traditional acoustic features, can improve the accuracy of pitch accent detection by about 2%.
Junhong Zhao, Weiqiang Zhang 0001, Jia Liu 0001, Shanhong Xia
J. Zhejiang Univ. Sci. C3
2013 Rapid speaker adaptation using compressive sensing
Dan Qu 0003, Weiqiang Zhang 0001, Bi-Cheng Li
Speech Commun.3
2012 Bayesian Speaker Adaptation Based on a New Hierarchical Probabilistic Model
abstract
In this paper, a new hierarchical Bayesian speaker adaptation method called HMAP is proposed that combines the advantages of three conventional algorithms, maximum a posteriori (MAP), maximum-likelihood linear regression (MLLR), and eigenvoice, resulting in excellent performance across a wide range of adaptation conditions. The new method efficiently utilizes intra-speaker and inter-speaker correlation information through modeling phone and speaker subspaces in a consistent hierarchical Bayesian way. The phone variations for a specific speaker are assumed to be located in a low-dimensional subspace. The phone coordinate, which is shared among different speakers, implicitly contains the intra-speaker correlation information. For a specific speaker, the phone variation, represented by speaker-dependent eigenphones, are concatenated into a supervector. The eigenphone supervector space is also a low dimensional speaker subspace, which contains inter-speaker correlation information. Using principal component analysis (PCA), a new hierarchical probabilistic model for the generation of the speech observations is obtained. Speaker adaptation based on the new hierarchical model is derived using the maximum a posteriori criterion in a top-down manner. Both batch adaptation and online adaptation schemes are proposed. With tuned parameters, the new method can handle varying amounts of adaptation data automatically and efficiently. Experimental results on a Mandarin Chinese continuous speech recognition task show good performance under all testing conditions.
Weiqiang Zhang 0001, Bi-Cheng Li, Dan Qu 0003, Michael T. Johnson
IEEE Trans. Speech Audio Process.2
2011 Speaker adaptation based on speaker-dependent eigenphone estimation
abstract
Based on speaker dependent eigenphone estimation, a novel speaker adaptation technique is proposed in this paper. Different from conventional speaker adaptation approaches, the proposed method explicitly models the phone variations for each speaker through subspace modeling in the phone space. The phone coordinate, which is shared by all speakers, contains correlation information between different phones. During speaker adaptation, two schemes for estimation of the new speaker specific phone variation bases (namely eigenphones) are derived under maximum likelihood (ML) criterion and maximum a posteriori (MAP) criterion respectively. Supervised speaker adaptation experiments on a Mandarin Chinese continuous speech recognition task show that the new method outperforms both eigenvoice and maximum likelihood linear regression (MLLR) methods when sufficient adaptation data is available.
Weiqiang Zhang 0001, Bi-Cheng Li
ASRU2
2011 Robust Audio Fingerprinting Based on Local Spectral Luminance Maxima Scheme
Yongzhe Shi, Weiqiang Zhang 0001, Jia Liu 0001
INTERSPEECH2
2011 Time-Frequency Cepstral Features and Heteroscedastic Linear Discriminant Analysis for Language Recognition
abstract
The shifted delta cepstrum (SDC) is a widely used feature extraction for language recognition (LRE). With a high context width due to incorporation of multiple frames, SDC outperforms traditional delta and acceleration feature vectors. However, it also introduces correlation into the concatenated feature vector, which increases redundancy and may degrade the performance of backend classifiers. In this paper, we first propose a time-frequency cepstral (TFC) feature vector, which is obtained by performing a temporal discrete cosine transform (DCT) on the cepstrum matrix and selecting the transformed elements in a zigzag scan order. Beyond this, we increase discriminability through a heteroscedastic linear discriminant analysis (HLDA) on the full cepstrum matrix. By utilizing block diagonal matrix constraints, the large HLDA problem is then reduced to several smaller HLDA problems, creating a block diagonal HLDA (BDHLDA) algorithm which has much lower computational complexity. The BDHLDA method is finally extended to the GMM domain, using the simpler TFC features during re-estimation to provide significantly improved computation speed. Experiments on NIST 2003 and 2007 LRE evaluation corpora show that TFC is more effective than SDC, and that the GMM-based BDHLDA results in lower equal error rate (EER) and minimum average cost (Cavg) than either TFC or SDC approaches.
Weiqiang Zhang 0001, Liang He 0003, Jia Liu 0001, Michael T. Johnson
IEEE Trans. Speech Audio Process.1
2010 Combining Chinese spoken term detection systems via side-information conditioned linear logistic regression
Sha Meng, Weiqiang Zhang 0001, Jia Liu 0001
INTERSPEECH2
2010 A fast query by humming system based on notes
abstract
Query by humming (QBH), a content-based retrieval method, is an efficient way to search the song from a large database. The frame-based systems can achieve a good performance, but it is time-consuming. In this paper, we proposed an efficient note-based system, which is mainly comprised of noted-based linear scaling (NLS) and noted-based recursive align (NRA). The system after post-processing can achieve 96.1 % in Top5 and 0.211s in time. Index Terms: query by humming (QBH), musical information retrieval (MIR), note-based linear scaling (NLS), note-based recursive align (NRA) 1.
Jia Liu 0001, Weiqiang Zhang 0001
INTERSPEECH3
2010 Variant time-frequency cepstral features for speaker recognition
Weiqiang Zhang 0001, Liang He 0003, Jia Liu 0001
INTERSPEECH1
2008 An Equalized Heteroscedastic Linear Discriminant Analysis Algorithm
abstract
Heteroscedastic linear discriminant analysis (HLDA) is a widely used feature extraction algorithm. This method, however, suffers from unbalanced training data in some cases. In this letter, we equalize the objective function and statistics of HLDA and present an equalized HLDA algorithm, which balances the training data according to the class prior probability. Simulations as well as experimental results for the task of language identification are used to demonstrate the effectiveness of the proposed method.
Weiqiang Zhang 0001, Jia Liu 0001
IEEE Signal Process. Lett.1
2007 Two-Stage Method for Specific Audio Retrieval
abstract
Specific audio retrieval, also referred as similarity-based audio retrieval, means to detect and locate a given query audio segment in a long stored audio signal. In this paper, we proposed a two-stage method for specific audio retrieval. In the first stage, the histogram pruning algorithm is used for coarse detection. In the second stage, the partial distance technique is used for fine verification and localization. Experimental results show that the two-stage coarse-to-fine method offers fast search speed and improves the robustness to additive noise and compression encoding.
Weiqiang Zhang 0001, Jia Liu 0001
ICASSP (4)1