VLDB 2026 Research / reviewers in the wild / expert
Dong Wang 0013
dblp:40/3934-13
· DBLP profile ↗
89ranked-venue papers
17as first author
31since 2021 · last 2025
0000-0002-1286-0644ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 70 · 14 first-author · 26 since 2021Artificial intelligence and machine learning · 58 · 10 first-author · 22 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo LanguageabstractWe present a novel Automatic Speech Recognition (ASR) dataset for the Oromo language, a widely spoken language in Ethiopia and neighboring regions. The dataset was collected through a crowdsourcing initiative, encompassing a diverse range of speakers and phonetic variations. It consists of 100 hours of real-world audio recordings paired with transcriptions, covering read speech in both clean and noisy environments. This dataset addresses the critical need for ASR resources for the Oromo language which is underrepresented. To show its applicability for the ASR task, we conducted experiments using the Conformer model, achieving a Word Error Rate (WER) of 15.32% with hybrid CTC and AED loss and WER of 18.74% with pure CTC loss. Additionally, fine-tuning the Whisper model resulted in a significantly improved WER of 10.82%. These results establish baselines for Oromo ASR, highlighting both the challenges and the potential for improving ASR performance in Oromo. The dataset is publicly available at https://github.com/turinaf/sagalee and we encourage its use for further research and development in Oromo speech processing. Turi Abu, Ying Shi 0001, Thomas Fang Zheng, Dong Wang 0013 |
ICASSP | 4 |
| 2025 | CNVSRC 2024: The Second Chinese Continuous Visual Speech Recognition Challenge
Zehua Liu, Xiaolou Li, Chen Chen 0075, Lantian Li, Dong Wang 0013 |
INTERSPEECH | 5 |
| 2025 | Dual Orthogonality Sub-center Loss for Enhanced Anomalous Sound Detection
Dong Wang 0013, Jiqing Han 0001, Tieran Zheng, Guibin Zheng, Yongjun He 0002 |
INTERSPEECH | 1 |
| 2025 | Adaptive Across-Subcenter Representation Learning for Imbalanced Anomalous Sound Detection
Dong Wang 0013, Jiqing Han 0001, Guibin Zheng, Tieran Zheng, Yongjun He 0002 |
INTERSPEECH | 1 |
| 2025 | Neural Scoring: A Refreshed End-to-End Approach for Speaker Verification in Complex ConditionsabstractModern speaker verification systems primarily rely on speaker embeddings, followed by verification based on cosine similarity between the embedding vectors of the enrollment and test utterances. While effective, these methods struggle with multi-talker speech due to the unidentifiability of embedding vectors. In this paper, we propose Neural Scoring, a refreshed end-to-end framework that directly estimates verification posterior probabilities without relying on test-side embeddings, making it more robust to complex conditions, e.g., with multiple talkers. To make the training of such an end-to-end model more efficient, we introduce a large-scale trial e2e training strategy, where each test utterance pairs with a set of enrolled speakers, thus enabling processing of large-scale verification trials per batch. Experiments on VoxCeleb dataset demonstrate that Neural Scoring consistently outperforms both the baseline and competitive methods across various conditions, achieving an overall 70.36% reduction in Equal Error Rate compared to the baseline. Wan Lin, Junhui Chen, Lantian Li, Dong Wang 0013 |
IEEE Signal Process. Lett. | 6 |
| 2024 | An Investigation of Distribution Alignment in Multi-Genre Speaker RecognitionabstractMulti-genre speaker recognition is becoming increasingly popular due to its ability to better represent the complexities of real-world applications. However, a major challenge is the significant shift in the distribution of speaker vectors across different genres. While distribution alignment is a common approach to address this challenge, previous studies have mainly focused on aligning a source domain with a target domain, and the performance of multi-genre data is unknown.This paper presents a comprehensive study of mainstream distribution alignment methods on multi-genre data, where multiple distributions need to be aligned. We analyze various methods both qualitatively and quantitatively. Our experiments on the CN-Celeb dataset show that within-between distribution alignment (WBDA) performs relatively better. However, we also found that none of the investigated methods consistently improved performance in all test cases. This suggests that solely aligning the distributions of speaker vectors may not fully address the challenges posed by multi-genre speaker recognition. Further investigation is necessary to develop a more comprehensive solution. Junhui Chen, Namin Wang, Lantian Li, Dong Wang 0013 |
ICASSP | 5 |
| 2024 | Serialized Output Training by Learned Dominance
Ying Shi 0001, Lantian Li, Dong Wang 0013, Jiqing Han 0001 |
INTERSPEECH | 4 |
| 2024 | CNVSRC 2023: The First Chinese Continuous Visual Speech Recognition Challenge
Chen Chen 0075, Zehua Liu, Xiaolou Li, Lantian Li, Dong Wang 0013 |
INTERSPEECH | 5 |
| 2024 | Zero-Shot Fake Video Detection by Audio-Visual Consistency
Xiaolou Li, Zehua Liu, Chen Chen 0075, Lantian Li, Li Guo 0004, Dong Wang 0013 |
INTERSPEECH | 6 |
| 2024 | SE/BN Adapter: Parametric Efficient Domain Adaptation for Speaker Recognition
Lantian Li, Dong Wang 0013 |
INTERSPEECH | 3 |
| 2024 | Few-Shot Keyword Spotting from Mixed SpeechabstractFew-shot keyword spotting (KWS) aims to detect unknown keywords with limited training samples.A commonly used approach is the pre-training and fine-tuning framework.While effective in clean conditions, this approach struggles with mixed keyword spotting -simultaneously detecting multiple keywords blended in an utterance, which is crucial in real-world applications.Previous research has proposed a Mix-Training (MT) approach to solve the problem, however, it has never been tested in the few-shot scenario.In this paper, we investigate the possibility of using MT and other relevant methods to solve the two practical challenges together: few-shot and mixed speech.Experiments conducted on the LibriSpeech and Google Speech Command corpora demonstrate that MT is highly effective on this task when employed in either the pre-training phase or the fine-tuning phase.Moreover, combining SSL-based large-scale pre-training (HuBert) and MT fine-tuning yields very strong results in all the test conditions. Junming Yuan, Ying Shi 0001, Lantian Li, Dong Wang 0013, Askar Hamdulla |
INTERSPEECH | 4 |
| 2024 | A Comprehensive Investigation on Speaker Augmentation for Speaker RecognitionabstractData augmentation (DA) has played a pivotal role in the success of deep speaker recognition.Current DA techniques primarily focus on speaker-preserving augmentation, which does not change the speaker trait of the speech and does not create new speakers.Recent research has shed light on the potential of speaker augmentation, which generates new speakers to enrich the training dataset.In this study, we delve into two speaker augmentation approaches: speed perturbation (SP) and vocal tract length perturbation (VTLP).Despite the empirical utilization of both methods, a comprehensive investigation into their efficacy is lacking.Our study, conducted using two public datasets, VoxCeleb and CN-Celeb, revealed that both SP and VTLP are proficient at generating new speakers, leading to significant performance improvements in speaker recognition.Furthermore, they exhibit distinct properties in sensitivity to perturbation factors and data complexity, hinting at the potential benefits of their fusion.Our research underscores the substantial potential of speaker augmentation, highlighting the importance of in-depth exploration and analysis. Shibiao Xu, Lantian Li, Dong Wang 0013 |
INTERSPEECH | 5 |
| 2024 | On evaluation trials in speaker verification
Lantian Li, Di Wang 0039, Andrew Abel, Dong Wang 0013 |
Appl. Intell. | 4 |
| 2024 | Maximum Gaussianality training for deep speaker vector normalization
Yunqi Cai, Lantian Li, Andrew Abel, Xiaoyan Zhu 0001, Dong Wang 0013 |
Pattern Recognit. | 5 |
| 2024 | Keyword Guided Target Speech RecognitionabstractThis letter presents a new target speech recognition problem, where the target speech is defined by a keyword. For instance, when a person speaks “Hey Google” or “Help Me”, we hope the model can recognize the entire contextual speech of that person, even with strong interference speech from other people. The new problem is denoted by target content ASR (TC-ASR). The core challenge of TC-ASR is that the model needs to simultaneously detect the existence of the keyword from heavily mixed speech and recognize the target speech component using the information of the detected keyword segment. Surprisingly, our experiments show that an attention encoder-decoder (AED) model augmented with a keyword encoder can solve this problem pretty well. We also defined a key content spotting (KCS) task and tested the proposed model on it. Our experiments on the LibriMix dataset demonstrated that our approach could address the KCS task with a promising accuracy, outperforming two baseline models by a large margin. Further analysis shows that the proposed model identifies the target speech by a timbre cue, i.e., ensuring that the identified speech is coherent in speaker trait. Ying Shi 0001, Lantian Li, Dong Wang 0013, Jiqing Han 0001 |
IEEE Signal Process. Lett. | 3 |
| 2023 | CN-CVS: A Mandarin Audio-Visual Dataset for Large Vocabulary Continuous Visual to Speech SynthesisabstractResearch on Video to Speech Synthesis (VTS) surges recently and the focus is gradually shifting from small-vocabulary short-phrase VTS to large-vocabulary continuous VTS (LVC-VTS). A large-scale dataset with sufficient speakers and utterances is a prerequisite for such research, and the database is certainly language dependent.In this paper, we introduce CN-CVS, a large-scale Mandarin continuous visual-speech dataset, to support LVC-VTS research. The dataset contains about 200k utterances from more than 2500 individuals, amounting to more than 300 hours of visual-speech data. We built a state-of-the-art VTS model with the new dataset and conducted preliminary studies. Our results show that models that achieve good performance on small vocabulary tasks may perform very poor on CN-CVS, indicating that continuous VTS is indeed a challenging task, and the main challenge comes from the unconstrained vocabulary. The dataset and baseline code can be downloaded for free from http://cncvs.cslt.org. Chen Chen 0075, Dong Wang 0013, Thomas Fang Zheng |
ICASSP | 2 |
| 2023 | Spot Keywords From Very Noisy and Mixed Speech
Ying Shi 0001, Dong Wang 0013, Lantian Li, Jiqing Han 0001 |
INTERSPEECH | 2 |
| 2023 | Visualizing Data Augmentation in Deep Speaker Recognition
Pengqi Li, Lantian Li, Askar Hamdulla, Dong Wang 0013 |
INTERSPEECH | 4 |
| 2023 | CN-Celeb-AV: A Multi-Genre Audio-Visual Dataset for Person Recognition
Lantian Li, Xiaolou Li, Chen Chen 0075, Ruihai Hou, Dong Wang 0013 |
INTERSPEECH | 6 |
| 2023 | Ordered and Binary Speaker Embedding
Namin Wang, Lantian Li, Dong Wang 0013 |
INTERSPEECH | 5 |
| 2023 | A Multi-Scale Attentive Transformer for Multi-Instrument Symbolic Music Generation
Xipin Wei, Junhui Chen, Zirui Zheng, Li Guo 0004, Lantian Li, Dong Wang 0013 |
INTERSPEECH | 6 |
| 2023 | Random Cycle Loss and Its Application to Voice ConversionabstractSpeech disentanglement aims to decompose independent causal factors of speech signals into separate codes. Perfect disentanglement benefits to a broad range of speech processing tasks. This paper presents a simple but effective disentanglement approach based on cycle consistency loss and random factor substitution. This leads to a novel random cycle (RC) loss that enforces analysis-and-resynthesis consistency, a main principle of reductionism. We theoretically demonstrate that the proposed RC loss can achieve independent codes if well optimized, which in turn leads to superior disentanglement when combined with information bottleneck (IB). Extensive simulation experiments were conducted to understand the properties of the RC loss, and experimental results on voice conversion further demonstrate the practical merit of the proposal. Source code and audio samples can be found on the webpage http://rc.cslt.org. Dong Wang 0013, Lantian Li, Chen Chen 0075, Thomas Fang Zheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Real Additive Margin Softmax for Speaker VerificationabstractThe additive margin softmax (AM-Softmax) loss has delivered remarkable performance in speaker verification. A supposed behavior of AM-Softmax is that it can shrink within-class variation by putting emphasis on target logits, which in turn improves margin between target and non-target classes. In this paper, we conduct a careful analysis on the behavior of AM-Softmax loss, and show that this loss does not implement real max-margin training. Based on this observation, we present a Real AM-Softmax loss which involves a true margin function in the softmax training. Experiments conducted on VoxCeleb1, SITW and CNCeleb demonstrated that the corrected AM-Softmax loss consistently outperforms the original one. The code has been released at https://gitlab.com/csltstu/sunine. Lantian Li, Ruiqian Nai, Dong Wang 0013 |
ICASSP | 3 |
| 2022 | Reliable Visualization for Deep Speaker RecognitionabstractIn spite of the impressive success of convolutional neural networks (CNNs) in speaker recognition, our understanding to CNNs' internal functions is still limited. A major obstacle is that some popular visualization tools are difficult to apply, for example those producing saliency maps. The reason is that speaker information does not show clear spatial patterns in the temporal-frequency space, which makes it hard to interpret the visualization results, and hence hard to confirm the reliability of a visualization tool. In this paper, we conduct an extensive analysis on three popular visualization methods based on CAM: Grad-CAM, Score-CAM and Layer-CAM, to investigate their reliability for speaker recognition tasks. Experiments conducted on a state-of-the-art ResNet34SE model show that the Layer-CAM algorithm can produce reliable visualization, and thus can be used as a promising tool to explain CNN-based speaker models. The source code and examples are available in our project page: http://project.cslt.org/. Pengqi Li, Lantian Li, Askar Hamdulla, Dong Wang 0013 |
INTERSPEECH | 4 |
| 2022 | Oriental Language Recognition (OLR) 2021: Summary and AnalysisabstractThe fifth Oriental Language Recognition (OLR) Challenge focuses on language recognition in a variety of complex environments to promote its development.The OLR 2020 Challenge includes three tasks: (1) cross-channel language identification, (2) dialect identification, and (3) noisy language identification.We choose Cavg as the principle evaluation metric, and the Equal Error Rate (EER) as the secondary metric.There were 58 teams participating in this challenge and one third of the teams submitted valid results.Compared with the best baseline, the Cavg values of Top 1 system for the three tasks were relatively reduced by 82%, 62% and 48%, respectively.This paper describes the three tasks, the database profile, and the final results.We also outline the novel approaches that improve the performance of language recognition systems most significantly, such as the utilization of auxiliary information. Binling Wang, Wenxuan Hu, Qiulin Wang, Dong Wang 0013, Lin Li 0032, Qingyang Hong |
INTERSPEECH | 6 |
| 2022 | CN-Celeb: Multi-genre speaker recognition
Lantian Li, Jiawen Kang 0002, Yunqi Cai, Ravichander Vipperla, Thomas Fang Zheng, Dong Wang 0013 |
Speech Commun. | 9 |
| 2022 | A Principle Solution for Enroll-Test Mismatch in Speaker RecognitionabstractMismatch between enrollment and test conditions causes serious performance degradation on speaker recognition systems. This paper presents a statistics decomposition (SD) approach to solve this problem. This approach decomposes the PLDA score into three components that corresponding to enrollment, prediction and normalization respectively. Given that correct statistics are used in each component, the resultant score is theoretically optimal. A comprehensive experimental study was conducted on three datasets with different types of mismatch: (1) physical channel mismatch, (2) long-term speaker characteristics mismatch, (3) near-far recording mismatch. The results demonstrated that the proposed SD approach is highly effective, and outperforms the ad-hoc multi-condition training approach that is commonly adopted but not optimal in theory. Lantian Li, Dong Wang 0013, Jiawen Kang 0002, Renyu Wang, Zhendong Gao, Xiao Chen 0012 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Squeezing Value of Cross-Domain Labels: A Decoupled Scoring Approach for Speaker VerificationabstractDomain mismatch often occurs in real applications and causes serious performance reduction on speaker verification systems. The common wisdom is to collect cross-domain data and train a multi-domain PLDA model, with the hope to learn a domain-independent speaker subspace. In this paper, we firstly present an empirical study to show that simply adding cross-domain data does not help performance in conditions with enrollment-test mismatch. Careful analysis shows that this striking result is caused by the incoherent statistics between the enrollment and test conditions. Based on this analysis, we present a decoupled scoring approach that can maximally squeeze the value of cross-domain labels and obtain optimal verification scores in the enrollment-test mismatch condition. When the statistics are coherent, the new formulation falls back to the conventional PLDA. Experimental results on cross-channel test show that the proposed approach is highly effective and is a principal solution to domain mismatch. Lantian Li, Yang Zhang 0052, Jiawen Kang 0002, Thomas Fang Zheng, Dong Wang 0013 |
ICASSP | 5 |
| 2021 | Oriental Language Recognition (OLR) 2020: Summary and AnalysisabstractThe fifth Oriental Language Recognition (OLR) Challenge focuses on language recognition in a variety of complex environments to promote its development. The OLR 2020 Challenge includes three tasks: (1) cross-channel language identification, (2) dialect identification, and (3) noisy language identification. We choose Cavg as the principle evaluation metric, and the Equal Error Rate (EER) as the secondary metric. There were 58 teams participating in this challenge and one third of the teams submitted valid results. Compared with the best baseline, the Cavg values of Top 1 system for the three tasks were relatively reduced by 82%, 62% and 48%, respectively. This paper describes the three tasks, the database profile, and the final results. We also outline the novel approaches that improve the performance of language recognition systems most significantly, such as the utilization of auxiliary information. Binling Wang, Yiming Zhi, Lin Li 0032, Qingyang Hong, Dong Wang 0013 |
Interspeech | 7 |
| 2021 | Can We Trust Deep Speech Prior?abstractRecently, speech enhancement (SE) based on deep speech prior has attracted much attention, such as the variational auto-encoder with non-negative matrix factorization (VAE-NMF) architecture. Compared to conventional approaches that represent clean speech by shallow models such as Gaussians with a low-rank covariance, the new approach employs deep generative models to represent the clean speech, which often provides a better prior. Despite the clear advantage in theory, we argue that deep priors must be used with much caution, since the likelihood produced by a deep generative model does not always coincide with the speech quality. We designed a comprehensive study on this issue and demonstrated that based on deep speech priors, a reasonable SE performance can be achieved, but the results might be suboptimal. A careful analysis showed that this problem is deeply rooted in the disharmony between the flexibility of deep generative models and the nature of the maximum-likelihood (ML) training. Ying Shi 0001, Zhiyuan Tang, Lantian Li, Dong Wang 0013, Jiqing Han 0001 |
SLT | 5 |
| 2021 | Deep Normalization for Speaker VectorsabstractDeep speaker embedding has demonstrated state-of-the-art performance in speaker recognition tasks. However, one potential issue with this approach is that the speaker vectors derived from deep embedding models tend to be non-Gaussian for each individual speaker, and non-homogeneous for distributions of different speakers. These irregular distributions can seriously impact speaker recognition performance, especially with the popular PLDA scoring method, which assumes homogeneous Gaussian distribution. In this article, we argue that deep speaker vectors require deep normalization, and propose a deep normalization approach based on a novel discriminative normalization flow (DNF) model. We demonstrate the effectiveness of the proposed approach with experiments using the widely used SITW and CNCeleb corpora. In these experiments, the DNF-based normalization delivered substantial performance gains and also showed strong generalization capability in out-of-domain tests. Yunqi Cai, Lantian Li, Andrew Abel, Xiaoyan Zhu 0001, Dong Wang 0013 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2020 | CN-Celeb: A Challenging Chinese Speaker Recognition DatasetabstractRecently, researchers set an ambitious goal of conducting speaker recognition in unconstrained conditions where the variations on ambient, channel and emotion could be arbitrary. However, most publicly available datasets are collected under constrained environments, i.e., with little noise and limited channel variation. These datasets tend to deliver over-optimistic performance and do not meet the request of research on speaker recognition in unconstrained conditions.In this paper, we present CN-Celeb, a large-scale speaker recognition dataset collected ‘in the wild’. This dataset contains more than 130,000 utterances from 1,000 Chinese celebrities, and covers 11 different genres in real world. Experiments conducted with two state-of-the-art speaker recognition approaches (i-vector and x-vector) show that the performance on CN-Celeb is far inferior to the one obtained on Vox-Celeb, a widely used speaker recognition dataset. This result demonstrates that in real-life conditions, the performance of existing techniques might be much worse than it was thought. Our database is free for researchers and can be downloaded from http://project.cslt.org. Jiawen Kang 0002, Lantian Li, Kaicheng Li, Sitong Cheng, Pengyuan Zhang, Ziya Zhou, Yunqi Cai, Dong Wang 0013 |
ICASSP | 10 |
| 2020 | A Robust Audio-Visual Speech Enhancement ModelabstractMost existing audio-visual speech enhancement (AVSE) methods work well in conditions with strong noise, however when applied to conditions with a medium SNR, serious performance degradations are often observed. These degradations can be partly attributed to the feature-fusion(early fusion etc.) architecture that tightly couples the audio information that is very strong and the visual information that is relatively weak. In this paper, we present a safe AVSE approach that can make the visual stream contribute to audio speech enhancment(ASE) safely in conditions of various SNRs by late fusion.The key novelty is two-fold: Firstly, we define power binary masks (PBMs) as a rough representation of speech signals. This rough representation admits the weakness of the visual information and so can be easily predicted from the visual stream. Secondly, we design a posterior augmentation architecture that integrate the visual-derived PBMs to the audio-derived masks via a gating network. By this architecture, the entire performance is lower-bounded by the audio-based component. Our experiments on the Grid dataset demonstrated that this new approach consistently outperforms the audio-based system in all noise conditions, confirming that it is a safe way to incorporate visual knowledge in speech enhancement. Wupeng Wang, Dong Wang 0013, Xiao Chen 0012, Fengyu Sun |
ICASSP | 3 |
| 2020 | ASR-Free Pronunciation AssessmentabstractMost of the pronunciation assessment methods are based on local features derived from automatic speech recognition (ASR), e.g., the Goodness of Pronunciation (GOP) score. In this paper, we investigate an ASR-free scoring approach that is derived from the marginal distribution of raw speech signals. The hypothesis is that even if we have no knowledge of the language (so cannot recognize the phones/words), we can still tell how good a pronunciation is, by comparatively listening to some speech data from the target language. Our analysis shows that this new scoring approach provides an interesting correction for the phone-competition problem of GOP. Experimental results on the ERJ dataset demonstrated that combining the ASR-free score and GOP can achieve better performance than the GOP baseline. Sitong Cheng, Lantian Li, Zhiyuan Tang, Dong Wang 0013, Thomas Fang Zheng |
INTERSPEECH | 5 |
| 2020 | Domain-Invariant Speaker Vector Projection by Model-Agnostic Meta-LearningabstractDomain generalization remains a critical problem for speaker recognition, even with the state-of-the-art architectures based on deep neural nets. For example, a model trained on reading speech may largely fail when applied to scenarios of singing or movie. In this paper, we propose a domain-invariant projection to improve the generalizability of speaker vectors. This projection is a simple neural net and is trained following the Model-Agnostic Meta-Learning (MAML) principle, for which the objective is to classify speakers in one domain if it had been updated with speech data in another domain. We tested the proposed method on CNCeleb, a new dataset consisting of single-speaker multi-condition (SSMC) data. The results demonstrated that the MAML-based domain-invariant projection can produce more generalizable speaker vectors, and effectively improve the performance in unseen domains. Jiawen Kang 0002, Lantian Li, Yunqi Cai, Dong Wang 0013, Thomas Fang Zheng |
INTERSPEECH | 5 |
| 2020 | Neural Discriminant Analysis for Deep Speaker EmbeddingabstractProbabilistic Linear Discriminant Analysis (PLDA) is a popular tool in open-set classification/verification tasks.However, the Gaussian assumption underlying PLDA prevents it from being applied to situations where the data is clearly non-Gaussian.In this paper, we present a novel nonlinear version of PLDA named as Neural Discriminant Analysis (NDA).This model employs an invertible deep neural network to transform a complex distribution to a simple Gaussian, so that the linear Gaussian model can be readily established in the transformed space.We tested this NDA model on a speaker recognition task where the deep speaker vectors (x-vectors) are presumably non-Gaussian.Experimental results on two datasets demonstrate that NDA consistently outperforms PLDA, by handling the non-Gaussian distributions of the x-vectors. Lantian Li, Dong Wang 0013, Thomas Fang Zheng |
INTERSPEECH | 2 |
| 2019 | Gaussian-constrained Training for Speaker VerificationabstractNeural models, in particular the d-vector and x-vector architectures, have produced state-of-the-art performance on many speaker verification tasks. However, two potential problems of these neural models deserve more investigation. Firstly, both models suffer from `information leak', which means that some parameters participating in model training will be discarded during inference, i.e, the layers that are used as the classifier. Secondly, these models do not regulate the distribution of the derived speaker vectors. This `unconstrained distribution' may degrade the performance of the subsequent scoring component, e.g., PLDA. This paper proposes a Gaussian-constrained training approach that (1) discards the parametric classifier, and (2) enforces the distribution of the derived speaker vectors to be Gaussian. Our experiments on the VoxCeleb and SITW databases demonstrated that this new training approach produced more representative and regular speaker embeddings, leading to consistent performance improvement. Lantian Li, Zhiyuan Tang, Ying Shi 0001, Dong Wang 0013 |
ICASSP | 4 |
| 2019 | VAE-Based Regularization for Deep Speaker EmbeddingabstractDeep speaker embedding has achieved state-of-the-art performance in speaker recognition. A potential problem of these embedded vectors (called `x-vectors') are not Gaussian, causing performance degradation with the famous PLDA back-end scoring. In this paper, we propose a regularization approach based on Variational Auto-Encoder (VAE). This model transforms x-vectors to a latent space where mapped latent codes are more Gaussian, hence more suitable for PLDA scoring. Yang Zhang 0052, Lantian Li, Dong Wang 0013 |
INTERSPEECH | 3 |
| 2018 | Full-Info Training for Deep Speaker Feature LearningabstractIn recent studies, it has shown that speaker patterns can be learned from very short speech segments (e.g., 0.3 seconds) by a carefully designed convolutional & time-delay deep neural network (CT-DNN) model. By enforcing the model to discriminate the speakers in the training data, frame-level speaker features can be derived from the last hidden layer. In spite of its good performance, a potential problem of the present model is that it involves a parametric classifier, i.e., the last affine layer, which may consume some discriminative knowledge, thus leading to `information leak' for the feature learning. This paper presents a full-info training approach that discards the parametric classifier and enforces all the discriminative knowledge learned by the feature net. Our experiments on the Fisher database demonstrate that this new training scheme can produce more coherent features, leading to consistent and notable performance improvement on the speaker verification task. Lantian Li, Zhiyuan Tang, Dong Wang 0013, Thomas Fang Zheng |
ICASSP | 3 |
| 2018 | Deep Factorization for Speech SignalabstractVarious informative factors mixed in speech signals, leading to great difficulty when decoding any of the factors. An intuitive idea is to factorize each speech frame into individual informative factors, though it turns out to be highly difficult. Recently, we found that speaker traits, which were assumed to be long-term distributional properties, are actually short-time patterns, and can be learned by a carefully designed deep neural network (DNN). This discovery motivated a cascade deep factorization (CDF) framework that will be presented in this paper. The proposed framework infers speech factors in a sequential way, where factors previously inferred are used as conditional variables when inferring other factors. We will show that this approach can effectively factorize speech signals, and using these factors, the original speech spectrum can be recovered with a high accuracy. This factorization and reconstruction approach provides potential values for many speech processing tasks, e.g., speaker recognition and emotion recognition, as will be demonstrated in the paper. Lantian Li, Dong Wang 0013, Yixiang Chen 0003, Ying Shi 0001, Zhiyuan Tang, Thomas Fang Zheng |
ICASSP | 2 |
| 2018 | Human and Machine Speaker Recognition Based on Short Trivial EventsabstractHuman speech often has events that we will call trivial events, e.g., cough, laugh and sniff. Compared to regular speech, these trivial events are usually short and variable, thus generally regarded as not speaker discriminative and so are largely ignored by present speaker recognition research. However, these trivial events are highly valuable in some particular circumstances such as forensic examination, as they are less subjected to intentional change, so can be used to discover the genuine speaker from disguised speech. In this paper, we collect a trivial event speech database that involves 75 speakers and 6 types of events, and report preliminary speaker recognition results on this database, by both human listeners and machines. Particularly, the deep feature learning technique recently proposed by our group is utilized to analyze and recognize the trivial events, leading to acceptable equal error rates (EERs) ranging from 5% to 15% despite the extremely short durations (0.2-0.5 seconds) of these events. Comparing different types of events, `hmm' seems more speaker discriminative. Xiaofei Kang, Lantian Li, Zhiyuan Tang, Haisheng Dai, Dong Wang 0013 |
ICASSP | 7 |
| 2018 | Phonetic Temporal Neural Model for Language IdentificationabstractDeep neural models, particularly the long short-term memory recurrent neural network (LSTM-RNN) model, have shown great potential for language identification (LID). However, the use of phonetic information has been largely overlooked by most existing neural LID methods, although this information has been used very successfully in conventional phonetic LID systems. We present a phonetic temporal neural model for LID, which is an LSTM-RNN LID system that accepts phonetic features produced by a phone-discriminative DNN as the input, rather than raw acoustic features. This new model is similar to traditional phonetic LID methods, but the phonetic knowledge here is much richer: It is at the frame level and involves compacted information of all phones. Our experiments conducted on the Babel database and the AP16-OLR database demonstrate that the temporal phonetic neural approach is very effective, and significantly outperforms existing acoustic neural models. It also outperforms the conventional i-vector approach on short utterances and in noisy conditions. Zhiyuan Tang, Dong Wang 0013, Yixiang Chen 0003, Lantian Li, Andrew Abel |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Flexible and Creative Chinese Poetry Generation Using Neural MemoryabstractJiyuan Zhang, Yang Feng, Dong Wang, Yang Wang, Andrew Abel, Shiyue Zhang, Andi Zhang. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017. Jiyuan Zhang 0001, Yang Feng 0004, Dong Wang 0013, Andrew Abel, Shiyue Zhang 0001, Andi Zhang 0002 |
ACL (1) | 3 |
| 2017 | Memory-augmented Neural Machine TranslationabstractNeural machine translation (NMT) has achieved notable success in recent times, however it is also widely recognized that this approach has limitations with handling infrequent words and word pairs.This paper presents a novel memoryaugmented NMT (M-NMT) architecture, which stores knowledge about how words (usually infrequently encountered ones) should be translated in a memory and then utilizes them to assist the neural model.We use this memory mechanism to combine the knowledge learned from a conventional statistical machine translation system and the rules learned by an NMT system, and also propose a solution for out-of-vocabulary (OOV) words based on this framework.Our experiments on two Chinese-English translation tasks demonstrated that the M-NMT architecture outperformed the NMT baseline by 9.0 and 2.7 BLEU points on the two tasks, respectively.Additionally, we found this architecture resulted in a much more effective OOV treatment compared to competitive methods. Yang Feng 0004, Shiyue Zhang 0001, Andi Zhang 0002, Dong Wang 0013, Andrew Abel |
EMNLP | 4 |
| 2017 | Memory visualization for gated recurrent neural networks in speech recognitionabstractRecurrent neural networks (RNNs) have shown clear superiority in sequence modeling, particularly the ones with gated units, such as long short-term memory (LSTM) and gated recurrent unit (GRU). However, the dynamic properties behind the remarkable performance remain unclear in many applications, e.g., automatic speech recognition (ASR). This paper employs visualization techniques to study the behavior of LSTM and GRU when performing speech recognition tasks. Our experiments show some interesting patterns in the gated memory, and some of them have inspired simple yet effective modifications on the network structure. We report two of such modifications: (1) lazy cell update in LSTM, and (2) shortcut connections for residual learning. Both modifications lead to more comprehensible and powerful networks. Zhiyuan Tang, Ying Shi 0001, Dong Wang 0013, Yang Feng 0004, Shiyue Zhang 0001 |
ICASSP | 3 |
| 2017 | Deep Speaker Feature Learning for Text-Independent Speaker VerificationabstractRecently deep neural networks (DNNs) have been used to learn speaker features.However, the quality of the learned features is not sufficiently good, so a complex back-end model, either neural or probabilistic, has to be used to address the residual uncertainty when applied to speaker verification, just as with raw features.This paper presents a convolutional timedelay deep neural network structure (CT-DNN) for speaker feature learning.Our experimental results on the Fisher database demonstrated that this CT-DNN can produce highquality speaker features: even with a single feature (0.3 seconds including the context), the EER can be as low as 7.68%.This effectively confirmed that the speaker trait is largely a deterministic short-time property rather than a long-time distributional pattern, and therefore can be extracted from just dozens of frames. Lantian Li, Yixiang Chen 0003, Ying Shi 0001, Zhiyuan Tang, Dong Wang 0013 |
INTERSPEECH | 5 |
| 2017 | A Study on Replay Attack and Anti-Spoofing for Automatic Speaker VerificationabstractFor practical automatic speaker verification (ASV) systems, replay attack poses a true risk.By replaying a pre-recorded speech signal of the genuine speaker, ASV systems tend to be easily fooled.An effective replay detection method is therefore highly desirable.In this study, we investigate a major difficulty in replay detection: the over-fitting problem caused by variability factors in speech signal.An F-ratio probing tool is proposed and three variability factors are investigated using this tool: speaker identity, speech content and playback & recording device.The analysis shows that device is the most influential factor that contributes the highest over-fitting risk.A frequency warping approach is studied to alleviate the over-fitting problem, as verified on the ASV-spoof 2017 database. Lantian Li, Yixiang Chen 0003, Dong Wang 0013, Thomas Fang Zheng |
INTERSPEECH | 3 |
| 2017 | Collaborative Joint Training With Multitask Recurrent Model for Speech and Speaker RecognitionabstractAutomatic speech and speaker recognition are traditionally treated as two independent tasks and are studied separately. The human brain in contrast deciphers the linguistic content, and the speaker traits from the speech in a collaborative manner. This key observation motivates the work presented in this paper. A collaborative joint training approach based on multitask recurrent neural network models is proposed, where the output of one task is backpropagated to the other tasks. This is a general framework for learning collaborative tasks and fits well with the goal of joint learning of automatic speech and speaker recognition. Through a comprehensive study, it is shown that the multitask recurrent neural net models deliver improved performance on both automatic speech and speaker recognition tasks as compared to single-task systems. The strength of such multitask collaborative learning is analyzed, and the impact of various training configurations is investigated. Zhiyuan Tang, Lantian Li, Dong Wang 0013, Ravichander Vipperla |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Recurrent neural network training with dark knowledge transferabstractRecurrent neural networks (RNNs), particularly long short-term memory (LSTM), have gained much attention in automatic speech recognition (ASR). Although some successful stories have been reported, training RNNs remains highly challenging, especially with limited training data. Recent research found that a well-trained model can be used as a teacher to train other child models, by using the predictions generated by the teacher model as supervision. This knowledge transfer learning has been employed to train simple neural nets with a complex one, so that the final performance can reach a level that is infeasible to obtain by regular training. In this paper, we employ the knowledge transfer learning approach to train RNNs (precisely LSTM) using a deep neural network (DNN) model as the teacher. This is different from most of the existing research on knowledge transfer learning, since the teacher (DNN) is assumed to be weaker than the child (RNN); however, our experiments on an ASR task showed that it works fairly well: without applying any tricks on the learning scheme, this approach can train RNNs successfully even with limited training data. Zhiyuan Tang, Dong Wang 0013, Zhiyong Zhang 0001 |
ICASSP | 2 |
| 2016 | Chinese Song Iambics Generation with Neural Attention-Based Model
Tianyi Luo, Dong Wang 0013 |
IJCAI | 3 |
| 2016 | Improving Short Utterance Speaker Recognition by Modeling Speech Unit ClassesabstractShort utterance speaker recognition (SUSR) is highly challenging due to the limited enrollment and/or test data. We argue that the difficulty can be largely attributed to the mismatched prior distributions of the speech data used to train the universal background model (UBM) and those for enrollment and test. This paper presents a novel solution that distributes speech signals into a multitude of acoustic subregions that are defined by speech units, and models speakers within the subregions. To avoid data sparsity, a data-driven approach is proposed to cluster speech units into speech unit classes, based on which robust subregion models can be constructed. Further more, we propose a model synthesis approach based on maximum likelihood linear regression (MLLR) to deal with no-data speech unit classes. The experiments were conducted on a publicly available database SUD12. The results demonstrated that on a text-independent speaker recognition task where the test utterances are no longer than 2 seconds and mostly shorter than 0.5 seconds, the proposed subregion modeling offered a 21.51% relative reduction in equal error rate (EER), compared with the standard GMM-UBM baseline. In addition, with the model synthesis approach, the performance can be greatly improved in scenarios where no enrollment data are available for some speech unit classes. Lantian Li, Dong Wang 0013, Thomas Fang Zheng |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Similar Word Model for Unfrequent Word Enhancement in Speech RecognitionabstractThe popular n-gram language model (LM) is weak for unfrequent words. Conventional approaches such as class-based LMs pre-define some sharing structures (e.g., word classes) to solve the problem. However, defining such structures requires prior knowledge, and the context sharing based on these structures is generally inaccurate. This paper presents a novel similar word model to enhance unfrequent words. In principle, we enrich the context of an unfrequent word by borrowing context information from some “similar words.” Compared to conventional class-based methods, this new approach offers a fine-grained context sharing by referring to words that best match the target word, and it is more flexible as no sharing structures need to be defined by hand. Experiments on a large-scale Chinese speech recognition task demonstrated that the similar word approach can improve performance on unfrequent words significantly, while keeping the performance on general tasks almost unchanged. Xi Ma, Dong Wang 0013, Javier Tejedor |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Stochastic Top-k ListNetabstractListNet is a well-known listwise learning to rank model and has gained much attention in recent years.A particular problem of ListNet, however, is the high computation complexity in model training, mainly due to the large number of object permutations involved in computing the gradients.This paper proposes a stochastic ListNet approach which computes the gradient within a bounded permutation subset.It significantly reduces the computation complexity of model training and allows extension to Top-k models, which is impossible with the conventional implementation based on full-set permutations.Meanwhile, the new approach utilizes partial ranking information of human labels, which helps improve model quality.Our experiments demonstrated that the stochastic ListNet method indeed leads to better ranking performance and speeds up the model training remarkably. Tianyi Luo, Dong Wang 0013, Yiqiao Pan |
EMNLP | 2 |
| 2015 | Lasso-based reverberation suppression in automatic speech RecognitionabstractFar-field automatic speech recognition (ASR) is challenging, mainly attributed to the high reverberation in the recordings. A novel linear sparse prediction model has been proposed to estimate and suppress reverberation. This model considers reverberation as a mixture of early and late reflections of the direct signal and estimates the late reflection with Lasso. It has been demonstrated that this approach is promising in improving perceptual intelligibility, however it is unknown if the improvement can be propagated to ASR tasks. This paper applies the Lasso-based dereverberation approach to far-field speech recognition, and shows that it can deliver significant performance improvement for ASR based on deep neural networks (DNN). Particularly, we demonstrated that an utterance-based Lasso is sufficient to obtain good performance, which is important for applying the Lasso-based dereverberation to real-time ASR systems. Yiye Lin, Dong Wang 0013 |
ICASSP | 3 |
| 2015 | Recognize foreign low-frequency words with similar pairsabstractLow-frequency words place a major challenge for automatic speech recognition (ASR). The probabilities of these words, which are often important name entities, are generally underestimated by the language model (LM) due to their limited occurrences in the training data. Recently, we proposed a wordpair approach to deal with the problem, which borrows information of frequent words to enhance the probabilities of lowfrequency words. This paper presents an extension to the wordpair method by involving multiple ‘predicting words’ to produce better estimation for low-frequency words. We also employ this approach to deal with out-of-language words in the task of multi-lingual speech recognition. Xi Ma, Xiaoxi Wang, Dong Wang 0013, Zhiyong Zhang 0001 |
INTERSPEECH | 3 |
| 2015 | Learning speech rate in speech recognitionabstractA significant performance reduction is often observed in speech recognition when the rate of speech (ROS) is too low or too high. Most of present approaches to addressing the ROS variation focus on the change of speech signals in dynamic properties caused by ROS, and accordingly modify the dynamic model, e.g., the transition probabilities of the hidden Markov model (HMM). However, an abnormal ROS changes not only the dynamic but also the static property of speech signals, and thus can not be compensated for purely by modifying the dynamic model. This paper proposes an ROS learning approach based on deep neural networks (DNN), which involves an ROS feature as the input of the DNN model and so the spectrum distortion caused by ROS can be learned and compensated for. The experimental results show that this approach can deliver better performance for too slow and too fast utterances, demonstrating our conjecture that ROS impacts both the dynamic and the static property of speech. In addition, the proposed approach can be combined with the conventional HMM transition adaptation method, offering additional performance gains. Dong Wang 0013 |
INTERSPEECH | 3 |
| 2015 | Normalized Word Embedding and Orthogonal Transform for Bilingual Word TranslationabstractWord embedding has been found to be highly powerful to translate words from one language to another by a simple linear transform.However, we found some inconsistence among the objective functions of the embedding and the transform learning, as well as the distance measurement.This paper proposes a solution which normalizes the word vectors on a hypersphere and constrains the linear transform as an orthogonal transform.The experimental results confirmed that the proposed solution can offer better performance on a word similarity task and an English-to-Spanish word translation task. Dong Wang 0013, Chao Liu 0066, Yiye Lin |
HLT-NAACL | 2 |
| 2015 | Detection and reconstruction of clipped speech for speaker recognition
Fanhu Bie, Dong Wang 0013, Jun Wang 0073, Thomas Fang Zheng |
Speech Commun. | 2 |
| 2014 | Pruning deep neural networks by optimal brain damage
Chao Liu 0066, Zhiyong Zhang 0001, Dong Wang 0013 |
INTERSPEECH | 3 |
| 2014 | Feature analysis for discriminative confidence estimation in spoken term detection
Javier Tejedor, Doroteo T. Toledano, Dong Wang 0013, Simon King 0001, José Colás Pasamontes |
Comput. Speech Lang. | 3 |
| 2013 | Subspace models for bottleneck featuresabstractThe bottleneck (BN) feature, particularly based on deep structures, has gained significant success in automatic speech recognition (ASR). However, applying the BN feature to small/medium-scale tasks is nontrivial. An obvious reason is that the limited training data prevent from training a complicated deep network; another reason, which is more subtle, is that the BN feature tends to possess high inter-dimensional correlation, thus being inappropriate to be modeled by the conventional diagonal Gaussian mixture model (GMM). This difficulty can be mitigated by increasing the number of Gaussian components and/or employing full covariance matrices. These approaches, however, are not applicable for small/medium-scale tasks for which only a limited amount of training data is available. In this paper, we study the subspace Gaussian mixture model (SGMM) for BN features. The SGMM assumes full but shared covariance matrices, and hence can address the interdimensional correlation in a parsimonious way. This is particularly attractive for the BN feature, especially on small/mediumscale tasks, where the inter-dimensional correlation is high but the full covariance modeling is not affordable due to the limited training data. Our preliminary experiments on the Resource Management (RM) database demonstrate that the SGMM can deliver significant performance improvement for ASR systems based on BN features. Jun Qi 0002, Dong Wang 0013, Javier Tejedor |
INTERSPEECH | 2 |
| 2013 | Bottleneck features based on gammatone frequency cepstral coefficientsabstractRecent work demonstrates impressive success of the bottleneck (BN) feature in speech recognition, particularly with deep networks plus appropriate pre-training. A widely admitted advantage associated with the BN feature is that the network structure can learn multiple environmental conditions with abundant training data. For tasks with limited training data, however, this multi-condition training is unavailable, and so the networks tend to be over-fitted and sensitive to acoustic condition changes. A possible solution is to base the BN features on a channel-robust primary feature. In this paper, we propose to derive the BN feature based on Gammatone frequency cepstral coefficients (GFCCs). The GFCC feature has shown nice robustness against acoustic change, due to its capability of simulating the auditory system of humans. The idea is to integrate the advantage of the GFCC feature in acoustic robustness and the advantage of the BN feature in signal representation, so that the BN feature can be improved in the condition of mismatched training/test channels. This is particularly useful for small-scale tasks for which the training data are often limited. The experiments are conducted on the WSJCAM0 database, where the test utterances are mixed with noises at various SNR levels to simulate the channel change. The results confirm that the GFCC-based BN feature is much more robust than the BN features based on the MFCC and the PLP. Furthermore, the primary GFCC feature and the GFCC-based BN feature can be concatenated, leading to a more robust combined feature which provides considerable performance gains in all the tested noise conditions. Jun Qi 0002, Dong Wang 0013, Javier Tejedor |
INTERSPEECH | 2 |
| 2013 | Sequential model adaptation for speaker verificationabstractGMM-UBM-based speaker verification heavily relies on well-trained UBMs.In practice, it is not often easy to obtain a UBM that fully matches the acoustic channel in operation.In a previous study, we proposed to address this problem by a novel sequential UBM adaptation approach based on MAP.This work extends the study by applying the sequential approach to speaker model adaptation.In addition, we investigate a new feature-space sequential adaptation approach based on feature MAP linear regression (fMAPLR) and compare it with the previously proposed model-space MAP approach.We find that these two approaches are complementary and can be combined to deliver additional performance gains.The experiments conducted on a time-varying speech database demonstrate that the proposed MAP-fMAPLR approach leads to significant EER reduction with two mismatched UBMs (25% and 39% respectively). Jun Wang 0073, Dong Wang 0013, Thomas Fang Zheng, Javier Tejedor |
INTERSPEECH | 2 |
| 2013 | Auditory features based on Gammatone filters for robust speech recognitionabstractA major challenge for automatic speech recognition (ASR) relates to significant performance reduction in noisy environments. Recent research has shown that auditory features based on Gammatone filters are promising to improve robustness of ASR systems against noise, though the research is far from extensive and generalizability of the new features is unknown. This paper presents our implementation of the Gamma-tone filter-based feature and the experimental results on Mandarin speech data. By some thorough designs, we obtained significant performance gains with the new feature in various noise conditions when compared with the widely used MFCC and PLP features. A particular novelty of our implementation is that the filter design is purely in the time domain. This means that the channel signals are obtained with a set of Gammatone filters applied directly on the speech signals in time domain, which is totally different from the commonly adopted frequency-domain design that first converts signals to spectra and then applies the filter banks upon them. The time-domain implementation on the one hand avoids the approximation introduced by short-time spectral analysis and hence is more precise; and on the other hand, it avoids the complex spectral computation and hence simplifies hardware realization. Jun Qi 0002, Dong Wang 0013, Runsheng Liu |
ISCAS | 2 |
| 2013 | Evolutionary discriminative confidence estimation for spoken term detection
Javier Tejedor, Alejandro Echeverría, Dong Wang 0013, Ravichander Vipperla |
Multim. Tools Appl. | 3 |
| 2012 | Speech overlap detection and attribution using convolutive non-negative sparse codingabstractOverlapping speech is known to degrade speaker diarization performance with impacts on speaker clustering and segmentation. While previous work made important advances in detecting overlapping speech intervals and in attributing them to relevant speakers, the problem remains largely unsolved. This paper reports the first application of convolutive non-negative sparse coding (CNSC) to the overlap problem. CNSC aims to decompose a composite signal into its underlying contributory parts and is thus naturally suited to overlap detection and attribution. Experimental results on NIST RT data show that the CNSC approach gives comparable results to a state-of-the-art hidden Markov model based overlap detector. In a practical diarization system, CNSC based speaker attribution is shown to reduce the speaker error by over 40% relative in overlapping segments. Ravichander Vipperla, Jürgen T. Geiger, Simon Bozonnet, Dong Wang 0013, Nicholas W. D. Evans, Björn W. Schuller, Gerhard Rigoll |
ICASSP | 4 |
| 2012 | N-gram FST Indexing for Spoken Term DetectionabstractAn efficient indexing scheme is essentially important for spoken term detection (STD) on large databases, particularly for phone-based systems that have been widely adopted to achieve vocabulary-independent detection. While the finite state transducer (FST) composition provides a standard indexing approach, the n-gram reverse indexing is more flexible in connectivity representation and confidence measuring and therefore may result in better performance than searching within the original lattices or the equivalent FSTs. In this paper we present an n-gram FST indexing approach which combines the flexibility of n-gram indexing and the efficiency of FST indexing. Specifically, we employ the n-gram indexing to relax connectivity in original lattices and then formalize the indices into an FST for online search. We demonstrate this approach with a phone-based STD task where the lattice is sparse due to strong language models. The results show that n-gram FST indexing provides not only better detection performance than lattice search, but also a faster detection than both conventional n-gram and FST indexing. Index Terms: spoken term indexing, finite state transducer, spoken term detection, speech recognition Chao Liu 0066, Dong Wang 0013, Javier Tejedor |
INTERSPEECH | 2 |
| 2012 | Heterogeneous Convolutive Non-Negative Sparse CodingabstractConvolutive non-negative matrix factorization (CNMF) and its sparse version, convolutive non-negative sparse coding (CNSC), exhibit great success in speech processing. A particular limitation of the current CNMF/CNSC approaches is that the convolution ranges of the bases in learning are identical, resulting in patterns covering the same time span. This is obvious unideal as most of sequential signals, for example speech, involve patterns with a multitude of time spans. This paper extends the CMNF/CNSC algorithm and presents a heterogeneous learning approach which can learn bases with non-uniformed convolution ranges. The validity of this extension is demonstrated with a simple speech separation task Dong Wang 0013, Javier Tejedor |
INTERSPEECH | 1 |
| 2012 | Term-Dependent Confidence Normalisation for Out-of-Vocabulary Spoken Term Detection
Dong Wang 0013, Javier Tejedor, Simon King 0001, Joe Frankel |
J. Comput. Sci. Technol. | 1 |
| 2012 | A Comparative Study of Bottom-Up and Top-Down Approaches to Speaker DiarizationabstractThis paper presents a theoretical framework to analyze the relative merits of the two most general, dominant approaches to speaker diarization involving bottom-up and top-down hierarchical clustering. We present an original qualitative comparison which argues how the two approaches are likely to exhibit different behavior in speaker inventory optimization and model training: bottom-up approaches will capture comparatively purer models and will thus be more sensitive to nuisance variation such as that related to the speech content; top-down approaches, in contrast, will produce less discriminative speaker models but, importantly, models which are potentially better normalized against nuisance variation. We report experiments conducted on two standard, single-channel NIST RT evaluation datasets which validate our hypotheses. Results show that competitive performance can be achieved with both bottom-up and top-down approaches (average DERs of 21% and 22%), and that neither approach is superior. Speaker purification, which aims to improve speaker discrimination, gives more consistent improvements with the top-down system than with the bottom-up system (average DERs of 19% and 25%), thereby confirming that the top-down system is less discriminative and that the bottom-up system is less stable. Finally, we report a new combination strategy that exploits the merits of the two approaches. Combination delivers an average DER of 17% and confirms the intrinsic complementary of the two approaches. Nicholas W. D. Evans, Simon Bozonnet, Dong Wang 0013, Corinne Fredouille, Raphaël Troncy |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | Direct posterior confidence for out-of-vocabulary spoken term detectionabstractSpoken term detection (STD) is a key technology for spoken information retrieval. As compared to the conventional speech transcription and keyword spotting, STD is an open-vocabulary task and has to address out-of-vocabulary (OOV) terms. Approaches based on subword units, for example phones, are widely used to solve the OOV issue; however, performance on OOV terms is still substantially inferior to that of in-vocabulary (INV) terms. The performance degradation on OOV terms can be attributed to a multitude of factors. One particular factor we address in this article is the unreliable confidence estimation caused by weak acoustic and language modeling due to the absence of OOV terms in the training corpora. We propose a direct posterior confidence derived from a discriminative model, such as multilayer perceptron (MLP). The new confidence considers a wide-range acoustic context which is usually important for speech recognition and retrieval; moreover, it localizes on detected speech segments and therefore avoids the impact of long-span word context which is usually unreliable for OOV term detection. In this article, we first develop an extensive discussion about the modeling weakness problem associated with OOV terms, and then propose our approach to address this problem based on direct poster confidence. Our experiments carried out on spontaneous and conversational multiparty meeting speech, demonstrate that the proposed technique provides a significant improvement in STD performance as compared to conventional lattice-based confidence, in particular for OOV terms. Furthermore, the new confidence estimation approach is fused with other advanced techniques for OOV treatment, such as stochastic pronunciation modeling and discriminative confidence normalization. This leads to an integrated solution for OOV term detection that results in a large performance improvement. Dong Wang 0013, Simon King 0001, Joe Frankel, Ravichander Vipperla, Nicholas W. D. Evans, Raphaël Troncy |
ACM Trans. Inf. Syst. | 1 |
| 2011 | Linguistic influences on bottom-up and top-down clustering for speaker diarizationabstractWhile bottom-up approaches have emerged as the standard, default approach to clustering for speaker diarization we have always found the top-down approach gives equivalent or superior performance. Our recent work shows that significant gains in performance can be obtained when cluster purification is applied to the output of top down systems but that it can degrade performance when applied to the output of bottom-up systems. This paper demonstrates that these observations can be accounted for by factors unrelated to the speaker and that they can impact more strongly on the performance of bottom-up clustering strategies than top-down strategies. Experimental results confirm that clusters produced through top-down clustering are better normalized against phone variation than those produced through bottom-up clustering and that this accounts for the observed inconsistencies in purification performance. The work highlights the need for marginalization strategies which should encourage convergence toward different speakers rather than toward nuisance factors such as that those related to the linguistic content. Simon Bozonnet, Dong Wang 0013, Nicholas W. D. Evans, Raphaël Troncy |
ICASSP | 2 |
| 2011 | Handling overlaps in spoken term detectionabstractSpoken term detection (STD) systems usually arrive at many overlapping detections which are often addressed with some pragmatic approaches, e.g. choosing the best detection to represent all the overlaps. In this paper we present a theoretical study based on a concept of acceptance space. In particular, we present two confidence estimation approaches based on Bayesian and evidence perspectives respectively. Analysis shows that both approaches possess respective ad vantages and shortcomings, and that their combination has the potential to provide an improved confidence estimation. Experiments conducted on meeting data confirm our analysis and show considerable performance improvement with the combined approach, in particular for out-of-vocabulary spoken term detection with stochastic pronunciation modeling. Dong Wang 0013, Nicholas W. D. Evans, Raphaël Troncy, Simon King 0001 |
ICASSP | 1 |
| 2011 | Online Pattern Learning for Non-Negative Convolutive Sparse CodingabstractThe unsupervised learning of spectro-temporal speech patterns is relevant in a broad range of tasks. Convolutive non-negative matrix factorization (CNMF) and its sparse version, convolutive non-negative sparse coding (CNSC), are powerful, related tools. A particular difficulty of CNMF/CNSC, however, is the high demand on computing power and memory, which can prohibit their application to large scale tasks. In this paper, we propose an online algorithm for CNMF and CNSC, which processes input data piece-by-piece and updates the learned patterns after the processing of each piece by using accumulated sufficient statistics. The online CNSC algorithm remarkably increases converge speed of the CNMF/CNSC pattern learning, thereby enabling its application to large scale tasks. Dong Wang 0013, Ravichander Vipperla, Nicholas W. D. Evans |
INTERSPEECH | 1 |
| 2011 | Parallel and Hierarchical Decision Making for Sparse Coding in Speech RecognitionabstractSparse coding exhibits promising performance in speech processing, mainly due to the large number of bases that can be used to represent speech signals. However, the high demand for computational power represents a major obstacle in the case of large datasets, as does the difficulty in utilising information scattered sparsely in high dimensional features. This paper reports the use of an online dictionary learning technique, proposed recently by the machine learning community, to learn large scale bases efficiently, and proposes a new parallel and hierarchical architecture to make use of the sparse information in high dimensional features. The approach uses multilayer perceptrons (MLPs) to model sparse feature subspaces and make local decisions accordingly; the latter are integrated by additional MLPs in a hierarchical way for making global decisions. Experiments on the WSJ database show that the proposed approach not only solves the problem of prohibitive computation with large-dimensional sparse features, but also provides better performance in a frame-level phone prediction task. Dong Wang 0013, Ravichander Vipperla, Nicholas W. D. Evans |
INTERSPEECH | 1 |
| 2011 | Letter-to-Sound Pronunciation Prediction Using Conditional Random FieldsabstractPronunciation prediction, or letter-to-sound (LTS) conversion, is an essential task for speech synthesis, open vocabulary spoken term detection and other applications dealing with novel words. Most current approaches (at least for English) employ data-driven methods to learn and represent pronunciation “rules” using statistical models such as decision trees, hidden Markov models (HMMs) or joint-multigram models (JMMs). The LTS task remains challenging, particularly for languages with a complex relationship between spelling and pronunciation such as English. In this paper, we propose to use a conditional random field (CRF) to perform LTS because it avoids having to model a distribution over observations and can perform global inference, suggesting that it may be more suitable for LTS than decision trees, HMMs or JMMs. One challenge in applying CRFs to LTS is that the phoneme and grapheme sequences of a word are generally of different lengths, which makes CRF training difficult. To solve this problem, we employed a joint-multigram model to generate aligned training exemplars. Experiments conducted with the AMI05 dictionary demonstrate that a CRF significantly outperforms other models, especially if n-best lists of predictions are generated. Dong Wang 0013, Simon King 0001 |
IEEE Signal Process. Lett. | 1 |
| 2011 | Stochastic Pronunciation Modeling for Out-of-Vocabulary Spoken Term DetectionabstractSpoken term detection (STD) is the name given to the task of searching large amounts of audio for occurrences of spoken terms, which are typically single words or short phrases. One reason that STD is a hard task is that search terms tend to contain a disproportionate number of out-of-vocabulary (OOV) words. The most common approach to STD uses subword units. This, in conjunction with some method for predicting pronunciations of OOVs from their written form, enables the detection of OOV terms but performance is considerably worse than for in-vocabulary terms. This performance differential can be largely attributed to the special properties of OOVs. One such property is the high degree of uncertainty in the pronunciation of OOVs. We present a stochastic pronunciation model (SPM) which explicitly deals with this uncertainty. The key insight is to search for all possible pronunciations when detecting an OOV term, explicitly capturing the uncertainty in pronunciation. This requires a probabilistic model of pronunciation, able to estimate a distribution over all possible pronunciations. We use a joint-multigram model (JMM) for this and compare the JMM-based SPM with the conventional soft match approach. Experiments using speech from the meetings domain demonstrate that the SPM performs better than soft match in most operating regions, especially at low false alarm probabilities. Furthermore, SPM and soft match are found to be complementary: their combination provides further performance gains. Dong Wang 0013, Simon King 0001, Joe Frankel |
IEEE Trans. Speech Audio Process. | 1 |
| 2010 | Stochastic pronunciation modelling and soft match for out-of-vocabulary spoken term detectionabstractA major challenge faced by a spoken term detection (STD) system is the detection of out-of-vocabulary (OOV) terms. Although a subword-based STD system is able to detect OOV terms, performance reduction is always observed compared to in-vocabulary terms. One challenge that OOV terms bring to STD is the pronunciation uncertainty. A commonly used approach to address this problem is a soft matching procedure, and the other is the stochastic pronunciation modelling (SPM) proposed by the authors. In this paper we compare these two approaches, and combine them using a discriminative decision strategy. Experimental results demonstrated that SPM and soft match are highly complementary, and their combination gives significant performance improvement to OOV term detection. Dong Wang 0013, Simon King 0001, Joe Frankel, Peter Bell 0001 |
ICASSP | 1 |
| 2010 | An integrated top-down/bottom-up approach to speaker diarizationabstractInternational audience Simon Bozonnet, Nicholas W. D. Evans, Corinne Fredouille, Dong Wang 0013, Raphaël Troncy |
INTERSPEECH | 4 |
| 2010 | Augmented set of features for confidence estimation in spoken term detectionabstractDiscriminative confidence estimation along with confidence normalisation have been shown to construct robust decision maker modules in spoken term detection (STD) systems. Discriminative confidence estimation, making use of termdependent features, has been shown to improve the widely used lattice-based confidence estimation in STD. In this work, we augment the set of these term-dependent features and show a significant improvement in the STD performance both in terms of ATWV and DET curves in experiments conducted on a Spanish geographical corpus. This work also proposes a multiple linear regression analysis to carry out the feature selection. Next, the most informative features derived from it are used within the discriminative confidence on the STD system. Javier Tejedor, Doroteo T. Toledano, Miguel Bautista, Simon King 0001, Dong Wang 0013, José Colás Pasamontes |
INTERSPEECH | 5 |
| 2010 | CRF-based stochastic pronunciation modeling for out-of-vocabulary spoken term detectionabstractOut-of-vocabulary (OOV) terms present a significant challenge to spoken term detection (STD). This challenge, to a large ex-tent, lies in the high degree of uncertainty in pronunciations of OOV terms. In previous work, we presented a stochastic pro-nunciation modeling (SPM) approach to compensate for this uncertainty. A shortcoming of our original work, however, is that the SPM was based on a joint-multigram model (JMM), which is suboptimal. In this paper, we propose to use con-ditional random fields (CRFs) for letter-to-sound conversion, which significantly improves quality of the predicted pronun-ciations. When applied to OOV STD, we achieve consider-able performance improvement with both a 1-best system and an SPM-based system. Index Terms: speech recognition, spoken term detection, con-ditional random field, joint multigram model Dong Wang 0013, Simon King 0001, Nicholas W. D. Evans, Raphaël Troncy |
INTERSPEECH | 1 |
| 2009 | Posterior-based confidence measures for spoken term detectionabstractConfidence measures play a key role in spoken term detection (STD) tasks. The confidence measure expresses the posterior probability of the search term appearing in the detection period, given the speech. Traditional approaches are based on the acoustic and language model scores for candidate detections found using automatic speech recognition, with Bayes' rule being used to compute the desired posterior probability. In this paper, we present a novel direct posterior-based confidence measure which, instead of resorting to the Bayesian formula, calculates posterior probabilities from a multi-layer perceptron (MLP) directly. Compared with traditional Bayesian-based methods, the direct-posterior approach is conceptually and mathematically simpler. Moreover, the MLP-based model does not require assumptions to be made about the acoustic features such as their statistical distribution and the independence of static and dynamic co-efficients. Our experimental results in both English and Spanish demonstrate that the proposed direct posterior-based confidence improves STD performance. Dong Wang 0013, Javier Tejedor, Joe Frankel, Simon King 0001, José Colás Pasamontes |
ICASSP | 1 |
| 2009 | A posterior probability-based system hybridisation and combination for spoken term detectionabstractSpoken term detection (STD) is a fundamental task for multimedia information retrieval. To improve the detection performance, we have presented a direct posterior-based confidence measure generated from a neural network. In this paper, we propose a detection-independent confidence estimation based on the direct posterior confidence measure, in which the decision making is totally separated from the term detection. Based on this idea, we first present a hybrid system which conducts the term detection and confidence estimation based on different sub-word units and then propose a combination method which merges detections from heterogeneous term detectors based on the direct posterior-based confidence. Experimental results demonstrated that the proposed methods improved system performance considerably for both English and Spanish. Index Terms: speech recognition, spoken term detection, confidence estimation, grapheme Javier Tejedor, Dong Wang 0013, Simon King 0001, Joe Frankel, José Colás Pasamontes |
INTERSPEECH | 2 |
| 2009 | Stochastic pronunciation modelling for spoken term detectionabstractA major challenge faced by a spoken term detection (STD) system is the detection of out-of-vocabulary (OOV) terms. Although a subword-based STD system is able to detect OOV terms, performance reduction is always observed compared to in-vocabulary terms. Current approaches to STD do not acknowledge the particular properties of OOV terms, such as pronunciation uncertainty. In this paper, we use a stochastic pronunciation model to deal with the uncertain pronunciations of OOV terms. By considering all possible term pronunciations, predicted by a joint-multigram model, we observe a significant performance improvement. Index Terms: joint-multigram, pronunciation model, spoken term detection, speech recognition Dong Wang 0013, Simon King 0001, Joe Frankel |
INTERSPEECH | 1 |
| 2009 | Term-dependent confidence for out-of-vocabulary term detectionabstractWithin a spoken term detection (STD) system, the decision maker plays an important role in retrieving reliable detections. Most of the state-of-the-art STD systems make decisions based on a confidence measure that is term-independent, which poses a serious problem for out-of-vocabulary (OOV) term detection. In this paper, we study a term-dependent confidence measure based on confidence normalisation and discriminative modelling, particularly focusing on its remarkable effectiveness for detecting OOV terms. Experimental results indicate that the term-dependent confidence provides much more significant improvement for OOV terms than terms in-vocabulary. Index Terms: confidence estimation, spoken term detection, speech recognition Dong Wang 0013, Simon King 0001, Joe Frankel, Peter Bell 0001 |
INTERSPEECH | 1 |
| 2008 | A comparison of phone and grapheme-based spoken term detectionabstractWe propose grapheme-based sub-word units for spoken term detection (STD). Compared to phones, graphemes have a number of potential advantages. For out-of-vocabulary search terms, phone- based approaches must generate a pronunciation using letter-to-sound rules. Using graphemes obviates this potentially error-prone hard decision, shifting pronunciation modelling into the statistical models describing the observation space. In addition, long-span grapheme language models can be trained directly from large text corpora. We present experiments on Spanish and English data, comparing phone and grapheme-based STD. For Spanish, where phone and grapheme-based systems give similar transcription word error rates (WERs), grapheme-based STD significantly outperforms a phone- based approach. The converse is found for English, where the phone- based system outperforms a grapheme approach. However, we present additional analysis which suggests that phone-based STD performance levels may be achieved by a grapheme-based approach despite lower transcription accuracy, and that the two approaches may usefully be combined. We propose a number of directions for future development of these ideas, and suggest that if grapheme-based STD can match phone-based performance, the inherent flexibility in dealing with out-of-vocabulary terms makes this a desirable approach. Dong Wang 0013, Joe Frankel, Javier Tejedor, Simon King 0001 |
ICASSP | 1 |
| 2008 | Growing bottleneck features for tandem ASR
Joe Frankel, Dong Wang 0013, Simon King 0001 |
INTERSPEECH | 2 |
| 2008 | A posterior approach for microphone array based speech recognitionabstractAutomatic speech recognition (ASR) becomes rather difficult in meetings domains because of the adverse acoustic conditions, including more background noise, more echo and reverberation and frequent cross-talking. Microphone arrays have been demonstrated able to boost ASR performance dramatically in such noisy and reverberant environments, with various beamforming algorithms. However, almost all existing beamforming measures work in the acoustic domain, resorting to signal processing theories and geometric explanation. This limits their application, and induces significant performance degradation when the geometric property is unavailable or hard to estimate, or if heterogenous channels exist in the audio system. In this paper, we preset a new posterior-based approach for array-based speech recognition. The main idea is, instead of enhancing speech signals, we try to enhance the posterior probabilities that frames belonging to recognition units, e.g., phones. These enhanced posteriors are then transferred to posterior probability based features and are modeled by HMMs, leading to a tandem ANN-HMM hybrid system presented by Hermansky et al.. Experimental results demonstrated the validity of this posterior approach. With the posterior accumulation or enhancement, significant improvement was achieved over the single channel baseline. Moreover, we can combine the acoustic enhancement and posterior enhancement together, leading to a hybrid acoustic-posterior beamforming approach, which works significantly better than just the acoustic beamforming, especially in the scenario with moving-speakers. Dong Wang 0013, Ivan Himawan, Joe Frankel, Simon King 0001 |
INTERSPEECH | 1 |
| 2008 | A comparison of grapheme and phoneme-based units for Spanish spoken term detection
Javier Tejedor, Dong Wang 0013, Joe Frankel, Simon King 0001, José Colás Pasamontes |
Speech Commun. | 2 |