EDBT 2026 Demo / reviewers in the wild / expert
Yonghong Yan 0002
dblp:96/2570-2
· DBLP profile ↗
201ranked-venue papers
7as first author
56since 2021 · last 2026
0000-0001-6907-5770ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 158 · 5 first-author · 38 since 2021Artificial intelligence and machine learning · 122 · 4 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LAES: A local adaptive edge-enhanced spectrogram method for unsupervised anomalous sound detection
Jiyu Lu, Wenbo Guan, Ming Zhang 0035, Ta Li, Yonghong Yan 0002 |
Signal Process. | 5 |
| 2026 | Multilevel contextual prompting for conversational ASR: unifying conversation history and hotwords with speech LLM
Gaofeng Cheng, Xuyang Wang 0002, Qingwei Zhao, Yonghong Yan 0002 |
Speech Commun. | 5 |
| 2026 | Data-Efficient Semi-Supervised Few-Shot Speaker Verification via Prototype Space OptimizationabstractSpeaker verification technology has widespread applications across many domains, benefiting from deep learning advancements. However, due to the high cost of acquiring labeled data, semi-supervised learning has emerged as a prominent research focus. Current semi-supervised learning frameworks commonly suffer from two limitations: (1) the labeled data distribution is often restricted, and (2) they still rely on a considerable amount of labeled data. To address these issues, we propose three different distribution scenarios of labeled data and construct a general semi-supervised framework. Furthermore, to enhance the guidance efficacy of limited labeled data, we innovatively employ prototype space optimization to strengthen the model's discriminative capability under low-resource scenarios. Experimental results demonstrate that on the Vox1-o test set, our approach achieves a 41.7% relative reduction in equal error rate compared to self-supervised baselines, and a 28.9% improvement over conventional semi-supervised framework baselines. Zhenduo Zhao, Shisong Wu, Pengyuan Zhang, Xueshuai Zhang, Yonghong Yan 0002 |
IEEE Signal Process. Lett. | 7 |
| 2026 | SIM-IDE-FFM: Similarity-Based Residual-Identity Feature Fusion for Speaker VerificationabstractRecent advances in speaker verification have focused on two primary areas: deepening or widening backbone networks and introducing different attention mechanisms in the main residual branch to enhance discriminative power. However, the shortcut branches in residual architectures are still treated as simple identity mappings, leaving their representational potential underexplored. In this paper, we propose SIM-IDE-FFM, a similarity-based identity feature enhancement framework which enables the trivial shortcut branch to explore correlations between identity feature maps through a similarity-based attention mechanism. SIM-IDE introduces an attention mechanism at the shortcut branch that models inter-channel similarity of identity inputs, enhancing channel-wise representations. A non-linear Feature Fusion Module (FFM) is further designed to efficiently combine original and enhanced features through nonlinear depthwise interactions. Without altering the backbones, SIM-IDE-FFM can be seamlessly integrated into ResNet, EcapaTDNN, and other state-of-the-art (SOTA) architectures. Experiments on VoxCeleb1 and cross-age subsets demonstrate consistent relative improvements across different frameworks, validating its robustness and generalization. Baizhu Li, Xueshuai Zhang, Pengyuan Zhang, Yonghong Yan 0002 |
IEEE Signal Process. Lett. | 5 |
| 2025 | From External Similarity to Internal Consistency: An Enhanced Retrieval-Based Method for LLMs' Reliable Content GenerationabstractArtificial Intelligence Generated Content (AIGC) has emerged as a mainstream research direction with the development of Large Language Models (LLMs). The hallucination of LLMs, however, always interweaves the generated content with outdated or fabricated information, making it hard to be fully trusted and severely hindering LLMs form being widely applied in real-life scenarios. To address this problem, Retrieval Augmented Generation (RAG) has been proposed, which incorporates external knowledge to assist LLMs with content generation and significantly alleviates the hallucination problem. Nonetheless, the vanilla RAG uses similarity as the sole criterion for selecting external knowledge, neglecting the problem of internal inconsistency within the knowledge itself, which may distract LLMs from focusing on the most important information during the content generation process and, therefore, has a negative impact on the generated content's reliability. In this paper, we propose a novel metric, Entropy-based Internal Consistency (EIC), to measure the internal consistency of the external knowledge which is then integrated with similarity to mutually determine the knowledge's importance. Experimental results demonstrate that the proposed metric can provide a more fine-grained signal for external knowledge selection, thereby enhancing the reliability of generated content. Wenbo Guan, Hangchen Liu, Jun Zhou 0024, Yonghong Yan 0002 |
CSCWD | 5 |
| 2025 | ZCS-CDiff: A Zero-Shot Code-Switching TTS System with Conformer-Based Diffusion Model
Yonghong Yan 0002 |
ICASSP | 4 |
| 2025 | Automatic Text Pronunciation Correlation Generation and Application for Contextual BiasingabstractEffectively distinguishing the pronunciation correlations between different written texts is a significant issue in linguistic acoustics. Traditionally, such pronunciation correlations are obtained through manually designed pronunciation lexicons. In this paper, we propose a data-driven method to automatically acquire these pronunciation correlations, called automatic text pronunciation correlation (ATPC). The supervision required for this method is consistent with the supervision needed for training end-to-end automatic speech recognition (E2E-ASR) systems, i.e., speech and corresponding text annotations. First, the iteratively-trained timestamp estimator (ITSE) algorithm is employed to align the speech with their corresponding annotated text symbols. Then, a speech encoder is used to convert the speech into speech embeddings. Finally, we compare the speech embeddings distances of different text symbols to obtain ATPC. Experimental results on Mandarin show that ATPC enhances E2E-ASR performance in contextual biasing and holds promise for dialects or languages lacking artificial pronunciation lexicons. Gaofeng Cheng, Haitian Lu, Chengxu Yang, Xuyang Wang 0002, Ta Li, Yonghong Yan 0002 |
ICASSP | 6 |
| 2025 | Debiased Training For Semi-supervised Sound Event DetectionabstractRecently, semi-supervised sound event detection has attracted increasing research interest due to the scarcity of labeled data. However, traditional semi-supervised learning methods can lead to training instability and confirmation bias because of potentially incorrect pseudo labels. To address this issue, we propose the debiased training, a novel approach to reduce the inherent bias of pseudo labels. Debiased training can effectively decouple the generation and utilization of pseudo labels to mitigate the error accumulation and promote model’s robustness against biased pseudo labels. In addition, we introduce the channel restruction module (CRM) to decrease redundant computing and facilitate representation ability. Experimental results on DCASE 2023 task4 dataset show that the proposed methods significantly enhance the performance of semi-supervised methods while maintaining relatively low computational complexity. Shengchang Xiao, Xueshuai Zhang, Pengyuan Zhang, Yonghong Yan 0002 |
ICASSP | 4 |
| 2025 | TV-MDiff: A Zero-Shot Text-To-Speech and Zero-Shot Voice Conversion System with Mamba-Based Diffusion ModelabstractMany studies have proposed zero-shot (ZS) speaker adaptation methods for Text-To-Speech (TTS) and Voice Conversion (VC) to synthesize speech for an unseen speaker from a reference speech segment without additional training. However, the models developed using these methods do not fully utilize their learned ZS speaker adaptation capability because they can only perform either ZS TTS or ZS VC. Additionally, the TTS task restricts the model from using untranscribed data for training. This limitation prevents the model from leveraging untranscribed data to further enhance its overall performance. In this paper, we propose TV-MDiff, a ZS TTS and ZS VC system with the Mamba-based diffusion model. We decouple speech features and use the diffusion model to model these decoupled attributes, thus achieving ZS TTS and ZS VC while allowing for further training with untranscribed data to enhance the model’s performance. Moreover, we introduce a Mamba-based WaveNet as the denoising network for the diffusion model, which improves sample generation quality while keeping the model parameters relatively small. The experiments demonstrate that the speech generated by our model achieves good results in terms of speaker similarity and quality for both ZS TTS and ZS VC tasks. Further exploration reveals that its ZS TTS capability is highly transferable and can be easily adapted to another language. Audio samples are available: https://echohzh.github.io/TV-MDiff/. Yonghong Yan 0002 |
IJCNN | 3 |
| 2025 | A comprehensive validation study on the influencing factors of cough-based COVID-19 detection through multi-center data with abundant metadata
Jiakun Shen, Xueshuai Zhang, Yanfen Tang, Pengyuan Zhang, Yonghong Yan 0002, Pengfei Ye, Shaoxing Zhang |
J. Biomed. Informatics | 5 |
| 2025 | UnitDiff: A Unit-Diffusion Model for Code-Switching Speech SynthesisabstractGiven the scarcity of Code-Switching (CS) datasets, most researchers synthesize CS speech using multiple monolingual datasets. However, this approach presents challenges in synthesizing CS speech, such as difficulty controlling the speaker's identity and causing low intelligibility of the generated speech. This letter proposes UnitDiff, a CS speech synthesis model based on the unit-diffusion framework. The model employs the self-supervised high-level representation ’soft unit' extracted from soft HuBERT to directly predict a clean mel-spectrogram$x_{0}$. This approach enhances control over speaker identity. A language tagging method is also introduced to improve speech intelligibility. Evaluation results validate the model's effectiveness in improving the intelligibility, speaker similarity, and speaker consistency of the generated CS speech. Yonghong Yan 0002 |
IEEE Signal Process. Lett. | 4 |
| 2025 | GALD-SE: Guided Anisotropic Lightweight Diffusion for Efficient Speech EnhancementabstractSpeech enhancement is designed to enhance the intelligibility and quality of speech across diverse noise conditions. Recently, diffusion models have gained lots of attention in speech enhancement area, achieving competitive results. Current diffusion-based methods blur the distribution of the signal with isotropic Gaussian noise and recover clean speech distribution from the prior. However, these methods often suffer from a substantial computational burden. We argue that the computational inefficiency partially stems from the oversight that speech enhancement is not purely a generative task; it primarily involves noise reduction and completion of missing information, while the clean clues in the original mixture do not need to be regenerated. In this paper, we propose a method that introduces noise with anisotropic guidance during the diffusion process, allowing the neural network to preserve clean clues within noisy recordings. This approach substantially reduces computational complexity while exhibiting robustness against various forms of noise and speech distortion. Experiments demonstrate that the proposed method achieves state-of-the-art results with only approximately 4.5 million parameters, a number significantly lower than that required by other diffusion methods. This effectively narrows the model size disparity between diffusion-based and predictive speech enhancement approaches. Additionally, the proposed method performs well in very noisy scenarios, demonstrating its potential for applications in highly challenging environments. Chengzhong Wang, Jianjun Gu 0005, Dingding Yao, Yonghong Yan 0002 |
IEEE Signal Process. Lett. | 5 |
| 2025 | Multi-Branch Coordinate Attention With Channel Dynamic Difference for Speaker VerificationabstractIn prior studies, researchers have proved the excellent performance of various deep neural networks on speaker verification (SV). However, most of the improvements of SV systems are aimed at modifying the specific network structure to enhance its robustness but with limited flexibility. In this paper, MCA-CDD, which is a novel universal residual block module and can easily replace the original residual block without increasing the block number, is proposed with adaptive multi-branch coordinate attention (MCA) and channel-level dynamic difference (CDD). The design of multiple branches enables CA to model time and frequency at multiple scales. In addition, CDD fusion is applied into the feature fusion process of the residual block. The CDD fusion at shallow positions of the model enables the model to learn the detailed speaker-related dynamic texture information of speech. Experiments are conducted on several ResNet backbones and the results on different seen and unseen test sets show significant improvements, outperforming the baseline by about relatively 15-20%. Zhenduo Zhao, Shisong Wu, Xueshuai Zhang, Pengyuan Zhang, Yonghong Yan 0002 |
IEEE Signal Process. Lett. | 6 |
| 2024 | Fetal Heart Sounds Classification Using Time-Cyclic Frequency Spectrogram and Hybrid Attention NetworkabstractFetal heart monitoring is a crucial method for assessing fetal health status. However, commonly used techniques may pose potential medical risks and are not suitable for long-term monitoring. In this paper, we propose utilizing fetal heart sounds (FHS) for classifying fetal health status, taking advantage of its non-invasive, safe, straightforward, and cost-effective properties. Firstly, we introduce a novel acoustic feature for fetal heart sounds, termed the time-cyclic frequency spectrogram. This feature emphasizes the periodicity of heartbeats and effectively captures the changes in fetal heart rate. Additionally, we implement a frequency band energy-weighted algorithm to mitigate interference from periodic noises. Secondly, we propose a hybrid attention network that integrates both global-local attention and time-cyclic frequency attention. This network leverages medical prior knowledge to focus on the most critical aspects of the time-cyclic frequency spectrogram. Experimental results demonstrate that the proposed feature effectively characterizes variations in fetal heart rate, and the hybrid attention network can accurately capture spectral line variations, leading to improved classification of fetal health conditions. Shengchang Xiao, Xueshuai Zhang, Pengyuan Zhang, Yonghong Yan 0002 |
BIBM | 4 |
| 2024 | Snore Sound Features Based on Percussive Enhancing and Positional Encoding Combined with Multi-Task Learning for Osahs DetectionabstractObstructive sleep apnea hypopnea syndrome (OSAHS) is a serious sleep disorder. As the typical symptom of OSAHS, snoring has been proved effective in OSAHS diagnosis and potential to replace the current laborious and expensive polysomnography. However, the lack of analysis on the characteristics of pathological snoring sounds limits the diagnosing performance. In this paper, we propose novel sound features for the classification of OSA, hypopnea and normal snores. The proposed features are based on percussive enhancing and positional encoding as the snores exhibit different percussive properties and temporal traits due to the disease generation mechanisms. To enhance the classification performance, we propose a multi-task learning framework to aid the main classification task by simultaneous learning of two related simple tasks. Experiments on real-recorded snoring sounds show that the proposed methods can greatly improve the classification AUC and ACC and the proposed system performs better than those in other literatures. Aolin Hu, Xueshuai Zhang, Shaoxing Zhang, Pengyuan Zhang, Pengfei Ye, Qingwei Zhao, Yonghong Yan 0002 |
ICASSP | 8 |
| 2024 | One-Epoch Training with Single Test Sample in Test Time for Better Generalization of Cough-Based Covid-19 Detection ModelabstractThe outbreak of COVID-19 has raised researchers’ attention to audio-based rapid disease detection. Most of the previous studies have obtained competitive detection performance. However, these results are usually obtained by testing data from the same source offline. When making cross-dataset testing, the performance may deteriorate dramatically due to inconsistent data distribution between different datasets. In addition, in practical application, the model has to make a prediction for the current test audio without any prior information, which requires good model generalization under limited training data. To address the above issues, we adopt a test-time training framework to achieve a cough-based COVID-19 detection model with better generalizability. In the model development stage, resnet18 serves as the backbone network and a self-supervised learning branch is added as an auxiliary task. In testing stage, the model parameters are first fine-tuned by the self-supervised branch with the single test audio as input, and then the classification head outputs predictions. The proposed method is validated on three open-source datasets using a variety of hyperparameters. In cross-dataset testing, AUC and UAR increase by 3.65% and 3% on average absolutely, respectively. The results show that the proposed framework is applicable to improve the model performance in practical application. Jiakun Shen, Xueshuai Zhang, Pengyuan Zhang, Yonghong Yan 0002, Qingwei Zhao, Ta Li, Yanfen Tang, Shaoxing Zhang |
ICASSP | 4 |
| 2024 | CosDiff: Code-Switching TTS Model Based on A Multi-Task DDIMabstractAlthough existing Text-To-Speech (TTS) synthesizers are able to generate high-quality speech in most cases, their overall performance is still affected by the distribution of the training data. When processing tasks that involve complex data distributions, such as code-switching TTS, these models might generate speech that sounds unnatural or has low speaker similarity. In this paper, we propose CosDiff, a Code-Switching TTS model based on a multi-task Denoising Diffusion Implicit Model (DDIM), which integrates Voice Conversion (VC) and TTS functionalities. Utilizing the VC function, we construct a single-speaker bilingual dataset for training, achieving a superior code-switching synthesis performance compared to the outcomes of the speaker encoder method, which trains with multiple single-speaker monolingual datasets. In addition, we employ strategies of directly predicting the clean data x0and progressive diffusion distillation, further accelerating the model’s sampling process. The experimental results demonstrate the efficacy of this method in improving the quality of generation, increasing sampling speed, and distilling the model. Kexin Lu, Yonghong Yan 0002 |
ICME | 4 |
| 2024 | MCRSpell: A metric learning of correct representation for Chinese spelling correction
Chengzhang Li, Ming Zhang 0035, Xuejun Zhang 0002, Yonghong Yan 0002 |
Expert Syst. Appl. | 4 |
| 2024 | A novel semi-blind source separation framework towards maximum signal-to-interference ratio
Jianjun Gu 0005, Dingding Yao, Yonghong Yan 0002 |
Signal Process. | 4 |
| 2024 | Conversational Short-Phrase Speaker Diarization via Self-Adjusting Speech Segmentation and Embedding ExtractionabstractConversational short-phrase speaker diarization focuses on diarizing the phrases that are short in duration. Nonetheless, conventional speaker diarization systems fail to give enough importance to conversational short phrases. This letter proposed a novel speaker diarization system to address this issue. Firstly, we employ an RNN-T model for joint speech recognition and speaker change detection. The speech recognition results can be utilized directly in downstream tasks while the speaker change points serve as guidance for the following steps. Secondly, we introduce self-adjusting speech segmentation, which dynamically adjusts segment lengths based on the temporal distribution of speaker change points. Thirdly, we introduce self-adjusting embedding extraction, which employs speaker encoders trained under different speech duration conditions by projecting them to the same embedding space. Our method achieves a major reduction of Diarization Error Rate (DER) and Conversational Diarization Error Rate (CDER) on the MagicData-RAMC and Mixer 6 datasets. Haitian Lu, Gaofeng Cheng, Yonghong Yan 0002 |
IEEE Signal Process. Lett. | 3 |
| 2024 | Interrelate Training and Clustering for Online Speaker DiarizationabstractIn clustering-based speaker diarization systems, the embedding clusters for distinctive speakers exhibit wide variability in size and density, posing difficulty for clustering accuracy. In spite of this, with the assistance of the overall distance relationships among speaker embeddings, most of the embeddings can be grouped to the correct cluster by sophisticated offline clustering algorithms. However, in online scenarios, such a complete distance relationships of the embeddings can not be obtained due to the incremental arrival of embeddings. Consequently, determining the number of clusters and then correctly grouping the embeddings become challenging in an online fashion. Furthermore, errors would accumulate quickly over time if the online clustering algorithm assigns the embeddings into clusters erroneously in the beginning. To address these problems, we designed a novel framework for online clustering. To reduce the high variability of speaker embeddings, we proposed the clustering guided embedding extractor training (CGEET) algorithm to encourage similarity between the size of the embedding space for different speakers in attempt to simplify the distance relationships of embeddings. The CGEET algorithm can grasp the distance information of the entire speaker embedding space and provide it to the online clustering algorithm. With this preliminary information, the distance thresholds guided online clustering (DTGOC) algorithm then processes incoming embeddings using a divide-and-conquer approach. It first handles the embeddings with explicit distance relationships and then searches for possible path combination they have with remaining embeddings in an online fashion. Moreover, in order to utilize the distance relationships of embeddings that are far apart in time, an online re-clustering strategy is incorporated in our DTGOC algorithm, which can alleviate error accumulation during online clustering. By implementing the above innovations, our proposed online clustering system achieves 14.00% DER with collar 0.25 at 2.5 s latency on the AISHELL-4, while the DER of the offline agglomerative hierarchical clustering system is 14.54%. Gaofeng Cheng, Runyan Yang, Pengyuan Zhang, Yonghong Yan 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | Boosting Cross-Domain Speech Recognition With Self-SupervisionabstractThe cross-domain performance of automatic speech recognition (ASR) could be severely hampered due to the mismatch between training and testing distributions. Since the target domain usually lacks labeled data, and domain shifts exist at acoustic and linguistic levels, it is challenging to perform unsupervised domain adaptation (UDA) for ASR. Previous work has shown that self-supervised learning (SSL) or pseudo-labeling (PL) is effective in UDA by exploiting the self-supervisions of unlabeled data. However, these self-supervisions also face performance degradation in mismatched domain distributions, which previous work fails to address. This work presents a systematic UDA framework to fully utilize the unlabeled data with self-supervision in the pre-training and fine-tuning paradigm. On the one hand, we apply continued pre-training and data replay techniques to mitigate the domain mismatch of the SSL pre-trained model. On the other hand, we propose a domain-adaptive fine-tuning approach based on the PL technique with three unique modifications: Firstly, we design a dual-branch PL method to decrease the sensitivity to the erroneous pseudo-labels; Secondly, we devise an uncertainty-aware confidence filtering strategy to improve pseudo-label correctness; Thirdly, we introduce a two-step PL approach to incorporate target domain linguistic knowledge, thus generating more accurate target domain pseudo-labels. Experimental results on various cross-domain scenarios demonstrate that the proposed approach effectively boosts the cross-domain performance and significantly outperforms previous approaches. Han Zhu 0004, Gaofeng Cheng, Jindong Wang 0001, Wenxin Hou, Pengyuan Zhang, Yonghong Yan 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2023 | Piecewise Position Encoding in Convolutional Neural Network for Cough-Based Covid-19 DetectionabstractA fast and efficient COVID-19 detection method is of vital importance to control the spread of the epidemic. Many studies have achieved good performance on cough-based COVID19 detection in the past two years. However, the effect of position information in time-frequency features of cough audio has been less considered in previous studies. Even the convolutional neural networks that are capable to learn position information may be affected by small transformations of input features. Therefore, we propose piecewise position encoding added to time-frequency features to provide supplementary position information explicitly. Considering the differences in recording devices among different people, we use modified instance normalization to achieve better generalization. The proposed methods are validated on three open-sourced datasets and achieve significant improvements in AUC and UAR. The proposed model also shows competitive results in detecting asymptomatic patients. Jiakun Shen, Xueshuai Zhang, Pengyuan Zhang, Yonghong Yan 0002, Shaoxing Zhang, Yanfen Tang, Fujie Zhang, Aijun Sun |
ICASSP | 4 |
| 2023 | Reminding the incremental language model via data-free self-distillation
Han Wang 0033, Ruiliu Fu, Chengzhang Li, Xuejun Zhang 0002, Jun Zhou 0024, Xing Bai, Yonghong Yan 0002, Qingwei Zhao |
Appl. Intell. | 7 |
| 2023 | SFA: Searching faster architectures for end-to-end automatic speech recognition models
Ta Li, Pengyuan Zhang, Yonghong Yan 0002 |
Comput. Speech Lang. | 4 |
| 2023 | The effect of source sparsity on independent vector analysis for blind source separation
Jianjun Gu 0005, Longbiao Cheng, Dingding Yao, Yonghong Yan 0002 |
Signal Process. | 5 |
| 2023 | First coarse, fine afterward: A lightweight two-stage complex approach for monaural speech enhancement
Feng Dang, Hangting Chen, Pengyuan Zhang, Yonghong Yan 0002 |
Speech Commun. | 5 |
| 2023 | Alternative Pseudo-Labeling for Semi-Supervised Automatic Speech RecognitionabstractWhen labeled data is insufficient, semi-supervised learning with the pseudo-labeling technique can significantly improve the performance of automatic speech recognition. However, pseudo-labels are often noisy, containing numerous incorrect tokens. Taking noisy labels as ground-truth in the loss function results in suboptimal performance. Previous works attempted to mitigate this issue by either filtering out the nosiest pseudo-labels or improving the overall quality of pseudo-labels. While these methods are effective to some extent, it is unrealistic to entirely eliminate incorrect tokens in pseudo-labels. In this work, we propose a novel framework named alternative pseudo-labeling to tackle the issue of noisy pseudo-labels from the perspective of the training objective. The framework comprises several components. Firstly, a generalized CTC loss function is introduced to handle noisy pseudo-labels by accepting alternative tokens in the positions of incorrect tokens. Applying this loss function in pseudo-labeling requires detecting incorrect tokens in the predicted pseudo-labels. In this work, we adopt a confidence-based error detection method that identifies the incorrect tokens by comparing their confidence scores with a given threshold, thus necessitating the confidence score to be discriminative. Hence, the second proposed technique is the contrastive CTC loss function that widens the confidence gap between the correctly and incorrectly predicted tokens, thereby improving the error detection ability. Additionally, obtaining satisfactory performance with confidence-based error detection typically requires extensive threshold tuning. Instead, we propose an automatic thresholding method that uses labeled data as a proxy for determining the threshold, thus saving the pain of manual tuning. Experiments demonstrate that alternative pseudo-labeling outperforms existing pseudo-labeling approaches on datasets in various domains and languages. Han Zhu 0004, Dongji Gao, Gaofeng Cheng, Daniel Povey, Pengyuan Zhang, Yonghong Yan 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2022 | Interrelate Training and Searching: A Unified Online Clustering Framework for Speaker DiarizationabstractFor online speaker diarization, samples arrive incrementally, and the overall distribution of the samples is invisible.Moreover, in most existing clustering-based methods, the training objective of the embedding extractor is not designed specially for clustering.To improve online speaker diarization performance, we propose a unified online clustering framework, which provides an interactive manner between embedding extractors and clustering algorithms.Specifically, the framework consists of two highly coupled parts: clustering-guided recurrent training (CGRT) and truncated beam searching clustering (TBSC).The CGRT introduces the clustering algorithm into the training process of embedding extractors, which could provide not only cluster-aware information for the embedding extractor, but also crucial parameters for the clustering process afterward.And with these parameters, which contain preliminary information of the metric space, the TBSC penalizes the probability score of each cluster, in order to output more accurate clustering results in online fashion with low latency.With the above innovations, our proposed online clustering system achieves 14.48% DER with collar 0.25 at 2.5s latency on the AISHELL-4, while the DER of the offline agglomerative hierarchical clustering is 14.57%. Qingxuan Li, Gaofeng Cheng, Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 6 |
| 2022 | Improving Streaming End-to-End ASR on Transformer-based Causal Models with Encoder States Revision StrategiesabstractThere is often a trade-off between performance and latency in streaming automatic speech recognition (ASR).Traditional methods such as look-ahead and chunk-based methods, usually require information from future frames to advance recognition accuracy, which incurs inevitable latency even if the computation is fast enough.A causal model that computes without any future frames can avoid this latency, but its performance is significantly worse than traditional methods.In this paper, we propose corresponding revision strategies to improve the causal model.Firstly, we introduce a real-time encoder states revision strategy to modify previous states.Encoder forward computation starts once the data is received and revises the previous encoder states after several frames, which is no need to wait for any right context.Furthermore, a CTC spike position alignment decoding algorithm is designed to reduce time costs brought by the proposed revision strategy.Experiments are all conducted on Librispeech datasets.Fine-tuning on the CTC-based wav2vec2.0model, our best method can achieve 3.7/9.2WERs on test-clean/other sets and brings 45% relative improvement for causal models, which is also competitive with the chunkbased methods and the knowledge distillation methods. Zehan Li, Haoran Miao, Keqi Deng, Gaofeng Cheng, Sanli Tian, Ta Li, Yonghong Yan 0002 |
INTERSPEECH | 7 |
| 2022 | NAS-SCAE: Searching Compact Attention-based Encoders For End-to-end Automatic Speech Recognition
Ta Li, Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 4 |
| 2022 | Knowledge Distillation For CTC-based Speech Recognition Via Consistent Acoustic Representation Learning
Sanli Tian, Keqi Deng, Zehan Li, Lingxuan Ye, Gaofeng Cheng, Ta Li, Yonghong Yan 0002 |
INTERSPEECH | 7 |
| 2022 | Open Source MagicData-RAMC: A Rich Annotated Mandarin Conversational(RAMC) Speech DatasetabstractThis paper introduces a high-quality rich annotated Mandarin conversational (RAMC) speech dataset called MagicData-RAMC.The MagicData-RAMC corpus contains 180 hours of conversational speech data recorded from native speakers of Mandarin Chinese over mobile phones with a sampling rate of 16 kHz.The dialogs in MagicData-RAMC are classified into 15 diversified domains and tagged with topic labels, ranging from science and technology to ordinary life.Accurate transcription and precise speaker voice activity timestamps are manually labeled for each sample.Speakers' detailed information is also provided.As a Mandarin speech dataset designed for dialog scenarios with high quality and rich annotations, MagicData-RAMC enriches the data diversity in the Mandarin speech community and allows extensive research on a series of speechrelated tasks, including automatic speech recognition, speaker diarization, topic detection, keyword search, text-to-speech, etc.We also conduct several relevant tasks and provide experimental results to help evaluate the dataset. Zehui Yang, Runyan Yang, Lingxuan Ye, Gaofeng Cheng, Yaohui Jin, Pengyuan Zhang, Lei Xie 0001, Yonghong Yan 0002 |
INTERSPEECH | 12 |
| 2022 | Improving Recognition of Out-of-vocabulary Words in E2E Code-switching ASR by Fusing Speech Generation Methods
Lingxuan Ye, Gaofeng Cheng, Runyan Yang, Zehui Yang, Sanli Tian, Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 7 |
| 2022 | Robust Cough Feature Extraction and Classification Method for COVID-19 Cough Detection Based on Vocalization CharacteristicsabstractA fast, efficient and accurate detection method of COVID-19 remains a critical challenge.Many cough-based COVID-19 detection researches have shown competitive results through artificial intelligence.However, the lack of analysis on vocalization characteristics of cough sounds limits the further improvement of detection performance.In this paper, we propose two novel acoustic features of cough sounds and a convolutional neural network structure for COVID-19 detection.First, a time-frequency differential feature is proposed to characterize dynamic information of cough sounds in time and frequency domain.Then, an energy ratio feature is proposed to calculate the energy difference caused by the phonation characteristics in different cough phases.Finally, a convolutional neural network with two parallel branches which is pre-trained on a large amount of unlabeled cough data is proposed for classification.Experiment results show that our proposed method achieves state-of-the-art performance on Coswara dataset for COVID-19 detection.The results on an external clinical dataset Virufy also show the better generalization ability of our proposed method. Xueshuai Zhang, Jiakun Shen, Jun Zhou 0024, Pengyuan Zhang, Yonghong Yan 0002, Yanfen Tang, Fujie Zhang, Shaoxing Zhang, Aijun Sun |
INTERSPEECH | 5 |
| 2022 | Wav2vec-S: Semi-Supervised Pre-Training for Low-Resource ASRabstractSelf-supervised pre-training could effectively improve the performance of low-resource automatic speech recognition (ASR).However, existing self-supervised pre-training are taskagnostic, i.e., could be applied to various downstream tasks.Although it enlarges the scope of its application, the capacity of the pre-trained model is not fully utilized for the ASR task, and the learned representations may not be optimal for ASR.In this work, in order to build a better pre-trained model for low-resource ASR, we propose a pre-training approach called wav2vec-S, where we use task-specific semi-supervised pretraining to refine the self-supervised pre-trained model for the ASR task thus more effectively utilize the capacity of the pretrained model to generate task-specific representations for ASR.Experiments show that compared to wav2vec 2.0, wav2vec-S only requires a marginal increment of pre-training time but could significantly improve ASR performance on in-domain, cross-domain and cross-lingual datasets.Average relative WER reductions are 24.5% and 6.6% for 1h and 10h fine-tuning, respectively.Furthermore, we show that semi-supervised pretraining could close the representation gap between the selfsupervised pre-trained model and the corresponding fine-tuned model through canonical correlation analysis. Han Zhu 0004, Gaofeng Cheng, Jindong Wang 0001, Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 6 |
| 2022 | Decoupled Federated Learning for ASR with Non-IID DataabstractAutomatic speech recognition (ASR) with federated learning (FL) makes it possible to leverage data from multiple clients without compromising privacy. The quality of FL-based ASR could be measured by recognition performance, communication and computation costs. When data among different clients are not independently and identically distributed (non-IID), the performance could degrade significantly. In this work, we tackle the non-IID issue in FL-based ASR with personalized FL, which learns personalized models for each client. Concretely, we propose two types of personalized FL approaches for ASR. Firstly, we adapt the personalization layer based FL for ASR, which keeps some layers locally to learn personalization models. Secondly, to reduce the communication and computation costs, we propose decoupled federated learning (DecoupleFL). On one hand, DecoupleFL moves the computation burden to the server, thus decreasing the computation on clients. On the other hand, DecoupleFL communicates secure high-level features instead of model parameters, thus reducing communication cost when models are large. Experiments demonstrate two proposed personalized FL-based ASR approaches could reduce WER by 2.3% - 3.4% compared with FedAvg. Among them, DecoupleFL has only 11.4% communication and 75% computation cost compared with FedAvg, which is also significantly less than the personalization layer based FL. Han Zhu 0004, Jindong Wang 0001, Gaofeng Cheng, Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 5 |
| 2022 | Modeling knowledge proficiency using multi-hierarchical capsule graph neural network
Wang Li 0007, Yonghong Yan 0002 |
Appl. Intell. | 3 |
| 2022 | Underwater Detection of Small-Volume Weak Target Echo in Harbor Scene Under Multisource InterferenceabstractOwing to the strong reverberation and various obstacle echoes interference in harbor scenes, the avtive detection performance of tracking methods for weak targets will be seriously reduced. Moreover, the conventional dereverberation and tracking methods generally cannot effectively separate the weak target under the overlapping echoes of multiple scattering sources. Focusing on solving the above problems, the correlation residual cumulative clustering (CRCC) algorithm is proposed to extract scattering features of weak target motion. There are two innovations in this paper: Firstly, the complex cepstrum filter is improved to suppress strong reverberation. Secondly, the correlation residual accumulation of adjacent frames is extracted to detect the reduction matrix containing the target trajectory. Finally, the FCM objective function is optimized via spatial-temporal constraint, thus the weak target can be effectively separated from the overlapping giant interference. The experimental results indicate the strong anti-jamming, low detection loss and short-term accumulation of our proposed model, which are remarkably superior to the traditional posterior probability models. Xingyue Zhou, Ning Wang 0107, Yonghong Yan 0002, Kunde Yang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | A Secondary Path-Decoupled Active Noise Control Algorithm Based on Deep LearningabstractActive noise control (ANC) systems are widely used to cancel unwanted noise. However, for high-level noise, the residual error signal cannot be fully eliminated because of the nonlinearity of the secondary path, resulting in the diverging of the adaptive filter. In this letter, we propose a secondary path-decoupled ANC (SPD-ANC) algorithm based on deep learning. Specifically, the secondary path decoupled module consisting of two time-domain convolutional recurrent networks, one for modeling the nonlinear secondary path and the other for modeling the reverse process, is employed to calculate the secondary path-decoupled (SPD) error signal. The control signal is then generated by an adaptive filter that is optimized towards minimizing the SPD error signal. Simulation results indicate that the proposed method outperforms the conventional ANC methods under different conditions. Daocheng Chen, Longbiao Cheng, Dingding Yao, Yonghong Yan 0002 |
IEEE Signal Process. Lett. | 5 |
| 2022 | An E2E-ASR-Based Iteratively-Trained Timestamp EstimatorabstractText-to-speech alignment, also known as time alignment, is essential for automatic speech recognition (ASR) systems used for speech retrieval tasks, such as keyword search and speech segment extraction. Previous works have used the Gaussian mixture model-hidden Markov model (GMM-HMM) forced alignment to improve the alignment performance. However, when used with end-to-end (E2E) ASR, GMM-HMM forced alignment causes extra reliance on expertise such as pronunciation lexica. It also increases the system complexity because GMM-HMMs are very dissimilar to E2E models. To tackle these two problems, we propose an E2E-ASR-based iteratively-trained timestamp estimator (ITSE), which performs alignment between token-level transcription and speech. We train ITSE first with coarse initial alignment targets generated using connectionist temporal classification (CTC) posteriors. During training, we iteratively perform realignment to update the targets. We attribute the effectiveness of the iterative training to ITSE’s two vital features. First, ITSE performs alignment using similarities between token and speech embeddings instead of frame-wise token classification posteriors. Second, ITSE uses speech embeddings that are aware of left context rather than global context. ITSE significantly outperforms CTC-based baselines in word alignment accuracy and is comparable to a GMM-HMM forced aligner. In short, ITSE is an accurate, lightweight text-to-speech alignment module implemented without expertise such as pronunciation lexica. Runyan Yang, Gaofeng Cheng, Pengyuan Zhang, Yonghong Yan 0002 |
IEEE Signal Process. Lett. | 4 |
| 2022 | ETEH: Unified Attention-Based End-to-End ASR and KWS ArchitectureabstractEven though attention-based end-to-end (E2E) automatic speech recognition (ASR) models have been yielding state-of-the-art recognition accuracy, they still fall behind many of the ASR models deployed in the industry in some crucial functionalities such as online processing and precise timestamps generating. This weakness prevents attention-based E2E ASR models from being applied in several essential speech tasks, such as online speech recognition and keyword searching (KWS). In this paper, we describe our proposed unified attention-based E2E ASR and KWS architecture–ETEH, which supports, in one model, both online and offline ASR decoding modes, thus allowing for precise and reliable KWS. “ETE” stands for attention-based E2E modeling, whereas “H” represents the hybrid gaussian mixture model and hidden Markov model (GMM-HMM). As a combination of both, ETEH is an attention-based E2E ASR architecture which utilizes the frame-wise time alignment (FTA) generated by GMM-HMM ASR models. This FTA is used to better the model in two ways: first, it helps the monotonic attentions of ETEH models to capture more accurate word time stamps, thus resulting in lower latency for online decoding; second, it helps ETEH models to provide accurate and reliable KWS results. Furthermore, we are able to combine both offline and online modes in one ETEH model and establish a concise system by adopt the universal training strategy. ETEH is functional and unique, and to the best of our knowledge, we can hardly find a comparable single attention-based E2E ASR system as the baseline. To evaluate ASR accuracy and latency for ETEH, we use our previously proposed monotonic truncated attention (MTA) based online CTC/attention (OCA) ASR models as baselines. Experimental results show that ETEH ASR models gain significant improvement in ASR latency compared to the baseline. To evaluate KWS performance, we compare ETEH models with CTC-based KWS models. Results demonstrate that our ETEH models achieve significantly better KWS performance compared to the CTC baselines. Gaofeng Cheng, Haoran Miao, Runyan Yang, Keqi Deng, Yonghong Yan 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2022 | Alleviating ASR Long-Tailed Problem by Decoupling the Learning of Representation and ClassificationabstractRecently, we have witnessed excellent improvement of end-to-end (E2E) automatic speech recognition (ASR). However, how to tackle the long-tailed data distribution problem while maintaining E2E ASR models' performance for high-frequency tokens is still challenging. To solve this challenge, we propose a novel decoupled ASR learning method for the sequence-to-sequence ASR architecture in this paper. Our method decouples the learning procedure of this model into two stages: representation learning and classification learning. In the representation learning stage, we use the encoder output of a pretrained language model as one of the ASR model’s learning targets, and propose threshold log cosine embedding loss (TLCE-loss) as the objective function. A frequency-mask cross-entropy loss (FMCE-loss) is also designed as an auxiliary loss. In the classification learning stage, we find that introducing a temperature into softmax function helps reduce the influence of negative samples on tail classes, thus mitigating the biased learning process for the classifier. Furthermore, we propose a weighted softmax (w-softmax) to adjust ASR posterior probabilities according to the token appearing frequency during inference. Additionally, we introduce tail word/character error rate (TWER / TCER) and head word/character error rate (HWER / HCER) that respectively evaluate the ASR accuracy for tail and head words/characters. Experimental results on the Switchboard and HKUST corpora show that our proposed method greatly outperforms the baseline, especially in TWER / TCER reduction. To the best of our knowledge, this is the first work to use a decoupled ASR learning method to alleviate the long-tailed problem in sequence-to-sequence ASR. Keqi Deng, Gaofeng Cheng, Runyan Yang, Yonghong Yan 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Self-Supervised Pre-Training for Attention-Based Encoder-Decoder ASR ModelabstractEnd-to-end (E2E) models, including the attention-based encoder-decoder (AED) models, have achieved promising performance on the automatic speech recognition (ASR) task. However, the supervised training process of the E2E model needs a large amount of speech-text paired data. In contrast, self-supervised pre-training can pre-train the model on the unlabeled data and then fine-tune it on the limited labeled data to realize better performance. Most of the previous self-supervised pre-training methods focus on learning hidden representations from speech but ignore how to utilize the unpaired text. As a result, previous works often pre-train an acoustic encoder and then fine-tune it as a classification based ASR model, such as Connectionist Temporal Classification (CTC) based model, rather than an AED model. In this paper, we propose a self-supervised pre-training method for the AED model (SP-AED). The SP-AED method contains acoustic pre-training for the encoder, linguistic pre-training for the decoder, and an adaptive combination fine-tuning for the whole system. We first design a linguistic pre-training method for decoder by utilizing the text-only data. The decoder will be pre-trained as a noise-condition language model to learn the prior distribution of the text. Then, we pre-train the AED encoder with the wav2vec2.0 method with some modifications. Finally, we combine the pre-trained encoder and decoder and fine-tune them on the limited labeled data. We design an adaptive combination method during fine-tuning by modifying the decoder’s input and output to prevent catastrophic forgetting. Experiments prove that compared with the random initialized models, the SP-AED pre-trained models can realize up to 17% relative improvement. And with similar model size or computational cost, we can get comparable results to other classification-based models on both English and Chinese corpus. Changfeng Gao, Gaofeng Cheng, Ta Li, Pengyuan Zhang, Yonghong Yan 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2021 | Far-Field Speech Recognition Based on Complex-Valued Neural Networks and Inter-Frame Similarity Difference MethodabstractFar-field automatic speech recognition (ASR) is a challenging task due to the background noise and reverberation. To address this issue, we introduce a novel end-to-end multi-channel far-field ASR architecture. First, we use a complex-valued CNN based architecture designed for speech tasks as a neural beamformer. Second, we propose an auxiliary mod-ule called absolute position regression module (APRM) with a position prediction loss to help the neural beamformer be better aware of the corresponding frequencies of each input time-frequency (T-F) bin. Third, inspired by the short-term stationarity of human speech, we propose an approach called the Inter-Frame Similarity Difference (IFSD) method to au-tomatically select useful channels as the inputs of the ASR backend from the outputs of the neural beamformer. We also implement a complex-valued attention module for the output channels of the neural beamformer to utilize each other's in-formation, thereby preventing the final outputs from information loss. With the above innovations, our proposed model achieves 9.7% and 11.1% relative WER reductions over a DNN-MVDR baseline on the CHiME4 dataset and a dataset simulated using the Librispeech corpus. Gaofeng Cheng, Pengyuan Zhang, Yonghong Yan 0002 |
ASRU | 5 |
| 2021 | SI-Net: Multi-Scale Context-Aware Convolutional Block for Speaker VerificationabstractUtilizing multi-scale information adequately is essential for building a high-performance speaker verification (SV) system. Biological research shows that the human auditory system employs a multi-timescale processing mode to extract information and has a mechanism of integrating multi-scale information to encode sound information. Inspired by this, we propose a novel block, named Split-Integration (SI), to explore multi-scale context-aware feature learning at a granular level for speaker verification. Our model involves a pair of operations, (i) multi-scale split, which is designed to imitate the multi-timescale processing mode, extracting multi-scale features by grouping and stacking different sizes of filters, and (ii) dynamic integration, which aims at reflecting analogy with the fusion mechanism, introducing KL divergence to measure the complementarity between multi-scale features such that the model fully integrates multi-scale features and produces better speaker-discriminative representation. Experiments are conducted on Voxceleb and Speakers in the Wild(SITW) datasets. Results demonstrate that our approach achieves a relative 10%–20% improvement on equal error rate (EER) over a strong baseline in the SV task. Ce Fang, Runqiu Xiao, Yonghong Yan 0002 |
ASRU | 5 |
| 2021 | Using Cognitive Interest Graph and Knowledge-activated Attention for Learning Resource RecommendationabstractRecently, a number of deep-learning based approaches have proved effective for personalized learning resource recommendation (PLR), but they ignore the impact of complex transitions in a learner’s interests and their cognitive ability. This undermines the accomplishment of personalized learning. This paper proposes a new method for learning session-based recommendation which uses a graph neural network (GNNs) and a cognitive interest graph that, for brevity, is referred to as CIGNN. It models the click sequence and exercise sequence in a learning session separately as a session graph and a cognitive graph, then creates a composite cognitive interest graph. CIGNN also uses a novel knowledge-activated attention mechanism that makes uses of the knowledge mastery of learners and their past behavior to actively adapt to their learning interests, so that representations of their preferences change as they acquire knowledge. A further novel feature of CIGNN is its use of an interest-aware semantic graph attention network to extract semantic information regarding different types of learning interests by drawing on meta-paths of different importance. Extensive experiments were conducted using real-world data and the results show that CIGNN can outperform existing baseline approaches for PLR-related tasks. We also provide some case studies that illustrate how CIGNN can help learners to master knowledge. Jianzong Kuang, Wang Li 0007, Yonghong Yan 0002 |
COMPSAC | 4 |
| 2021 | History Utterance Embedding Transformer LM for Speech RecognitionabstractHistory utterances contain rich contextual information; however, better extracting information from the history utterances and using it to improve the language model (LM) is still challenging. In this paper, we propose the history utterance embedding Transformer LM (HTLM), which includes an embedding generation network for extracting contextual information contained in the history utterances and a main Transformer LM for current prediction. In addition, the two-stage attention (TSA) is proposed to encode richer contextual information into the embedding of history utterances (h-emb) while supporting GPU parallel training. Furthermore, we combine the extracted h-emb and embedding of current utterance (c-emb) through the dot-product attention and a fusion method for HTLM's current prediction. Experiments are conducted on the HKUST dataset and achieve a 23.4% character error rate (CER) on the test set. Compared with the baseline, the proposed method yields 12.86 absolute perplexity reduction and 0.8% absolute CER reduction. Keqi Deng, Gaofeng Cheng, Haoran Miao, Pengyuan Zhang, Yonghong Yan 0002 |
ICASSP | 5 |
| 2021 | Pre-Training Transformer Decoder for End-to-End ASR Model with Unpaired Text DataabstractThis paper presents a method to pre-train transformer-based encoder-decoder automatic speech recognition (ASR) models using sufficient target-domain text. During pre-training, we train the transformer decoder as a conditional language model with empty or artifical states, rather than the real encoder states. By this pre-training strategy, the decoder can learn how to generate grammatical text sequence before learning how to generate correct transcriptions. Contrast to other methods which utilize text only data to improve the ASR performance, our method does not change the network architecture of the ASR model or introduce extra component like text-to-speech (TTS) or text-to-encoder (TTE). Experimental results on LibriSpeech corpus show that the proposed method can relatively reduce the word error rate over 10%, using 960 hours transcriptions. Changfeng Gao, Gaofeng Cheng, Runyan Yang, Han Zhu 0004, Pengyuan Zhang, Yonghong Yan 0002 |
ICASSP | 6 |
| 2021 | Residual Echo and Noise Cancellation with Feature Attention Module and Multi-Domain Loss Function
Jianjun Gu 0005, Longbiao Cheng, Xingwei Sun, Yonghong Yan 0002 |
Interspeech | 5 |
| 2021 | Incorporating Cross-Speaker Style Transfer for Multi-Language Text-to-Speech
Zengqiang Shang, Pengyuan Zhang, Yonghong Yan 0002 |
Interspeech | 5 |
| 2021 | LinearSpeech: Parallel Text-to-Speech with Linear Complexity
Zengqiang Shang, Pengyuan Zhang, Yonghong Yan 0002 |
Interspeech | 5 |
| 2021 | A unified system for multilingual speech recognition and language identification
Pengyuan Zhang, Yonghong Yan 0002 |
Speech Commun. | 4 |
| 2021 | FSCNet: Feature-Specific Convolution Neural Network for Real-Time Speech EnhancementabstractIn recent years, convolutional neural networks (CNNs) have been widely exploited in deep neural network (DNN)-based speech enhancement methods. However, the representation power of CNNs for speech modeling is limited because of the spatial-agnostic convolution kernels. This letter proposes a novel feature-specific convolution neural network (FSCNet) for real-time speech enhancement. In FSCNet, the encoder and decoder are adopted for forward and inverse feature space transformation, respectively. The denoising module based on the feature-specific convolution (FSC) is employed to enhance the generated deep features. Leveraging the long-term global contexts and considering the importance of each feature channel for speech modeling, the convolution kernels of FSC are dynamically parameterized in each time-frequency location. A function-constrained loss is further proposed to train the FSCNet, ensuring the encoder, denoising modules and decoder can function as expected. Experimental results show that the proposed FSCNet outperforms the state-of-the-art denoising algorithms in terms of five objective evaluation metrics and model size. Longbiao Cheng, Yonghong Yan 0002 |
IEEE Signal Process. Lett. | 3 |
| 2021 | Estimation Reliability Function Assisted Sound Source Localization With Enhanced Steering Vector Phase DifferenceabstractThe performance of the traditional direction-of-arrival (DOA) estimation algorithms greatly degrades in noisy and reverberant environments. Recently, deep learning has been applied to sound source localization and provided the substantial improvement in robustness for DOA estimation. In this paper, we propose a sound source localization approach using the deep learning-based steering vector phase difference enhancement. The steering vectors and their estimation reliability functions (ERFs) are first estimated under the guidance of the time-frequency masks that are predicted using deep neural network (DNN). The phase difference of the steering vectors is further enhanced with a second DNN model, which is trained with the ERF-weighted mean square error (MSE) loss. The DOA of the sound source is finally determined by the ERF-weighted histogram analysis. Experimental results with various types and levels of noise and various reverberant conditions show that the proposed approach outperforms the state-of-the-art sound source localization algorithms in utterance and frame-level DOA estimation. Longbiao Cheng, Xingwei Sun, Dingding Yao, Yonghong Yan 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2021 | Keyword Search Using Attention-Based End-to-End ASR and Frame-Synchronous Phoneme AlignmentsabstractAttention-based end-to-end (E2E) automatic speech recognition (ASR) architectures are now the state-of-the-art in terms of recognition performance. However, despite their effectiveness, they have not been widely applied in keyword search (KWS) tasks yet. In this paper, we propose the Att-E2E-KWS architecture, an attention-based E2E ASR framework for KWS that can afford accurate and reliable keyword retrieval results. First, we design a basic framework to carry out KWS based on attention-based E2E ASR. We adopt the connectionist temporal classification and attention (CTC/Att) joint E2E ASR architecture and exploit the spike posterior property of CTC to provide the keywords time stamps. Second, we introduce the frame-synchronous phonemes modeling and use the dynamic programming (DP) algorithm to provide alignments between E2E grapheme outputs and phoneme outputs. We call this alignment procedure dynamic time alignment (DTA), which can provide the proposed Att-E2E-KWS system with more accurate time stamps and reliable confidence scores. Third, we use the Transformer, a self-attention-based encoder-decoder neural network, in place of conventional recurrent neural networks in order to yield more parallelizable models and increased training speed. We conduct comprehensive experiments on English and Mandarin Chinese. To the best of our knowledge, this is the first practical Att-E2E-KWS framework, and experimental results on Switchboard and HKUST corpora show that our proposed Att-E2E-KWS systems significantly outperform the CTC E2E ASR based KWS baselines. Runyan Yang, Gaofeng Cheng, Haoran Miao, Ta Li, Pengyuan Zhang, Yonghong Yan 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2020 | Transformer-Based Online CTC/Attention End-To-End Speech Recognition ArchitectureabstractRecently, Transformer has gained success in automatic speech recognition (ASR) field. However, it is challenging to deploy a Transformer-based end-to-end (E2E) model for online speech recognition. In this paper, we propose the Transformer-based online CTC/attention E2E ASR architecture, which contains the chunk self-attention encoder (chunk-SAE) and the monotonic truncated attention (MTA) based self-attention decoder (SAD). Firstly, the chunk-SAE splits the speech into isolated chunks. To reduce the computational cost and improve the performance, we propose the state reuse chunk-SAE. Sencondly, the MTA based SAD truncates the speech features monotonically and performs attention on the truncated features. To support the online recognition, we integrate the state reuse chunk-SAE and the MTA based SAD into online CTC/attention architecture. We evaluate the proposed online models on the HKUST Mandarin ASR benchmark and achieve a 23.66% character error rate (CER) with a 320 ms latency. Our online model yields as little as 0.19% absolute CER degradation compared with the offline baseline, and achieves significant improvement over our prior work on Long Short-Term Memory (LSTM) based online E2E models. Haoran Miao, Gaofeng Cheng, Changfeng Gao, Pengyuan Zhang, Yonghong Yan 0002 |
ICASSP | 5 |
| 2020 | Improving generative adversarial networks for speech enhancement through regularization of latent representations
Risheng Xia, Yonghong Yan 0002 |
Speech Commun. | 5 |
| 2020 | Online Hybrid CTC/Attention End-to-End Automatic Speech Recognition ArchitectureabstractRecently, there has been increasing progress in end-to-end automatic speech recognition (ASR) architecture, which transcribes speech to text without any pre-trained alignments. One popular end-to-end approach is the hybrid Connectionist Temporal Classification (CTC) and attention (CTC/attention) based ASR architecture, which utilizes the advantages of both CTC and attention. The hybrid CTC/attention ASR systems exhibit performance comparable to that of the conventional deep neural network (DNN)/ hidden Markov model (HMM) ASR systems. However, how to deploy hybrid CTC/attention systems for online speech recognition is still a non-trivial problem. This article describes our proposed online hybrid CTC/attention end-to-end ASR architecture, which replaces all the offline components of conventional CTC/attention ASR architecture with their corresponding streaming components. Firstly, we propose stable monotonic chunk-wise attention (sMoChA) to stream the conventional global attention, and further propose monotonic truncated attention (MTA) to simplify sMoChA and solve the training-and-decoding mismatch problem of sMoChA. Secondly, we propose truncated CTC (T-CTC) prefix score to stream CTC prefix score calculation. Thirdly, we design dynamic waiting joint decoding (DWJD) algorithm to dynamically collect the predictions of CTC and attention in an online manner. Finally, we use latency-controlled bidirectional long short-term memory (LC-BLSTM) to stream the widely-used offline bidirectional encoder network. Experiments with LibriSpeech English and HKUST Mandarin tasks demonstrate that, compared with the offline CTC/attention model, our proposed online CTC/attention model improves the real time factor in human-computer interaction services and maintains its performance with moderate degradation. To the best of our knowledge, this is the first work to provide the full-stack online solution for CTC/attention end-to-end ASR architecture. Haoran Miao, Gaofeng Cheng, Pengyuan Zhang, Yonghong Yan 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | A Model Compression Method With Matrix Product Operators for Speech EnhancementabstractThe deep neural network (DNN) based speech enhancement approaches have achieved promising performance. However, the number of parameters involved in these methods is usually enormous for the real applications of speech enhancement on the device with the limited resources. This seriously restricts the applications. To deal with this issue, model compression techniques are being widely studied. In this paper, we propose a model compression method based on matrix product operators (MPO) to substantially reduce the number of parameters in DNN models for speech enhancement. In this method, the weight matrices in the linear transformations of neural network model are replaced by the MPO decomposition format before training. In experiment, this process is applied to the causal neural network models, such as the feedforward multilayer perceptron (MLP) and long short-term memory (LSTM) models. Both MLP and LSTM models with/without compression are then utilized to estimate the ideal ratio mask for monaural speech enhancement. The experimental results show that our proposed MPO-based method outperforms the widely-used pruning method for speech enhancement under various compression rates, and further improvement can be achieved with respect to low compression rates. Our proposal provides an effective model compression method for speech enhancement, especially in cloud-free application. Xingwei Sun, Ze-Feng Gao, Zhong-Yi Lu, Yonghong Yan 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2019 | An Audio Scene Classification Framework with Embedded Filters and a DCT-based Temporal ModuleabstractDeep convolutional neural network (DCNN) has recently improved the performance of acoustic scene classification. However, the input features of the network are usually based on predefined hand-tailored filters, which may not apply to the specific tasks. To overcome this, we propose a hybrid framework that jointly trains the front-end filters and the back-end DCNN. Also, a novel temporal module based on the discrete cosine transform (DCT) is inserted after the high-level feature map of the network, thus enabling us to utilize time information without a reduction of training samples. Our single system, composed of the fine-tuned wavelet front-end and the DCNN back-end, with the integrated DCT-based temporal module, has achieved an accuracy of 79.20% in the evaluation set in DCASE17, gaining around 3% and 8% accuracy improvement compared with scalogram-DCNN and FBank-DCNN systems, respectively. Hangting Chen, Pengyuan Zhang, Yonghong Yan 0002 |
ICASSP | 3 |
| 2019 | Self-attention Based Prosodic Boundary Prediction for Chinese Speech SynthesisabstractPredicting prosodic boundaries from input text plays an important role in Chinese text-to-speech (TTS) system, which directly influences the naturalness and intelligibility of synthesized speech. In this paper, we propose to combine self-attention with multitask learning for prosodic boundary prediction. Self-attention is used to capture the dependency between two arbitrary characters in the input sentence, while multitask learning models the relationships between prosodic boundaries and lexicon words by setting word segmentation as an auxiliary task. The proposed method can generate prosodic boundary labels directly from Chinese characters and achieve the whole process end-to-end. Experimental results show the effectiveness of our proposed model and prove that the performance can be further improved by pretraining the model with extra word segmentation data. Chunhui Lu, Pengyuan Zhang, Yonghong Yan 0002 |
ICASSP | 3 |
| 2019 | A Deep Learning Based Binaural Speech Enhancement Approach with Spatial Cues PreservationabstractThe studies of binaural hearing indicated considerable benefits of the spatial information of sound sources in speech understanding in noise. In this paper, we propose a binaural speech enhancement approach based on deep neural network. In this approach, the signals at the left and right channels are regarded as the real and imaginary parts of a monaural complex signal, a complex ideal ratio mask is accordingly introduced and then further estimated using the complex deep neural network, followed by applying to the monaural complex signal. Experimental results showed that the suggested binaural speech enhancement approach is able to effectively suppress multiple interfering signals and preserve the binaural cues of target signal. Xingwei Sun, Risheng Xia, Yonghong Yan 0002 |
ICASSP | 4 |
| 2019 | Multiple Temporal Scales Based Speaker Embeddings Learning for Text-dependent Speaker RecognitionabstractTo extract high speaker-sensitive embeddings from deep neural networks is still a challenge in the field of speaker recognition. This paper proposes a novel network that learns speaker embeddings from multiple temporal scales. This idea comes from the recent biological research that the human auditory system has a mechanism of fusing multi-timescale information together to encode sound information. A two-pathway neural network is presented, in which one pathway focuses on short-time (or local) traits and the other focuses on long-range (or global) scale. Both traits are fused into one feature vector and the utterance-level speaker embeddings are extracted from these features. Experimental results show that different timescale traits can complement each other. And their fusion, which refer to as t-vector, outperforms i-vector and other deep embeddings. Moreover, with the end-to-end training, t-vectors can obtain excellent performance even using simple scoring approach like cosine distance. Yonghong Yan 0002 |
ICASSP | 4 |
| 2019 | A Subband Energy Modification Method for Elevation Control in Median PlaneabstractElevation perception is crucial for binaural reproduction. A recent study proposed an elevation control method by modifying the energy of HRTFs in each auditory scale subband, such as the ERB and Mel subband. However, this subband division is designed based on auditory excitation patterns and may not be consistent with the elevation localization cues. To this end, this study proposes a novel subband division strategy which emphasizes the physiological information involved in elevation localization based on a statistical analysis of the HRTF. Then, the elevation controlled HRTFs are constructed by modifying the energy of the HRTF magnitudes in each subband. Results of the listening test demonstrate that our method with the proposed subband division strategy outperforms the method with ERB scale subdivision in terms of the accuracy for controlling the perceived elevation of sound image. Dingding Yao, Huaxing Xu, Risheng Xia, Yonghong Yan 0002 |
ICASSP | 5 |
| 2019 | Target Speaker Recovery and Recognition Network with Average x-Vector and Global Training
Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 3 |
| 2019 | Character-Aware Sub-Word Level Language Modeling for Uyghur and Turkish ASR
Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 4 |
| 2019 | Online Hybrid CTC/Attention Architecture for End-to-End Speech Recognition
Haoran Miao, Gaofeng Cheng, Pengyuan Zhang, Ta Li, Yonghong Yan 0002 |
INTERSPEECH | 5 |
| 2019 | A New Time-Frequency Attention Mechanism for TDNN and CNN-LSTM-TDNN, with Application to Language Identification
Xiaoxiao Miao, Ian McLoughlin 0001, Yonghong Yan 0002 |
INTERSPEECH | 3 |
| 2019 | Speaker-Invariant Feature-Mapping for Distant Speech Recognition via Adversarial Teacher-Student Learning
Long Wu, Hangting Chen, Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 5 |
| 2019 | Multi-Accent Adaptation Based on Gate MechanismabstractWhen only a limited amount of accented speech data is available, to promote multi-accent speech recognition performance, the conventional approach is accent-specific adaptation, which adapts the baseline model to multiple target accents independently. To simplify the adaptation procedure, we explore adapting the baseline model to multiple target accents simultaneously with multi-accent mixed data. Thus, we propose using accent-specific top layer with gate mechanism (AST-G) to realize multi-accent adaptation. Compared with the baseline model and accent-specific adaptation, AST-G achieves 9.8% and 1.9% average relative WER reduction respectively. However, in real-world applications, we can't obtain the accent category label for inference in advance. Therefore, we apply using an accent classifier to predict the accent label. To jointly train the acoustic model and the accent classifier, we propose the multi-task learning with gate mechanism (MTL-G). As the accent label prediction could be inaccurate, it performs worse than the accent-specific adaptation. Yet, in comparison with the baseline model, MTL-G achieves 5.1% average relative WER reduction. Han Zhu 0004, Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 4 |
| 2019 | Tailoring an Interpretable Neural Language ModelabstractNeural networks have shown great potential in language modeling. Currently, the dominant approach to language modeling is based on recurrent neural networks (RNNs) and convolutional neural networks (CNNs). Nonetheless, it is not clear why RNNs and CNNs are suitable for the language modeling task since these neural models are lack of interpretability. The goal of this paper is to tailor an interpretable neural model as an alternative to RNNs and CNNs for the language modeling task. This paper proposes a unified framework for language modeling, which can partly interpret the rationales behind existing language models (LMs). Based on the proposed framework, an interpretable neural language model (INLM) is proposed, including a tailored architectural structure and a tailored learning method for the language modeling task. The proposed INLM can be approximated as a parameterized auto-regressive moving average model and provides interpretability in two aspects: component interpretability and prediction interpretability. Experiments demonstrate that the proposed INLM outperforms some typical neural LMs on several language modeling datasets and on the switchboard speech recognition task. Further experiments also show that the proposed INLM is competitive with the state-of-the-art long short-term memory LMs on the Penn Treebank and WikiText-2 datasets. Pengyuan Zhang, Yonghong Yan 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | A Deep Neural Network Based Method of Source Localization in a Shallow Water EnvironmentabstractThis paper applies deep neural network (DNN) to source localization in a shallow water environment because of its powerful modeling capability and the little dependence on the prior knowledge of environmental parameters. The classical two-stage scheme is adopted, in which feature extraction and DNN analysis are independent steps. It firstly extracts the input feature from the observed signal received by underwater hydrophones. The eigenvectors associated with the modal signal space are decomposed from the covariance matrices of the data field at different frequencies, which are used as the input feature of DNN. The time delay neural network (TDNN) is exploited to model the long term feature representation and construct the regression model. The output is the source range-depth estimate. Several experiments using simulation and experimental data are conducted to evaluate the performance of the proposed method. The results demonstrate the effectiveness and potential of DNN for source localization. Particularly, experiments show that simulation data can be merged to train a general model for experimental data when lacking of sufficient training data in real-world environment. Zhaoqiong Huang, Zaixiao Gong, Yonghong Yan 0002 |
ICASSP | 5 |
| 2018 | Semi-Supervised Learning with Deep Neural Networks for Relative Transfer Function Inverse RegressionabstractPrior knowledge of the relative transfer function (RTF) is useful in many applications but remains little studied. In this paper, we propose a semi-supervised learning algorithm based on deep neural networks (DNNs) for RTF inverse regression, that is to generate the full-band RTF vector directly from the source-receiver pose (position and orientation). Two typical scenarios are discussed: training on labeled RTFs only, or on additional unlabeled RTFs. Both setups utilize the low-dimensional manifold property of RTF in stationary environments. With this property as an additional regularization term, a smooth mapping solution with respect to the manifold is obtained. Experimental simulations show that the proposed method achieves a lower mean prediction error than the free field model with few labeled RTFs, and the unlabeled RTFs are essential in improving the inverse regression performance. Yonghong Yan 0002, Emmanuel Vincent 0001 |
ICASSP | 3 |
| 2018 | On SDW-MWF and Variable Span Linear Filter with Application to Speech Recognition in Noisy EnvironmentsabstractNeural network based spectral mask estimation for acoustic beamforming, which consists of linear filtering and mask estimation, has shown to be a promising approach for robust speech recognition in noisy environments. Nevertheless, few improvements are made on the linear filtering. In this paper, we investigate the Speech Distortion Weighted Multichannel Wiener Filter (SDW-MWF) and the variable span linear filter, and prove that they can be linked by Generalized Eigenvalue Decomposition (GEVD) of the speech covariance matrix. The resulting GEVD based SDW-MWF largely reduces the word error rate and even achieves competitive recognition performance with the state-of-the-art generalized eigenvalue beamformer. Furthermore, we found that the recent signal approximation is no better than mask approximation when combined in calculating the linear filter coefficients. Yonghong Yan 0002 |
ICASSP | 4 |
| 2018 | Improving Multichannel Speech Recognition with Generalized Cross Correlation Inputs and Multitask LearningabstractAcoustic signals from microphone arrays are used to improve performance in distant speech recognition due to the availability of spatial information. And multichannel automatic speech recognition (ASR) systems often separate speech enhancement module from acoustic modeling, which may be not optimal for improving recognition accuracy. In this work, we propose to improve multichannel speech recognition by supplying the generalized cross correlation (GCC) between microphones, which encodes spatial information, as input features to a long short-term memory (LSTM) acoustic model in parallel with the regular acoustic features. Moreover, multitask learning architecture is incorporated and shows its ability to improve the robustness of the model. We performed experiments on the AMI and ICSI meeting corpora, with results indicating that the proposed model outperforms the model trained directly on the concatenation of multiple microphone outputs and the model trained on a beamformed channel. Pengyuan Zhang, Yonghong Yan 0002 |
ICASSP | 4 |
| 2018 | Deep Convolutional Neural Network with Scalogram for Audio Scene Modeling
Hangting Chen, Pengyuan Zhang, Haichuan Bai, Qingsheng Yuan, Xiuguo Bao, Yonghong Yan 0002 |
INTERSPEECH | 6 |
| 2018 | Output-Gate Projected Gated Recurrent Unit for Speech Recognition
Gaofeng Cheng, Daniel Povey, Sanjeev Khudanpur, Yonghong Yan 0002 |
INTERSPEECH | 6 |
| 2018 | Investigation on the Combination of Batch Normalization and Dropout in BLSTM-based Acoustic Modeling for ASR
Gaofeng Cheng, Fengpei Ge, Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 5 |
| 2018 | Cross-Lingual Multi-Task Neural Architecture for Spoken Language Understanding
Yujiang Li, Xuemin Zhao, Weiqun Xu, Yonghong Yan 0002 |
INTERSPEECH | 4 |
| 2018 | Multi-talker Speech Separation Based on Permutation Invariant Training and Beamforming
Risheng Xia, Yonghong Yan 0002 |
INTERSPEECH | 5 |
| 2018 | Improving Language Modeling with an Adversarial Critic for Automatic Speech Recognition
Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 3 |
| 2018 | Discriminating between Similar Languages on Imbalanced Conversational Texts
Junqing He, Xuemin Zhao, Yonghong Yan 0002 |
LREC | 5 |
| 2018 | Improved Conditional Generative Adversarial Net Classification For Spoken Language RecognitionabstractRecent research on generative adversarial nets (GAN) for language identification (LID) has shown promising results. In this paper, we further exploit the latent abilities of GAN networks to firstly combine them with deep neural network (DNN)-based i-vector approaches and then to improve the LID model using conditional generative adversarial net (cGAN) classification. First, phoneme dependent deep bottleneck features (DBF) combined with output posteriors of a pre-trained DNN for automatic speech recognition (ASR) are used to extract i-vectors in the normal way. These i-vectors are then classified using cGAN, and we show an effective method within the cGAN to optimize parameters by combining both language identification and verification signals as supervision. Results show firstly that cGAN methods can significantly outperform DBF DNN i-vector methods where 49-dimensional i-vectors are used, but not where 600-dimensional vectors are used. Secondly, training a cGAN discriminator network for direct classification has further benefit for low dimensional i-vectors as well as short utterances with high dimensional i-vectors. However, incorporating a dedicated discriminator network output layer for classification and optimizing both classification and verification loss brings benefits in all test cases. Xiaoxiao Miao, Ian McLoughlin 0001, Shengyu Yao, Yonghong Yan 0002 |
SLT | 4 |
| 2018 | Rank-1 constrained Multichannel Wiener Filter for speech recognition in noisy environments
Emmanuel Vincent 0001, Romain Serizel, Yonghong Yan 0002 |
Comput. Speech Lang. | 4 |
| 2017 | Deep neural network based wake-up-word speech recognition with two-stage detectionabstractThis paper presents a novel far-field voice trigger algorithm utilizing DNN with the objective function of state-level minimum Bayes risk for training, customizing the decoding network to absorb the ambient noise and background speech. We adopt a two-stage classification strategy to integrate the phonetic knowledge and model-based classification into detecting wake-up words. Experimental results of the online test show that it can provide a higher than 90% accuracy and meanwhile false alarms are less than once per nine hours in the noisy home environments where the sound pressure level is about 80dB. Fengpei Ge, Yonghong Yan 0002 |
ICASSP | 2 |
| 2017 | An Exploration of Dropout with LSTMs
Gaofeng Cheng, Vijayaditya Peddinti, Daniel Povey, Vimal Manohar, Sanjeev Khudanpur, Yonghong Yan 0002 |
INTERSPEECH | 6 |
| 2017 | Joint Training of Multi-Channel-Condition Dereverberation and Acoustic Modeling of Microphone Array Speech for Robust Distant Speech RecognitionabstractWe propose a novel data utilization strategy, called multichannel-condition learning, leveraging upon complementary information captured in microphone array speech to jointly train dereverberation and acoustic deep neural network (DNN) models for robust distant speech recognition. Experimental results, with a single automatic speech recognition (ASR) system, on the REVERB2014 simulated evaluation data show that, on 1-channel testing, the baseline joint training scheme attains a word error rate (WER) of 7.47%, reduced from 8.72% for separate training. The proposed multi-channel-condition learning scheme has been experimented on different channel data combinations and usage showing many interesting implications. Finally, training on all 8-channel data and with DNN-based language model rescoring, a state-of-the-art WER of 4.05% is achieved. We anticipate an even lower WER when combining more top ASR systems. Fengpei Ge, Kehuang Li, Bo Wu 0011, Sabato Marco Siniscalchi, Yonghong Yan 0002, Chin-Hui Lee 0001 |
INTERSPEECH | 5 |
| 2017 | Time Delay Histogram Based Speech Source Separation Using a Planar Array
Zhaoqiong Huang, Zhanzhong Cao, Dongwen Ying, Jielin Pan, Yonghong Yan 0002 |
INTERSPEECH | 5 |
| 2017 | Ideal Ratio Mask Estimation Using Deep Neural Networks for Monaural Speech Segregation in Noisy Reverberant Conditions
Yonghong Yan 0002 |
INTERSPEECH | 3 |
| 2017 | Attention-Based LSTM with Multi-Task Learning for Distant Speech Recognition
Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 3 |
| 2017 | Window-Dominant Signal Subspace Methods for Multiple Short-Term Speech Source LocalizationabstractSignal subspace has been widely exploited to localize multiple speech sources. However, most signal subspace methods cannot count the number of sources, and do not make use of speech sparsity in the frequency domain. This paper presents a grid search window-dominant signal subspace (GS-WDSS) method and a closed-form WDSS (CF-WDSS) method to localize short-term speech sources. Such methods are based upon the generalized sparsity assumption that each window containing some time-adjacent bins is dominated by one source, as opposed to the conventional assumption that each individual bin is dominated by one source. The generalized assumption enables the principal eigenvector of the spatial correlation matrix on each window to span the signal subspace of the window-dominant source. The direction-of-arrival (DOA) of the dominant source is estimated from the principal eigenvector. The DOAs and the number of sources are eventually summarized from the DOA histogram of all dominant sources. The conventional assumption is a special case of the generalized assumption. By using the generalized assumption, the performance in estimating DOAs of the window-dominant sources is significantly improved at the cost of acceptable masking effect. The superiority of the proposed methods is verified by simulated and real experiments. Dongwen Ying, Ruohua Zhou, Yonghong Yan 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2016 | Robust multiple speech source localization using time delay histogramabstractSpatial aliasing and spatial resolution are the two issues faced by most multiple speech source localization methods. The histogram of time delays is a simple but effective method to deal with these two issues on linear arrays. But few methods were capable of applying the time delay histogram to directional-of-arrivals (DOAs) estimation using a planar array. This paper proposes a novel method to estimate DOAs of multiple speech sources based on time delay histograms across all microphones of a planar array. The pairwise time delays of different sources are firstly obtained from each time delay histogram, and then, the time delays are identified with variant speech sources. Eventually, the DOA of each source is estimated by regression over its associated time delays. We conducted some experiments in both simulated and real environments to evaluate the proposed method using an eight-element circular array. The experimental results confirmed not only its high computational efficiency, but also its superiority in spatial resolution and spatial anti-aliasing. Zhaoqiong Huang, Ge Zhan, Dongwen Ying, Yonghong Yan 0002 |
ICASSP | 4 |
| 2016 | Effective utilization of multiple examples in query-by-example spoken term detectionabstractThis paper investigates the example utilization problem in query-by-example spoken term detection when multiple examples are provided for each query term. To achieve this goal, we propose three evaluation metrics to assess the quality of all the examples, namely posteriorgram stability score, pronunciation reliability score and local similarity score. We also present a clustering based example generation approach to creating better examples based on the original ones. Experiments conducted on a telephone speech corpus shows that it is better to use several representative examples selected by the quality assessment process than to simply use all the examples. Furthermore, even better results can be obtained if the generated examples are used. Yonghong Yan 0002 |
ICASSP | 3 |
| 2016 | Enhancing Link Prediction Using Gradient Boosting Features
Taisong Li, Manshu Tu, Yonghong Yan 0002 |
ICIC (2) | 5 |
| 2016 | Adaptive Group Sparsity for Non-Negative Matrix Factorization with Application to Unsupervised Source Separation
Xiaofei Wang 0007, Qiang Fu 0001, Yonghong Yan 0002 |
INTERSPEECH | 5 |
| 2016 | A DNN-HMM Approach to Non-Negative Matrix Factorization Based Speech Enhancement
Xiaofei Wang 0007, Qiang Fu 0001, Yonghong Yan 0002 |
INTERSPEECH | 5 |
| 2016 | An unsupervised vocabulary selection technique for Chinese automatic speech recognitionabstractThe vocabulary is a vital component of automatic speech recognition(ASR) systems. For a specific Chinese speech recognition task, using a large general vocabulary not only leads to a much longer time to decode, but also hurts the recognition accuracy. In this paper, we proposed an unsupervised algorithm to select task-specific words from a large general vocabulary. The out-of-vocabulary(OOV) rate is a measure of vocabularies, and it is related to the recognition accuracy. However, it is hard to compute OOV rate for a Chinese vocabulary, since OOVs are often segmented into single Chinese characters and most Chinese vocabularies contain all the single Chinese characters. To deal with this problem, we proposed a novel method to estimate the OOV rate of Chinese vocabularies. In experiments, we found that our estimated OOV rate is related to the character error rate(CER) of recognition. Our proposed vocabulary selection method provided both the lowest OOV rate and CER on two Chinese conversational telephone speech(CTS) evaluation sets compared to the general vocabulary and frequency based vocabulary selection method. In addition, our proposed method significantly reduced the size of the language model(LM) and the corresponding weighted finite state transducer(WFST) network, which led to a more efficient decoding. Pengyuan Zhang, Ta Li, Yonghong Yan 0002 |
SLT | 4 |
| 2016 | Structural Optimization and Online Evolutionary Learning for Spoken Dialog ManagementabstractDesigning dialog management (DM) policies that are robust to environmental noises is a nontrivial task. Approaches based on reinforcement learning (RL) are popular in academia and have been empirically shown to exhibit much better performance than handcrafted policies. However, the policies trained using RL are mostly incomprehensible, thus limiting the deployments for commercial applications. Policy optimization using genetic algorithm (GA) is a relatively new approach to spoken DM. The most notable advantage of this approach is that the trained policies can be directly interpreted by human experts. In this letter, we make several contributions to the GA-based framework. First, a structural policy learning procedure is presented. Second, a new fitness estimation method based on fitted policy evaluation is proposed. Finally, combining with these methods, an online evolutionary policy learning algorithm is designed which is much more data efficient than direct policy search using Monte Carlo simulations. These proposed approaches are empirically evaluated and compared with several state-of-the-art methods in a simulated environment. The experiments show favorable results for our approach. Yonghong Yan 0002 |
IEEE Signal Process. Lett. | 2 |
| 2016 | Cross Array and Rank-1 MUSIC Algorithm for Acoustic Highway Lane DetectionabstractA vehicle emits sound as it travels along the road, which can be used as a kind of robust feature for traffic monitoring. In this paper, an acoustic-based lane detection approach is introduced for a multilane traffic monitoring system. First, a microphone array is designed according to a typical Chinese highway configuration. The design is based on the cross-array structure, and the cross-correlation matrix from the two subarrays in the selected working frequency band is calculated for the subsequent traffic monitoring operations. Then, a cross section across the road is constructed by beamforming, in which the single-source assumption can be applied, and the passing vehicle azimuth is detected by the proposed rank-1 Multiple Signal Classification (MUSIC) algorithm. Finally, a Parzen-window-based technique is proposed to estimate the vehicle azimuth probability density function (pdf) from the individual azimuth observations. Lane centers and boundaries can be revealed from the peak and valley patterns of the estimated pdf. A prototype traffic monitoring system is developed, and several lane detection approaches are compared in both simulated and real-world environments in the developed system framework. The experimental results exhibit the efficiency of the proposed approach. Yueyue Na, Yanmeng Guo, Qiang Fu 0001, Yonghong Yan 0002 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2015 | Optimizing human-interpretable dialog management policy using genetic algorithmabstractAutomatic optimization of spoken dialog management policies that are robust to environmental noise has long been the goal for both academia and industry. Approaches based on reinforcement learning have been proved to be effective. However, the numerical representation of dialog policy is human-incomprehensible and difficult for dialog system designers to verify or modify, which limits its practical application. In this paper we propose a novel framework for optimizing dialog policies specified in domain language using genetic algorithm. The human-interpretable representation of policy makes the method suitable for practical employment. We present learning algorithms using user simulation and real human-machine dialogs respectively. Empirical experimental results are given to show the effectiveness of the proposed approach. Weiqun Xu, Yonghong Yan 0002 |
ASRU | 3 |
| 2015 | Two-stage ASGD framework for parallel training of DNN acoustic models using EthernetabstractDeep neural networks have shown significant improvements on acoustic modelling, pushing state-of-the-art performance in large vocabulary continuous speech recognition (LVCSR) tasks. However, training DNNs is very time-consuming on scaled data. In this paper, a data-parallel method, namely two-stage ASGD, is proposed. Two-stage ASGD is based on asynchronous stochastic gradient descent (ASGD) paradigm and is tuned for GPU-equipped computing cluster connected by 10Gbit/s Ethernet other than Infiniband. Several techniques, such as hierarchical learning rate control, double-buffering and order-locking are applied to optimise the communication-to-transmission ratio. The proposed framework is evaluated by training a DNN with 29.5M parameters using a 500-hours Chinese continuous telephone speech data set. By using 4 computer nodes and 8 GPU devices (2 devices used in each node), a 5.9 times acceleration is obtained over a single GPU with acceptable loss of accuracy (0.5% in average). A comparative experiment is done to compare the proposed two-stage ASGD with the parallel DNN training systems reported in prior work. Xingyu Na, Jielin Pan, Yonghong Yan 0002 |
ASRU | 5 |
| 2015 | How to Detect Communities in Large Networks
Yasong Jiang, Shengxiang Gao, Yonghong Yan 0002 |
ICIC (1) | 6 |
| 2015 | Robust localization of single sound source based on phase difference regressionabstractPhase difference regression (PDR) was widely utilized to estimate Direction-of-Arrival (DOA) for linear arrays because of its high time resolution and high computational efficiency. However, conventional regression methods were seldom reported to estimate DOA using planar arrays. This paper proposes a regression method to derive DOA from all phase differences on all frequencies for a planar array. The DOA is represented as the function of the array topology and phase differences between all microphones. Moreover, the proposed method considers another two problems that were often ignored by most regression methods. One is the problem about the period of phase difference in the regression cost function. The other is the signal enhancement that can effectively suppress the acoustic interference. We conducted some experiments in simulated environment to evaluate the proposed method using a 9-element circular array. The experimental results confirmed its superiority in both computational efficiency and robustness. Index Terms: Planar array, phase difference regression, signal enhancement, sound source localization. Zhaoqiong Huang, Ge Zhan, Dongwen Ying, Yonghong Yan 0002 |
INTERSPEECH | 4 |
| 2015 | Spectrographic speech mask estimation using the time-frequency correlation of speech presence
Ge Zhan, Zhaoqiong Huang, Dongwen Ying, Jielin Pan, Yonghong Yan 0002 |
INTERSPEECH | 5 |
| 2015 | A reverberation robust target speech detection method using dual-microphone in distant-talking scene
Xiaofei Wang 0007, Yanmeng Guo, Chao Wu 0011, Qiang Fu 0001, Yonghong Yan 0002 |
Speech Commun. | 5 |
| 2015 | Phonotactic language recognition using dynamic pronunciation and language branch discriminative information
Xianliang Wang, Yulong Wan, Lin Yang 0016, Ruohua Zhou, Yonghong Yan 0002 |
Speech Commun. | 5 |
| 2014 | Language recognition system using language branch discriminative informationabstractThis paper presents our study of using language branch discriminative information effectively for language recognition. Language branch variability (LBV) method based on factor analysis techniques is proposed. In LBV method, language branch variability factor is obtained by concatenating low-dimensional factors in the language branch variability spaces. Language models are trained within language branches and between languages. Experiments on NIST 2011 Language Recognition Evaluation (LRE) 30s, 10s and 03s tasks show the proposed LBV method provides stable improvement compared to the state-of-art total variability (TV) approach. In 30-second task, it gains relative improvement by 14.6% in equal error rate (EER) and 12.9% in minimum decision cost value (minDCF), and in new metrics of NIST 2011 LRE, it leads to relative improvement of 7.2%-17.7%. Xianliang Wang, Yulong Wan, Lin Yang 0016, Ruohua Zhou, Yonghong Yan 0002 |
ICASSP | 5 |
| 2014 | A robust step-size control algorithm for frequency domain acoustic echo cancellationabstractThe presence of near-end interferences and echo path changes make it essential for an adaptive filter to vary its learning rate according to corresponding conditions. In this paper, a robust step-size control algorithm which is based on the optimization of the square of the bin-wise a posteriori error is proposed. To prevent the adaptive filter from diverging in the presence of interferences, constraints on the filter update are applied. The learning rate expression is derived and then we extend the method to multidelay block frequency domain adaptive filter (MDF) so as to meet the demand of low delay in practical application. An updating strategy for the constraints is proposed as well. Experiments are carried out to demonstrate the superiority of the proposed approach, especially in double-talk and echo path change situations. Index Terms: Acoustic echo cancellation, step-size control, robust filtering. Chao Wu 0011, Kaiyu Jiang, Yanmeng Guo, Qiang Fu 0001, Yonghong Yan 0002 |
INTERSPEECH | 5 |
| 2014 | Direction-of-arrival estimation of multiple speakers using a planar array
Dongwen Ying, Ruohua Zhou, Jielin Pan, Yonghong Yan 0002 |
INTERSPEECH | 5 |
| 2014 | Markovian Discriminative Modeling for Dialog State TrackingabstractDiscriminative dialog state tracking has become a hot topic in dialog research com-munity recently. Compared to genera-tive approach, it has the advantage of be-ing able to handle arbitrary dependent fea-tures, which is very appealing. In this paper, we present our approach to the DSTC2 challenge. We propose to use dis-criminative Markovian models as a natu-ral enhancement to the stationary discrim-inative models. The Markovian structure allows the incorporation of ‘transitional’ features, which can lead to more effi-ciency and flexibility in tracking user goal changes. Results on the DSTC2 dataset show considerable improvements over the baseline, and the effects of the Markovian dependency is tested empirically. 1 Weiqun Xu, Yonghong Yan 0002 |
SIGDIAL Conference | 3 |
| 2014 | Markovian discriminative modeling for cross-domain dialog state trackingabstractDialog state tracking (DST), which infers user goals in the presence of noise, is important for spoken dialog systems. Recently it has attracted a lot of attention in the dialog research community. Several new tracking approaches have been proposed, especially in the series of DST Challenges (DSTC). But the problem of cross-domain generalization, i.e., whether trackers designed for one domain will perform similarly well on other domains, is still an open issue. This becomes the focus in DSTC3. To tackle this problem, we adopt domain-independent models and features. We extend our Markovian discriminative model with a joint feature space for effective parameter sharing, so as to accommodate the domain mismatch. In addition, a new two-step training procedure is used to mitigate the `label over-coupling' problem brought by the Markovian structure. When evaluated on the DSTC3 data, our system outperforms all the baselines. Weiqun Xu, Yonghong Yan 0002 |
SLT | 3 |
| 2014 | Acoustic Echo Control with Frequency-Domain Stage-Wise RegressionabstractThis letter introduces frequency domain stage-wise regression to acoustic echo control. By approximating the echo path as concatenate short segments in frequency domain, simple regression in the frequency domain can be carried out to estimate the echo contributed by consecutive far-end signal blocks stage by stage. A non-stationarity controlled smoothing factor is proposed alongside the regression procedure to mitigate the increasing variance of estimation when no significant echo but only near-end background noise is present. Experiments are carried out to demonstrate the superiority of the proposed approach, especially in unstable environment. Kaiyu Jiang, Chao Wu 0011, Yanmeng Guo, Qiang Fu 0001, Yonghong Yan 0002 |
IEEE Signal Process. Lett. | 5 |
| 2013 | A novel discriminative method for pronunciation quality assessmentabstractThis paper presents a novel method for automatic pronunciation quality assessment. Unlike the traditional “Goodness of Pronunciation” (GOP) method, we judged utterance's pronunciation quality directly by a discriminative method. Under this novel framework, we also designed an algorithm to calculate the assessment confidence. We decoded the student's utterance for two passes. The first-pass decoding was just for getting the phone time points, and the second-pass decoding was for differentiating the pronunciation quality for each triphone. In the second-pass decoding, we used a specially trained acoustic model (AM), where the triphones in different pronunciation qualities were trained as different units. The confidence of the phone-level scoring was also calculated, and the low confidence phone-level scores were excluded in calculating the word-level score. The experimental results shows that the scoring performance was increased significantly compared to the traditional GOP method. Fuping Pan, Bin Dong 0003, Yonghong Yan 0002 |
ICASSP | 4 |
| 2013 | Effect of linguistic masker on the intelligibility of Mandarin sentences
Fei Chen 0011, Lena L. N. Wong, Yonghong Yan 0002 |
INTERSPEECH | 4 |
| 2013 | Comparative investigation of objective speech intelligibility prediction measures for noise-reduced signals in Mandarin and JapaneseabstractIn this paper, eight state-of-the-art objective speech intelligibility prediction measures are comparatively investigated for noisy signals before and after noise-reduction processing between Mandarin and Japanese. Clean speech signals (Chinese words and Japanese words) were first corrupted by three types of noise at two signal-to-noise ratios and then processed by normal-hearing listeners for recognition, whose intelligibility was subsequently predicted by objective measures. Further investigations were conducted for objective measures in predicting speech intelligibility of noise-reduced signals between subjective evaluation scores and objective prediction results, and of noisy signals before and after noise-reduction processing, in terms of correlation analysis and prediction errors. Results showed that the majority of objective measures behave differently for Mandarin and Japanese in predicting the subjective ratings, and the STOI measure consistently provided the best ability in predicting the effect on speech intelligibility of the noise-reduction processing for both Mandarin and Japanese. Fei Chen 0011, Masato Akagi, Yonghong Yan 0002 |
INTERSPEECH | 4 |
| 2013 | Prefix tree based n-best list re-scoring for recurrent neural network language model used in speech recognition system
Yujing Si, Ta Li, Jielin Pan, Yonghong Yan 0002 |
INTERSPEECH | 5 |
| 2013 | Discriminative pronunciation modeling based on minimum phone error training
Meixu Song, Jielin Pan, Yonghong Yan 0002 |
INTERSPEECH | 4 |
| 2013 | Dialog State Tracking using Conditional Random Fields
Weiqun Xu, Yonghong Yan 0002 |
SIGDIAL Conference | 4 |
| 2013 | Robust and Fast Localization of Single Speech Source Using a Planar ArrayabstractHeavy computational load and acoustic interferences are two major problems to speech source localization in real applications. Conventional methods can mitigate one problem, but deteriorate the other. This letter proposes an algorithm of direction-of-arrival (DOA) estimation, which is both computationally efficient and robust in the presence of acoustic interferences. The robustness is considered in two aspects. One is the eigenanalysis-based enhancement to reduce acoustic interferences such as noise and reverberation. The other is the coefficients that weight the pairwise time delays to mitigate the effect of delay outliers on DOA. The high computational efficiency is achieved by making use of a concave cost function, from which, the optimal estimate of DOA is given by a closed-form solution. The grid-search method often adopted in conventional algorithms is no longer used in this algorithm. We conduct some experiments in both simulated and real environments with a 9-element circular array. The proposed algorithm runs about ten times faster than Steered Response Power PHAse Transform (SRP-PHAT), and outperforms SRP-PHAT in terms of robustness. Dongwen Ying, Yonghong Yan 0002 |
IEEE Signal Process. Lett. | 2 |
| 2013 | Noise Estimation Using a Constrained Sequential Hidden Markov Model in the Log-Spectral DomainabstractThe temporal correlation of speech presence/absence is widely used in noise estimation. The most popular technique for exploiting temporal correlation is the smoothing of noisy spectra using a time-recursive filter, in which the forgetting factor is controlled by speech presence probability. However, this technique is not unified into a theoretical framework that enables optimal noise estimation. In theory, hidden Markov models (HMMs) are superior to this technique in modeling temporal correlation. HMMs can model a time sequence of presence/absence of speech signal as a dynamic process of the transition between speech and non-speech states. Moreover, a number of methods, such as maximum likelihood, are available for optimal estimation of HMM parameters. This paper presents a constrained sequential HMM for modeling the log-power sequence on each frequency band. The emission probability of each HMM state is represented by a Gaussian model. The Gaussian mean of the non-speech state is considered as the optimal estimate of noise logarithmic power. The HMM parameter set is sequentially estimated from one frame to another on the basis of maximum likelihood. The proposed method is compared with well-established algorithms through various experiments. Our method delivers more accurate results and does not rely on the assumption of the “non-speech signal onset” as do most algorithms. Dongwen Ying, Yonghong Yan 0002 |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | A two-microphone based voice activity detection for distant-talking speech in wide range of direction of arrivalabstractIn this paper, a two-microphone based voice activity detection (VAD) algorithm is proposed to detect the distant-talking speech coming randomly from a wide range of direction of arrival (DOA). The long-term information of inter-channel phase difference (LTIPD) is introduced as a target speech existence measure, which describes the concentration degree of DOA estimations on a sound source with harmonic structure. The proposed algorithm performs robustly on distant-talking speech recorded in several real environments. Yanmeng Guo, Qiang Fu 0001, Yonghong Yan 0002 |
ICASSP | 4 |
| 2012 | Evaluation of objective intelligibility prediction measures for noise-reduced signals in mandarinabstractIn this paper, the performance of eight state-of-the-art objective measures is evaluated in terms of predicting speech intelligibility in Mandarin of the processed signals by noise-reduction algorithms. The speech signals were first corrupted by three types of noises at two signal-to-noise ratios and subsequently processed by four classes of noise reduction algorithms, followed by objective intelligibility prediction. The subjective intelligibility ratings were obtained through a set of listening tests. Further investigation was conducted for objective measures in predicting speech intelligibility of noisy signals before and after noise-reduction processing in terms of correlation analysis and prediction errors. The analysis results reported here do provide valuable hints for analyzing and optimizing noise-reduction algorithms for Mandarin. Risheng Xia, Masato Akagi, Yonghong Yan 0002 |
ICASSP | 4 |
| 2012 | Factor analysis of Laplacian approach for speaker recognitionabstractIn this study, we introduce a new factor analysis of Laplacian approach to speaker recognition under the support vector machine (SVM) framework. The Laplacian-projected supervector from our proposed Laplacian approach, which finds an embedding that preserves local information by locality preserving projections (LPP), is believed to contain speaker dependent information. The proposed method was compared with the state-of-the-art total variability approach on 2010 National Institute of Standards and Technology (NIST) Speaker Recognition Evaluation (SRE) corpus. According to the compared results, our proposed method is effective. Jinchao Yang, Chunyan Liang, Lin Yang 0016, Hongbin Suo, Yonghong Yan 0002 |
ICASSP | 6 |
| 2012 | Noise estimation using a constrained sequential HMM IN log-spectral domainabstractHow to utilize the time correlation of speech/nonspeech presence is a crucial problem faced by noise estimators. The popular technique of exploiting such correlation is to smooth noisy spectra by using a temporal recursive filter with a time-varying smoothing factor. But this technique cannot warrant the statistical optimality. In theory, hidden Markov model (HMM) is more desirable than this technique. It can give an elaborate description of speech/nonspeech transition. Moreover, some theoretical frameworks, such as maximum likelihood (ML), are available for optimal estimation. This paper presents a constrained sequential HMM to model the time correlation of speech/nonspeech presence of an individual log-power sequence. Its parameter set is on-line adapted to varying signals based on a ML framework. We compared its performance with that of well-established algorithms by speech enhancement experiments. The results confirmed its promising performance. Dongwen Ying, Xugang Lu, Yonghong Yan 0002, Jianwu Dang 0001, Frank K. Soong |
ICASSP | 4 |
| 2012 | Speaker Verification Using Neighborhood Preserving Embedding
Chunyan Liang, Jinchao Yang, Lin Yang 0016, Yonghong Yan 0002 |
INTERSPEECH | 4 |
| 2012 | Discriminative Decision Function Based Scoring Method in Joint Factor Analysis for Speaker Verification
Chunyan Liang, Xiang Zhang 0014, Lin Yang 0016, Yonghong Yan 0002 |
INTERSPEECH | 4 |
| 2012 | A Initial Attempt on Task-Specific Adaptation for Deep Neural Network-based Large Vocabulary Continuous Speech Recognition
Yeming Xiao, Shang Cai, Jielin Pan, Yonghong Yan 0002 |
INTERSPEECH | 5 |
| 2012 | Sparse Probabilistic Linear Discriminant Analysis for Speaker Verification
Chunyan Liang, Lin Yang 0016, Yonghong Yan 0002 |
INTERSPEECH | 5 |
| 2012 | Maximum A Posteriori Linear Regression for language recognition
Jinchao Yang, Xiang Zhang 0014, Hongbin Suo, Jianping Zhang 0001, Yonghong Yan 0002 |
Expert Syst. Appl. | 6 |
| 2012 | A Novel Similarity Measure to Induce Semantic Classes and Its Application for Language Model Adaptation in a Dialogue System
Weiqun Xu, Yonghong Yan 0002 |
J. Comput. Sci. Technol. | 3 |
| 2011 | Robust understanding of spoken Chinese through character-based tagging and prior knowledge exploitationabstractRobustness is one of the most challenging issues for spoken language understanding (SLU). In this paper we studied the semantic understanding of Chinese spoken language for a voice search dialogue system. We first simplified the problem of semantic understanding into a named entity recognition (NER) task, which was further formulated as sequential tagging. We carried out experiments to opt for character over word as the tagging unit. Then two approaches were proposed to exploit prior knowledge - in the form of a domain lexicon - into the character-based tagging framework. One enriched tagger features by incorporating more formal lexical features with a domain lexicon. The other made plain use of domain entities by simply adding them to the training data. Experiment results show that both approaches are effective. The best performance is achieved by combining the above two complimentary approaches. By exploiting prior knowledge we improved the NER performance from 75.27 to 90.24 in F1score on a field test set using speech recognizer output. Weiqun Xu, Changchun Bao, Jielin Pan, Yonghong Yan 0002 |
ASRU | 5 |
| 2011 | Speaker Verification Using Sparse Representations on Total Variability i-vectorsabstractIn this paper, the sparse representation computed by l1-minimization with quadratic constraints is employed to model the i-vectors in the low dimensional total variability space af-ter performing the Within-Class Covariance Normalization and Linear Discriminate Analysis channel compensation. First, we propose the background normalized l2 residual as a scoring cri-terion. Second, we demonstrate that the Tnorm can be effi-ciently achieved by using the Tnorm data as the non-target sam-ples in the over-complete dictionary. Finally, by fusing with the conventional i-vector based support vector machine (SVM) and cosine distance scoring system, we demonstrate overall system performance improvement. Experimental results show that the proposed fusion system achieved 4.05 % (male) and 5.25 % (fe-male) equal error rate (EER) after Tnorm on the single-single multi-language handheld telephone task of NIST SRE 2008 and outperformed the SVM baseline by yielding 7.1 % and 4.9 % rel-ative EER reduction for the male and female tasks, respectively. Index Terms: speaker verification, sparse representation i-vector modeling Ming Li 0026, Xiang Zhang 0014, Yonghong Yan 0002, Shri Narayanan |
INTERSPEECH | 3 |
| 2011 | Towards precise and robust automatic synchronization of live speech and its transcripts
Jie Gao 0020, Qingwei Zhao, Yonghong Yan 0002 |
Speech Commun. | 3 |
| 2011 | Voice Activity Detection Based on an Unsupervised Learning FrameworkabstractHow to construct models for speech/nonspeech discrimination is a crucial point for voice activity detectors (VADs). Semi-supervised learning is the most popular way for model construction in conventional VADs. In this correspondence, we propose an unsupervised learning framework to construct statistical models for VAD. This framework is realized by a sequential Gaussian mixture model. It comprises an initialization process and an updating process. At each subband, the GMM is firstly initialized using EM algorithm, and then sequentially updated frame by frame. From the GMM, a self-regulatory threshold for discrimination is derived at each subband. Some constraints are introduced to this GMM for the sake of reliability. For the reason of unsupervised learning, the proposed VAD does not rely on an assumption that the first several frames of an utterance are nonspeech, which is widely used in most VADs. Moreover, the speech presence probability in the time-frequency domain is a byproduct of this VAD. We tested it on speech from TIMIT database and noise from NOISEX-92 database. The evaluations effectively showed its promising performance in comparison with VADs such as ITU G.729B, GSM AMR, and a typical semi-supervised VAD. Dongwen Ying, Yonghong Yan 0002, Jianwu Dang 0001, Frank K. Soong |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2010 | Automatic Synchronization of live speech and its Transcripts based on a frame-synchronous likelihood ratio testabstractIn this paper, we present our initial efforts in the task of Automatically Synchronizing live spoken Utterances with their Transcripts (textual contents) (ASUT) when the texts are known. We treat it as a online speech-text alignment problem. And it is further simplified into the problem of on-the-fly detecting of the end time of a spoken utterance given its textual content. A general framework called frame-synchronous likelihood ratio test (FS-LRT) procedure is proposed for this end time detection task and explored with the hidden Markov models (HMMs). The property of FS-LRT is studied empirically. Extensive experiments indicate that our proposed approach shows satisfying performance. In addition, FS-LRT has been successfully applied in a subtitling system for live broadcast news. Jie Gao 0020, Qingwei Zhao, Yonghong Yan 0002 |
ICASSP | 3 |
| 2010 | Improved modeling for F0 generation and V/U decision in HMM-based TTSabstractThe HMM-based TTS can produce a highly intelligible and decent quality voice. However, sometimes the synthesized speech exhibits perceptibly annoying glitches due to F0 extraction errors in the training data and voiced/unvoiced swapping errors in F0 generation. In the conventional MSD based F0 modeling [10], the dual but incompatible two probabilistic spaces, the continuous probability density for voiced observations or the discrete probability for unvoiced observations, prevent us from using likelihood based frame occupancy to alleviate the deteriorating effect of F0 extraction errors in training a more robust model for synthesis. In this paper, we propose a new approach to improved modeling the piece-wise continuous F0 trajectory and v/u decision for HMM-based TTS. Voicing strength, characterized by the normalized correlation coefficient magnitude calculated in F0 feature extraction, is used as an additional feature in F0 modeling and for v/u decision. Experimental results show the new approach to F0 modeling and generation outperforms MSD-HMM method and a newly proposed GTD-HMM method [9] significantly. The improvements are both objectively measurable and subjectively perceivable. Frank K. Soong, Yao Qian, Zhijie Yan, Jielin Pan, Yonghong Yan 0002 |
ICASSP | 6 |
| 2010 | Maximum a posteriori linear regression for speaker recognitionabstractRecently, using maximum likelihood linear regression (MLLR) transforms as the features for SVM based speaker recognition has been proposed. This can achieve performance comparable to that obtained with state-of-the-art approaches. In this paper, we focus on calculating the transforms based on a GMM universal background model (UBM). Rather than estimating the transforms using maximum likelihood criterion, we describe a new feature extraction technique for speaker recognition based on maximum a posteriori linear regression (MAPLR). This work is enriched by a proposed multi-class technique, which clusters the Gaussian mixtures into regression classes and estimates a different transform for each class. All the transforms of all the classes for a given utterance are concatenated into a supervector for SVM classification. Experiments on a NIST 2008 SRE corpus show that the speaker recognition system using MAPLR outperforms MLLR, and the multi-class approach can also bring significant gains for MAPLR system. Xiang Zhang 0014, Jianping Zhang 0001, Yonghong Yan 0002 |
ICASSP | 5 |
| 2010 | Speech enhancement using improved generalized sidelobe canceller in frequency domain with multi-channel postfiltering
Qiang Fu 0001, Yonghong Yan 0002 |
INTERSPEECH | 3 |
| 2010 | Speaker recognition using the resynthesized speech via spectrum modeling
Xiang Zhang 0014, Chuan Cao, Lin Yang 0016, Hongbin Suo, Jianping Zhang 0001, Yonghong Yan 0002 |
INTERSPEECH | 6 |
| 2009 | Nonnative speech recognition based on bilingual model modificationabstractThis paper presents a novel bilingual model modification approach to improve nonnative speech recognition accuracy when the variations of accented pronunciations occur. Each state of baseline nonnative acoustic model is modified with several candidate states from the auxiliary acoustic model, which is trained on speakers' mother language. State mapping criterion and n-best candidates are investigated, and different numbers of Gaussian mixtures of the auxiliary acoustic model are compared based on a grammar-constrained speech recognition system. Using this bilingual model modification approach, compared to the nonnative acoustic model which has already been well trained by adaptation technique MAP, the Phrase Error Rate further achieves a 5.83% relative reduction, while only a small relative increase on Real Time Factor occurs. Jielin Pan, Shui-duen Chan, Yonghong Yan 0002 |
FUZZ-IEEE | 4 |
| 2009 | Online detecting end times of spoken utterances for synchronization of live speech and its transcripts
Jie Gao 0020, Qingwei Zhao, Yonghong Yan 0002 |
INTERSPEECH | 3 |
| 2009 | A one-step tone recognition approach using MSD-HMM for continuous speech
Changliang Liu, Fengpei Ge, Fuping Pan, Bin Dong 0003, Yonghong Yan 0002 |
INTERSPEECH | 5 |
| 2009 | Tonal articulatory feature for Mandarin and its application to conversational LVCSR
Jielin Pan, Yonghong Yan 0002 |
INTERSPEECH | 3 |
| 2009 | Physiologically-inspired feature extraction for emotion recognitionabstractIn this paper, we proposed a new feature extraction method for emotion recognition based on the knowledge of the emotion production mechanism in physiology. It was reported by physiacoustist that emotional speech is differently encoded from the normal speech in terms of articulation organs and that emotion information in speech is concentrated in different frequencies caused by the different movements of organs [4]. To apply these findings, in this paper, we first quantified the distribution of speech emotion information along with each frequency band by exploiting the Fisher’s F-Ratio and mutual information techniques, and then proposed a non-uniform sub-band processing method which is able to extract and emphasize the emotion features in speech. These extracted features are finally applied to emotional recognition. Experimental results in speech emotion recognition showed that the extracted features using our proposed non-uniform sub-band processing outperform the traditional (MFCC) features, and the average error reduction rate amounts to 16.8 % for speech emotion recognition. Index Terms: speech emotion recognition, feature extraction, non-uniform sub-band. Yanqing Sun, Jianping Zhang 0001, Yonghong Yan 0002 |
INTERSPEECH | 5 |
| 2009 | Simultaneous Synchronization of Text and Speech for Broadcast News Subtitling
Jie Gao 0020, Qingwei Zhao, Ta Li, Yonghong Yan 0002 |
ISNN (3) | 4 |
| 2009 | An SVM-Based Mandarin Pronunciation Quality Assessment System
Fengpei Ge, Fuping Pan, Changliang Liu, Bin Dong 0003, Shui-duen Chan, Yonghong Yan 0002 |
ISNN (4) | 7 |
| 2009 | Improving Voice Search Using Forward-Backward LVCSR System Combination
Ta Li, Changchun Bao, Weiqun Xu, Jielin Pan, Yonghong Yan 0002 |
ISNN (4) | 5 |
| 2009 | Dynamic Multiple Pronunciation Incorporation in a Refined Search Space for Reading Miscue Detection
Changliang Liu, Fuping Pan, Fengpei Ge, Bin Dong 0003, Shuiduen Chen, Yonghong Yan 0002 |
ISNN (4) | 6 |
| 2009 | A Novel Fuzzy-Based Automatic Speaker Clustering Algorithm
Xiang Zhang 0014, Hongbin Suo, Qingwei Zhao, Yonghong Yan 0002 |
ISNN (2) | 5 |
| 2009 | Nonnative Speech Recognition Based on Bilingual Model Modification at State Level
Jielin Pan, Shui-duen Chan, Yonghong Yan 0002 |
ISNN (4) | 4 |
| 2008 | Mandarin vowel pronunciation quality evaluation by a novel formant classification method and its combination with traditional algorithmsabstractThis paper discusses the vowel pronunciation quality assessment of our computer assisted Mandarin Chinese learning system. Under the speech recognition framework, phonetic pronunciation assessment is usually based on the phonetic posterior probability score, which may be computed by normalizing the frame-based posterior probability or be calculated on the phone segment directly. By the first method, we can achieve a human-machine scoring correlation coefficient (CC) of 0.832 for vowel; and by the second, the CC can be up to 0.847. In order to improve the performance, we suggest employing the formant feature of vowel. This paper proposes a novel method to utilize formant: we plot formant candidates of each frame on the time-frequency plane to form a bitmap, and then extract its Gabor feature for pattern classification. When we use the classification probability score for pronunciation assessment, we get a CC of 0.842. Finally we combine the three scores with various linear or nonlinear methods; the best CC of 0.913 is gotten by using neural network. Fuping Pan, Qingwei Zhao, Yonghong Yan 0002 |
ICASSP | 3 |
| 2008 | A novel speaker clustering algorithm via supervised affinity propagationabstractThis paper addresses the problem of speaker clustering in telephone conversations. Recently, a new clustering algorithm named affinity propagation (AP) is proposed. It exhibits fast execution speed and finds clusters with low error. However, AP is an unsupervised approach which may make the resulting number of clusters different from the actual one. This deteriorates the speaker purity dramatically. This paper proposes a modified method named supervised affinity propagation (SAP), which automatically reruns the AP procedure to make the final number of clusters converge to the specified number. Experiments are carried out to compare SAP with traditional k-means and agglomerative hierarchical clustering on 4-hour summed channel conversations in the NIST 2004 Speaker Recognition Evaluation. Experiment results show that the SAP method leads to a noticeable speaker purity improvement with slight cluster purity decrease compared with AP. Xiang Zhang 0014, Jie Gao 0020, Ping Lu 0009, Yonghong Yan 0002 |
ICASSP | 4 |
| 2008 | Mandarin-English bilingual Speech Recognition for real world music retrievalabstractThis paper presents our recent work on the development of a grammar-constrained, Mandarin-English bilingual Speech Recognition System (MESRS) for real world music retrieval. In order to balance the performance and the complexity of the bilingual SR system, an unified single set of bilingual acoustic models derived by phone clustering is developed. A novel Two-pass phone clustering method based on Confusion Matrix (TCM) is presented and compared with the log-likelihood measure method. In order to deal with the Mandarin accent in spoken English, different non-native adaptation approaches are investigated. With the effective incorporation of approaches on phone clustering and non-native adaptation, the Phrase Error Rate (PhrER) of MESRS for English utterances was reduced by 24.5% relatively compared to the baseline monolingual English system while the PhrER on Mandarin utterances was comparable to that of the baseline monolingual Mandarin system, and the performance for bilingual code-mixing utterances achieved 22.4% relative PhrER reduction. Jielin Pan, Yonghong Yan 0002 |
ICASSP | 3 |
| 2008 | Recognizing named entities in spoken Chinese dialogues with a character-level maximum entropy tagger
Changchun Bao, Weiqun Xu, Yonghong Yan 0002 |
INTERSPEECH | 3 |
| 2008 | An objective singing evaluation approach by relating acoustic measurements to perceptual ratings
Chuan Cao, Ming Li 0026, Yonghong Yan 0002 |
INTERSPEECH | 4 |
| 2008 | Robust speaker change detection using Kernel-Gaussian model
Jie Gao 0020, Xiang Zhang 0014, Qingwei Zhao, Yonghong Yan 0002 |
INTERSPEECH | 4 |
| 2008 | Forward optimal modeling of acoustic confusions in Mandarin CALL system
Fengpei Ge, Fuping Pan, Changliang Liu, Bin Dong 0003, Yonghong Yan 0002 |
INTERSPEECH | 5 |
| 2008 | Cochannel speech separation using multi-pitch estimation and model based voiced sequential groupingabstractIn this paper, a new cochannel speech separation algorithm us-ing multi-pitch extraction and speaker model based sequential grouping is proposed. After auditory segmentation based on on-set and offset analysis, robust multi-pitch estimation algorithm is performed on each segment and the corresponding voiced portions are segregated. Then speaker pair model based on support vector machine (SVM) is employed to determine the optimal sequential grouping alignments and group the speaker homogeneous segments into pure speaker streams. Systematic evaluation on the speech separation challenge database shows significant improvement over the baseline performance. Index Terms: Auditory scene analysis, cochannel speech, multi-pitch estimation, sequential grouping Ming Li 0026, Chuan Cao, Ping Lu 0009, Qiang Fu 0001, Yonghong Yan 0002 |
INTERSPEECH | 6 |
| 2008 | Towards vocabulary-independent speech indexing for large-scale repositories
Jian Shao 0001, Roger Peng Yu, Qingwei Zhao, Yonghong Yan 0002, Frank Seide |
INTERSPEECH | 4 |
| 2008 | A frequency domain approach for speech enhancement with directionality using compact microphone array
Qiang Fu 0001, Yonghong Yan 0002 |
INTERSPEECH | 3 |
| 2008 | Nonnative speech recognition based on state-candidate bilingual model modification
Ta Li, Jielin Pan, Yonghong Yan 0002 |
INTERSPEECH | 4 |
| 2007 | Audio Segmentation via Tri-Model Bayesian Information CriterionabstractThis paper addresses the problem of audio segmentation in practical media (e.g. TV series, movies and etc.) which usually consists of segments in various lengths with quite a portion of short ones. An unsupervised audio segmentation approach is presented, including a segmentation-stage to detect potential acoustic changes, and a refinement-stage to refine these candidate changes by a tri-model Bayesian information criterion. Experiments show that the proposed approach has good detectability of short segments and the novel tri-model BIC effectively improves the overall segmentation performance. Yunfeng Du, Wei Hu 0002, Yonghong Yan 0002, Tao Wang 0003, Yimin Zhang 0002 |
ICASSP (1) | 3 |
| 2007 | Mandarin Accent Analysis Based on Formant FrequenciesabstractAccent analysis for Mandarin Chinese based on formant frequencies is presented in this paper. Five monophthongs [a, o, e, i, u] of 430 speakers across eight accents were analyzed with univariance analysis of variance (UNIANOVA). The results show that accent has significant influence on the second formant frequency of monophthongs [o, i, u] and has no obvious influence on the formant frequencies of monophthong [a]. In addition, accent has no obvious influence on the first and the third formant frequencies of all five monophthongs. Zhiwei Shuang, Yong Qin 0001, Jianping Zhang 0001, Yonghong Yan 0002 |
ICASSP (4) | 5 |
| 2007 | Robust voice activity detection based on adaptive sub-band energy sequence analysis and harmonic detectionabstractVoice activity detection (VAD) in real-world noise is a very challenging task. In this paper, a two-step methodology is proposed to solve the problem. First, segments with non-stationary components, including speech and dynamic noise, are located using sub-band energy sequence analysis (SESA). Secondly, voice is detected within the selected segments employing the proposed method concerning its harmonic structure. Therefore, speech segments can be accurately detected by this rule-based framework. This algorithm is evaluated in several databases in terms of speech/non-speech discrimination and in terms of word accuracy rate when it is used as the front-end of automatic speech recognition (ASR) system. It provides a more reliable performance over the commonly used standard methods. Index Terms: voice activity detection, harmonic structure, noise robustness, automatic speech recognition Yanmeng Guo, Yonghong Yan 0002 |
INTERSPEECH | 3 |
| 2007 | Spoken language identification using score vector modeling and support vector machineabstractThe support vector machine (SVM) framework based on generalized linear discriminate sequence (GLDS) kernel has been shown effective and widely used in language identifica-tion tasks. In this paper, in order to compensate the distortions due to inter-speaker variability within the same language and solve the practical limitation of computer memory requested by large database training, multiple speaker group based discrim-inative classifiers are employed to map the cepstral features of speech utterances into discriminative language characterization score vectors (DLCSV). Furthermore, backend SVM classifiers are used to model the probability distribution of each target language in the DLCSV space and the output scores of back-end classifiers are calibrated as the final language recognition scores by a pair-wise posterior probability estimation algorithm. The proposed SVM framework is evaluated on 2003 NIST Lan-guage Recognition Evaluation databases, achieving an equal er-ror rate of 4.0 % in 30-second tasks, which outperformed the state-of-art SVM system by more than 30 % relative error re-duction. Index Terms: spoken language identification, support vector machine, score vector modeling Ming Li 0026, Hongbin Suo, Ping Lu 0009, Yonghong Yan 0002 |
INTERSPEECH | 5 |
| 2007 | Mandarin vowel pronunciation quality evaluation by using formant pattern recognition
Fuping Pan, Qingwei Zhao, Yonghong Yan 0002 |
INTERSPEECH | 3 |
| 2007 | A fast fuzzy keyword spotting algorithm based on syllable confusion network
Jian Shao 0001, Qingwei Zhao, Pengyuan Zhang, Zhaojie Liu, Yonghong Yan 0002 |
INTERSPEECH | 5 |
| 2007 | Contributions of temporal fine structure cues to Chinese speech recognition in cochlear implant simulation
Lin Yang 0016, Jianping Zhang 0001, Yonghong Yan 0002 |
INTERSPEECH | 3 |
| 2005 | Fast confidence measure algorithm for continuous speech recognition
Bin Dong 0003, Qingwei Zhao, Yonghong Yan 0002 |
INTERSPEECH | 3 |
| 2004 | Fusion based speech segmentation in DARPA SPINE2 taskabstractWe report a new fusion based segmentation approach using multiple filter bank coefficients. This approach takes advantage of current feature extraction procedure, with little additional computation cost. Another level of fusion was performed by combining several segmentation systems. Evaluation was conducted on the second Speech In Noisy Environments (SPINE2) task. Experiments show our fusion based approaches significantly reduced the WER compared to two classifier-based approaches. Compared to the manual segmentation, our approach only has 0.3% WER increase. Chengyi Zheng, Yonghong Yan 0002 |
ICASSP (1) | 2 |
| 2004 | Robust state clustering using phonetic decision trees
Chaojun Liu, Yonghong Yan 0002 |
Speech Commun. | 2 |
| 2004 | Speaker adaptation using constrained transformationabstractIn speech recognition research, transformation-based adaptation algorithms provide an effective way of adapting acoustic models to improve the recognition accuracy. However, when only limited amounts of adaptation data are available, the transformation is often poorly estimated, which may cause performance degradation. This paper presents the Markov Random Field Linear Regression (MRFLR) algorithm, which constrains the transformation-based adaptation by the correlations among acoustic parameters. The Markov Random Field theory is used to model the correlations. The correlations are estimated from the training corpus and hypothesized as prior knowledge of acoustic models. By explicitly incorporating them into adaptation, robust and fast adaptation can be achieved. The hypothesis is tested by comparing MRFLR with MLLR (Maximum Likelihood Linear Regression), a widely used transformation-based adaptation algorithm. Experimental results show that MRFLR outperforms MLLR when adaptation data are sparse, and converges to the MLLR performance when more adaptation data are available. Xintian Wu, Yonghong Yan 0002 |
IEEE Trans. Speech Audio Process. | 2 |
| 2003 | A dynamic cross-reference pruning strategy for multiple feature fusion at decoder run time
Yonghong Yan 0002, Chengyi Zheng, Jianping Zhang 0001, Jielin Pan, Jiang Han |
INTERSPEECH | 1 |
| 2002 | Run time information fusion in speech recognition
Chengyi Zheng, Yonghong Yan 0002 |
INTERSPEECH | 2 |
| 2001 | A context adaptation approach for building context dependent models in LVCSRabstractAbstract This paper introduces a new context adaptationframework for building context dependent HMM modelsin LVCSR. In this new framework, all states of eachcenter phone are clustered into groups by the decisiontree algorithm. All the tied states of context dependentHMM models were then derived by adapting theparameters of the multiple-mixture context independentmodel via data dependent MAP (maximum a posterioriprobability)method using the training vectorscorresponding to the tied state. An advantage of thisapproach is that it can maintain a high prediction andclassification power given limited training data thereforethe model trained in this framework is more reliable thanin conventional framework. Experimental results onWall Street Journal corpora demonstrate that theproposed approach leads to a significant improvement inrecognition performance. 1. Introduction Decision tree state tying based context modeling hasbecome increasingly popular for modeling speechvariations in large vocabulary speech recognition[1][2].In the conventional framework the stochastic classifierfor each tied state is trained using Baum-Welchalgorithm using the training data corresponding to thespecific tied state[3]. However, the context dependentclassifiers trained using this method are not so reliablefor the training data corresponding to each tied state islimited and model parameters are easily to be affectedby undesired sources of information such as speaker andchannel differences contained in the training data. Toattack this problem, We propose a new contextadaptation method to estimate the parameters of contextdependent models. In this method, a multiple-mixturecontext independent model is trained firstly. In decisiontree clustering, single mixture Gaussian models wereused to establish the state tying. After all the contextdependent states are clustered into groups, the clusteredstates of context dependent model were derived byadapting the parameters of the context independentmodel via data dependent MAP method using trainingdata corresponding to the tied state. We consider themulti-mixture context independent model as coveringthe space of more broad speaker and environmentclasses of speech signal, then adaptation is the contextdependent tuning of those speaker and environmentclasses observed in context’s training speech. Mixtureparameters for those speaker and environment classesnot observed in the training speech of the specific tiedstate are merely copied from the context independentmodel. This means that the model has higher predictionand classification power for the test data from speakerand environment classes unseen or rarely seen in thecontext’s training data. Experimental results on WallStreet Journal corpora demonstrate that the proposedapproaches lead to a significant improvement inrecognition performance. Xiaoxing Liu, Baosheng Yuan, Yonghong Yan 0002 |
INTERSPEECH | 3 |
| 2000 | Linear regression under maximum a posteriori criterion with Markov random field priorabstractSpeaker adaptation using linear transformations under the maximum a posteriori (MAP) criterion has been studied in this paper. The purpose is to improve the matrix estimation in the widely used maximum likelihood linear regression (MLLR) adaptation, which might generate poorly structured transform matrices when adaptation data are sparse. Unlike traditional MAP based adaptations, many known prior distributions of HMM parameters, such as normal-Washart priors, do not have a close form solution in the transform estimation. In Markov random field linear regression (MRFLR), the prior distribution of HMM parameters is modeled by Markov random field, which leads to a close form solution of estimating the linear transforms. Experimental results show that MRFLR outperforms MLLR when adaptation data are sparse, and converges to the MLLR performances when more adaptation data are available. Xintian Wu, Yonghong Yan 0002 |
ICASSP | 2 |
| 2000 | Keyword spotting in auto-attendant systemabstractIn this paper, an auto-attendant system using finite state grammar (FSG) based on a continuous speech recognition (CSR) model is introduced. However, by using two virtual garbage models, one is to match the leading extraneous speech before the key name and the other to match the tailing extraneous speech following the key name, we managed to reach a more flexible and robust auto-attendant system. The experiment result show that, in our auto attendant system (about 240 names), to the name only test set and the sentence test set 1 composed of sentences that FSG can recognize, the recognition rate of the keyword spotting system is almost the same as that of FSG. To the sentence test set 2 composed of sentences that undefined in the FSG the keyword spotting system outperforms the FSG system remarkably. Not affecting the recognition accuracy of name only test set and the sentence test set 1, task dependent keyword models cut off additional 20% of error rate comparing with task independent keyword models in the sentence test set 2. Yonghong Yan 0002, Baosheng Yuan, Qingwei Zhao |
INTERSPEECH | 2 |
| 2000 | Vocabulary-based acoustic model trim down and task adaptation
Yonghong Yan 0002, Baosheng Yuan, Ying Jia, Xiaoxing Liu |
INTERSPEECH | 2 |
| 2000 | Office message center - a spoken dialogue system
Jiang Han, Yonghong Yan 0002, Danjun Liu |
INTERSPEECH | 2 |
| 2000 | Dynamic threshold setting via Bayesian information criterion (BIC) in HMM training
Ying Jia, Yonghong Yan 0002, Baosheng Yuan |
INTERSPEECH | 2 |
| 2000 | Speaker change detection using minimum message length criterion
Chaojun Liu, Yonghong Yan 0002 |
INTERSPEECH | 2 |
| 2000 | An orthogonal GMM based speaker verification system
Xiaoxing Liu, Baosheng Yuan, Yonghong Yan 0002 |
INTERSPEECH | 3 |
| 2000 | Effective vector quantization for a highly compact acoustic model for LVCSR
Jielin Pan, Baosheng Yuan, Yonghong Yan 0002 |
INTERSPEECH | 3 |
| 2000 | Improvements in search algorithm for large vocabulary continuous speech recognition
Qingwei Zhao, Baosheng Yuan, Yonghong Yan 0002 |
INTERSPEECH | 4 |
| 2000 | Efficiently using speaker adaptation data
Chengyi Zheng, Yonghong Yan 0002 |
INTERSPEECH | 2 |
| 1999 | High accuracy acoustic modeling based on multi-stage decision tree
Chaojun Liu, Xintian Wu, Yonghong Yan 0002 |
EUROSPEECH | 4 |
| 1999 | High accuracy acoustic modeling using two-level decision-tree based state-tyingabstractThis paper addresses the problem of language modeling for the transcription of broadcast news data. Different approaches for language model training were explored and tested in the context of a complete transcription system. Language model efficiency was investigated for the following aspects: mixing of different training material (sources and epoch); approach for mixing (interpolation vs count merging); and using class-based language models. The experimental results indicate that judicious selection of the training source and epoch is important, and that given sufficient broadcast new transcriptions, newspaper and newswire texts are not necessary. Results are given in terms of perplexity and word error rates. The combined improvements in text selection, interpolation, 4-gram and class-based LMs led to a 20% reduction in the perplexity of the LM of the final pass (3-gram class interpolated with a word 4-gram) compared with the 3-gram LM used in the the LIMSI Nov’97 BN system. Chaojun Liu, Xintian Wu, Yonghong Yan 0002 |
EUROSPEECH | 3 |
| 1999 | Development of the 1998 OGI-FONIX broadcast news transcription system
Xintian Wu, Yonghong Yan 0002 |
EUROSPEECH | 2 |
| 1999 | Understanding speech recognition using correlation-generated neural network targetsabstractTraining neural networks with variable targets for speech recognition systems has been shown to be effective in improving word accuracy. In this correspondence, a new and simple method for estimating variable targets for a given training pattern is presented. It uses estimated correlations between different output nodes of a neural network to create a set of variable targets for each training pattern. Experimental results show that the word error is reduced by more than 20% when these new correlation-based targets are compared to more conventional zero/one targets with a squared-error cost function. Performance with these new targets approaches that of high-performance hidden Markov model (HMM) recognizers but requires far fewer parameters. Yonghong Yan 0002 |
IEEE Trans. Speech Audio Process. | 1 |
| 1998 | Accessible technology for interactive systems: a new approach to spoken language researchabstractIn this paper, we argue for a paradigm shift in spoken language technology, from transcription tasks to interactive systems. The current paradigm evaluates speech recognition technology in terms of word recognition accuracy on large vocabulary transcription tasks, such as telephone conversations or media broadcasts. Systems are evaluated in international competitions, with strict rules for participation and well-defined evaluation metrics. Participation in these competitions is limited to a few elite laboratories that have the resources to develop and field systems. We propose a new, more productive and more accessible paradigm for spoken language research, in which research advances are evaluated in the context of interactive systems that allow people to perform useful tasks, such as accessing information from the World Wide Web, while driving a car. These systems are made available for daily use by ordinary citizens through telephone networks or placement in easily accessible kiosks in public institutions. It has previously been argued that this new paradigm, which focuses on the goal of universal access to information for all people, better serves the needs of the research community, as well as the welfare of our citizens. We discuss the challenges and rewards of an interactive system approach to spoken language research, and discuss our initial attempts to stimulate a paradigm shift and engage a large community of researchers through free distribution of the CSLU toolkit. Ronald A. Cole, Stephen Sutton, Yonghong Yan 0002, Pieter J. E. Vermeulen, Mark A. Fanty |
ICASSP | 3 |
| 1998 | Universal speech tools: the CSLU toolkitabstractA set of freely available, universal speech tools is needed to accelerate progress in the speech technology. The CSLU Toolkit represents an effort to make the core technology and fundamental infrastructure accessible, affordable and easy to use. The CSLU Toolkit has been under development for five years. This paper describes recent improvements, additions and uses of the CSLU Toolkit. 1. INTRODUCTION Since 1993, the Center for Spoken Language Understanding (CSLU) has focused on incorporating state-of-the-art spokenlanguage technology into a portable, comprehensive and easyto -use software environment. The result of these efforts is the CSLU Toolkit. The toolkit integrates learning materials, authoring tools and core technologies such as speech recognition, text-to-speech synthesis, facial animation and speech reading. The toolkit is designed to support basic research, development and education activities related to spoken language systems and human-computer interfaces. What are our ... Stephen Sutton, Ronald A. Cole, Jacques de Villiers, Johan Schalkwyk, Pieter J. E. Vermeulen, Michael W. Macon, Yonghong Yan 0002, Edward C. Kaiser, Brian Rundle, Khaldoun Shobaki, John-Paul Hosom, Alexander Kain, Johan Wouters, Dominic W. Massaro, Michael M. Cohen |
ICSLP | 7 |
| 1997 | Speech recognition using neural networks with forward-backward probability generated targetsabstractNeural network training targets for speech recognition are estimated using a novel method. Rather than use zero and one, continuous targets are generated using forward-backward probabilities. Each training pattern has more than one class active. Experiments showed that the new method effectively decreased the error rate by 15% in a continuous digits recognition task. Yonghong Yan 0002, Mark A. Fanty, Ronald A. Cole |
ICASSP | 1 |
| 1997 | Matching training and testing criteria in hybrid speech recognition systems
Xin Tu, Yonghong Yan 0002, Ronald A. Cole |
EUROSPEECH | 2 |
| 1997 | Toward new language adaptation for language identification
Etienne Barnard, Yonghong Yan 0002 |
Speech Commun. | 2 |
| 1996 | Experiments for an approach to language identification with conversational telephone speechabstractThis paper presents work on language identification research using conversational speech (the LDC Conversational Telephone Speech Database). The baseline system used in this study is based on language-dependent phone recognition and phonotactic constraints. The system was trained using monologue data and obtained an error rate of around 9% on a commonly used nine-language monologue test set. While the system was used to process conversational speech from the same nine-language task, dramatic performance degradation (with an error rate of 40%) was observed. Based on our analysis of conversational speech, two methods: (1) pre-processing and, (2) post-processing, were proposed. Without the presence of training data from conversational speech database, the final system (the baseline system enhanced by the two proposed methods) obtained an error rate of 24%, a substantial improvement (with 41% error reduction) compared with the baseline system. Yonghong Yan 0002, Etienne Barnard |
ICASSP | 1 |
| 1996 | The contribution of consonants versus vowels to word recognition in fluent speechabstractThree perceptual experiments were conducted to test the relative importance of vowels vs. consonants to recognition of fluent speech. Sentences were selected from the TIMIT corpus to obtain approximately equal numbers of vowels and consonants within each sentence and equal durations across the set of sentences. In experiments 1 and 2, subjects listened to (a) unaltered TIMIT sentences; (b) sentences in which all of the vowels were replaced by noise; or (c) sentences in which all of the consonants were replaced by noise. The subjects listened to each sentence five times, and attempted to transcribe what they heard. The results of these experiments show that recognition of words depends more upon vowels than consonants-about twice as many words are recognized when vowels are retained in the speech. The effect was observed when occurrences of [1], [r], [w], [y] [m], [n], were included in the sentences (experiment 1) or replaced by noise (experiment 2). Experiment 3 tested the hypothesis that vowel boundaries contain more information about the neighboring consonants than vice versa. Ronald A. Cole, Yonghong Yan 0002, Brian Kan-Wing Mak, Mark A. Fanty, Troy Bailey |
ICASSP | 2 |
| 1996 | The influence of bigram constraints on word recognition by humans: implications for computer speech recognition
Ronald A. Cole, Yonghong Yan 0002, Troy Bailey |
ICSLP | 2 |
| 1996 | Development of an approach to automatic language identification based on phone recognition
Yonghong Yan 0002, Etienne Barnard, Ronald A. Cole |
Comput. Speech Lang. | 1 |
| 1995 | An approach to automatic language identification based on language-dependent phone recognitionabstractAn approach to language identification (LID) based on language-dependent phone recognition is presented. A variety of features and their combinations extracted by language-dependent recognizers were evaluated based on the same database. Two novel information sources for LID were introduced: (1) forward and backward bigram based language models, and (2) context-dependent duration models. An LID system using hidden Markov models and neural network was developed. The system was trained and evaluated using the OGLTS database. For a six-language task, the system performance (correct rate) for 45-second long utterances and 10-second long utterances reached 91-96% and 81-08% respectively. The experiments demonstrated the importance of detailed modeling and the method by which these information sources are combined. Yonghong Yan 0002, Etienne Barnard |
ICASSP | 1 |
| 1995 | An approach to language identification with enhanced language model
Yonghong Yan 0002, Etienne Barnard |
EUROSPEECH | 1 |