Yao Qian

dblp:67/2224 · DBLP profile ↗
← Back
98ranked-venue papers
28as first author
33since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 78 · 23 first-author · 21 since 2021Artificial intelligence and machine learning · 55 · 17 first-author · 16 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 RiskDiffuser: Continuous risk conditioning for diffusion planning in autonomous driving
Yao Qian, Yanfei Wu
Expert Syst. Appl.1
2025 SLM-S2ST: A multimodal language model for direct speech-to-speech translation
abstract
Speech-aware language models (LMs) have demonstrated capabilities in understanding spoken language while generating text-based responses. However, enabling them to produce speech output efficiently and effectively remains a challenge. In this paper, we present SLM-S2ST, a multimodal LM for direct speech-to-speech translation (S2ST), built on the opensource Phi4-MM model. SLM-S2ST extends its predecessor by generating translated speech using an audio transformer head that predicts audio tokens with a delay relative to text tokens, followed by a streaming vocoder for waveform synthesis. Our experimental results on the CVSS-C dataset demonstrate SLMS2ST’s superior performance, significantly surpassing existing baseline models trained on the same dataset. Furthermore, when we scale up the training data and the model size, SLM-S2ST reaches on-par performance with the current SOTA model.
Yuxuan Hu 0003, Ruchao Fan, Xiaofei Wang 0009, Yao Qian, Jinyu Li 0001
ASRU6
2025 Towards Efficient Speech-Text Jointly Decoding within One Speech Language Model
abstract
Speech language models (Speech LMs) enable end-to-end speech-text modeling within a single model, offering a promising direction for spoken dialogue systems. The choice of speech-text jointly decoding paradigm plays a critical role in performance, efficiency, and alignment quality. In this work, we systematically compare representative joint speech-text decoding strategies—including the interleaved, and parallel generation paradigms—under a controlled experimental setup using the same base language model, speech tokenizer and training data. Our results show that the interleaved approach achieves the best alignment. However it suffers from slow inference due to long token sequence length. To address this, we propose a novel early-stop interleaved (ESI) pattern that not only significantly accelerates decoding but also yields slightly better performance. Additionally, we curate high-quality question answering (QA) datasets to further improve speech QA performance.
Yuxuan Hu 0003, Ruchao Fan, Xiaofei Wang 0009, Ken'ichi Kumatani, Jianwei Yu 0001, Yao Qian, Jinyu Li 0001
ASRU10
2025 CoVoMix2: Advancing Zero-Shot Dialogue Generation with Fully Non-Autoregressive Flow Matching
abstract
Generating natural-sounding, multi-speaker dialogue is crucial for applications such as podcast creation, virtual agents, and multimedia content generation. However, existing systems struggle to maintain speaker consistency, model overlapping speech, and synthesize coherent conversations efficiently. In this paper, we introduce CoVoMix2, a fully non-autoregressive framework for zero-shot multi-talker dialogue generation. CoVoMix2 directly predicts mel-spectrograms from multi-stream transcriptions using a flow-matching-based generative model, eliminating the reliance on intermediate token representations. To better capture realistic conversational dynamics, we propose transcription-level speaker disentanglement, sentence-level alignment, and prompt-level random masking strategies. Our approach achieves state-of-the-art performance, outperforming strong baselines like MoonCast and Sesame in speech quality, speaker consistency, and inference speed. Notably, CoVoMix2 operates without requiring transcriptions for the prompt and supports controllable dialogue generation, including overlapping speech and precise timing control, demonstrating strong generalizability to real-world speech generation scenarios. Audio samples are available at https://www.microsoft.com/en-us/research/project/covomix/covomix2.
Leying Zhang, Yao Qian, Xiaofei Wang 0009, Manthan Thakker, Dongmei Wang, Jianwei Yu 0001, Yuxuan Hu 0003, Jinyu Li 0001, Yanmin Qian, Sheng Zhao 0002
NeurIPS2
2025 Transformer architecture with illumination aware mechanisms for low-light image enhancement via Retinex decomposition
Hanlin Xu, Yao Qian, Durga Prasad Bavirisetti
Eng. Appl. Artif. Intell.4
2025 GDNet: a low-light image enhancement network based on Ghost-Block and unique image decomposition
Yao Qian, Haojie Tang, Gaoqiang Wang, Durga Prasad Bavirisetti
J. Supercomput.3
2025 Infrared and visible image fusion method based on visual saliency objects and fuzzy region attributes
Jiebang Wang, Yao Qian
Vis. Comput.3
2024 Adapting Large Language Model with Speech for Fully Formatted End-to-End Speech Recognition
abstract
Most end-to-end (E2E) speech recognition models are composed of encoder and decoder blocks that perform acoustic and language modeling functions. Pretrained large language models (LLMs) have the potential to improve the performance of E2E ASR. However, integrating a pretrained language model into an E2E speech recognition model has shown limited benefits due to the mismatches between text-based LLMs and those used in E2E ASR. In this paper, we explore an alternative approach by adapting a pretrained LLMs to speech. Our experiments on fully-formatted E2E ASR transcription tasks across various domains demonstrate that our approach can effectively leverage the strengths of pretrained LLMs to produce more readable ASR transcriptions. Our model, which is based on the pretrained large language models with either an encoder-decoder or decoder-only structure, surpasses strong ASR models such as Whisper1, in terms of recognition error rate, considering formats like punctuation and capitalization as well.
Shaoshi Ling, Yuxuan Hu 0003, Shuangbei Qian, Guoli Ye, Yao Qian, Yifan Gong 0001, Ed Lin, Michael Zeng 0001
ICASSP5
2024 TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation
abstract
There is a rising interest and trend in research towards directly translating speech from one language to another, known as end-to-end speech-to-speech translation. However, most end-to-end models struggle to outperform cascade models, i.e., a pipeline framework by concatenating speech recognition, machine translation and text-to-speech models. The primary challenges stem from the inherent complexities involved in direct translation tasks and the scarcity of data. In this study, we introduce a novel model framework TransVIP that leverages diverse datasets in a cascade fashion yet facilitates end-to-end inference through joint probability. Furthermore, we propose two separated encoders to preserve the speaker’s voice characteristics and isochrony from the source speech during the translation process, making it highly suitable for scenarios such as video dubbing. Our experiments on the French-English language pair demonstrate that our model outperforms the current state-of-the-art speech-to-speech translation model.
Chenyang Le, Yao Qian, Dongmei Wang, Shujie Liu 0001, Xiaofei Wang 0009, Midia Yousefi, Yanmin Qian, Jinyu Li 0001, Sheng Zhao 0002, Michael Zeng 0001
NeurIPS2
2024 CoVoMix: Advancing Zero-Shot Speech Generation for Human-like Multi-talker Conversations
abstract
Recent advancements in zero-shot text-to-speech (TTS) modeling have led to significant strides in generating high-fidelity and diverse speech. However, dialogue generation, along with achieving human-like naturalness in speech, continues to be a challenge. In this paper, we introduce CoVoMix: Conversational Voice Mixture Generation, a novel model for zero-shot, human-like, multi-speaker, multi-round dialogue speech generation. CoVoMix first converts dialogue text into multiple streams of discrete tokens, with each token stream representing semantic information for individual talkers. These token streams are then fed into a flow-matching based acoustic model to generate mixed mel-spectrograms. Finally, the speech waveforms are produced using a HiFi-GAN model. Furthermore, we devise a comprehensive set of metrics for measuring the effectiveness of dialogue modeling and generation. Our experimental results show that CoVoMix can generate dialogues that are not only human-like in their naturalness and coherence but also involve multiple talkers engaging in multiple rounds of conversation. This is exemplified by instances generated in a single channel where one speaker's utterance is seamlessly mixed with another's interjections or laughter, indicating the latter's role as an attentive listener. Audio samples are enclosed in the supplementary.
Leying Zhang, Yao Qian, Shujie Liu 0001, Dongmei Wang, Xiaofei Wang 0009, Midia Yousefi, Yanmin Qian, Jinyu Li 0001, Lei He 0005, Sheng Zhao 0002, Michael Zeng 0001
NeurIPS2
2024 Investigating Neural Audio Codecs For Speech Language Model-Based Speech Generation
abstract
Neural audio codec tokens serve as the fundamental building blocks for speech language model (SLM)-based speech generation. However, there is no systematic understanding on how the codec system affects the speech generation performance of the SLM. In this work, we examine codec tokens within SLM framework for speech generation to provide insights for effective codec design. We retrain existing high-performing neural codec models on the same data set and loss functions to compare their performance in a uniform setting. We integrate codec tokens into two SLM systems: masked-based parallel speech generation system and an auto-regressive (AR) plus non-auto-regressive (NAR) model-based system. Our findings indicate that better speech reconstruction in codec systems does not guarantee improved speech generation in SLM. A high-quality codec decoder is crucial for natural speech production in SLM, while speech intelligibility depends more on quantization mechanism.
Jiaqi Li 0030, Dongmei Wang, Xiaofei Wang 0009, Yao Qian, Shujie Liu 0001, Midia Yousefi, Canrun Li, Chung-Hsien Tsai, Jun-Kun Chen, Sheng Zhao 0002, Jinyu Li 0001, Zhizheng Wu 0001, Michael Zeng 0001
SLT4
2024 DDTSE: Discriminative Diffusion Model for Target Speech Extraction
abstract
Diffusion models have gained attention in speech enhancement tasks, providing an alternative to conventional discriminative methods. However, research on target speech extraction under multispeaker noisy conditions remains relatively unexplored. Moreover, the superior quality of diffusion methods typically comes at the cost of slower inference speed. In this paper, we introduce the Discriminative Diffusion model for Target Speech Extraction (DDTSE). We apply the same forward process as diffusion models and utilize the reconstruction loss similar to discriminative methods. Furthermore, we devise a two-stage training strategy to emulate the inference process during model training. DDTSE not only works as a standalone system, but also can further improve the performance of discriminative models without additional retraining. Experimental results demonstrate that DDTSE not only achieves higher perceptual quality but also accelerates the inference process by 3 times compared to the conventional diffusion model.
Leying Zhang, Yao Qian, Linfeng Yu, Heming Wang, Hemin Yang, Shujie Liu 0001, Yanmin Qian
SLT2
2024 GAN-GA: infrared and visible image fusion generative adversarial network based on global awareness
Haojie Tang, Yao Qian
Appl. Intell.5
2024 RDGMEF: a multi-exposure image fusion framework based on Retinex decompostion and guided filter
Haojie Tang, Yao Qian, Jianchao Tang
Neural Comput. Appl.4
2024 CFNet: An infrared and visible image compression fusion network
Mengliang Xing, Haojie Tang, Yao Qian
Pattern Recognit.4
2023 i-Code: An Integrative and Composable Multimodal Learning Framework
abstract
Human intelligence is multimodal; we integrate visual, linguistic, and acoustic signals to maintain a holistic worldview. Most current pretraining methods, however, are limited to one or two modalities. We present i-Code, a self-supervised pretraining framework where users may flexibly combine the modalities of vision, speech, and language into unified and general-purpose vector representations. In this framework, data from each modality are first given to pretrained single-modality encoders. The encoder outputs are then integrated with a multimodal fusion network, which uses novel merge- and co-attention mechanisms to effectively combine information from the different modalities. The entire system is pretrained end-to-end with new objectives including masked modality unit modeling and cross-modality contrastive learning. Unlike previous research using only video for pretraining, the i-Code framework can dynamically process single, dual, and triple-modality data during training and inference, flexibly projecting different combinations of modalities into a single representation space. Experimental results demonstrate how i-Code can outperform state-of-the-art techniques on five multimodal understanding tasks and single-modality benchmarks, improving by as much as 11% and demonstrating the power of integrative multimodal pretraining.
Ziyi Yang 0011, Yuwei Fang, Chenguang Zhu 0001, Reid Pryzant, Dongdong Chen 0001, Yu Shi 0001, Yichong Xu, Yao Qian, Mei Gao, Liyang Lu, Yujia Xie, Robert Gmyr, Noel Codella, Naoyuki Kanda, Bin Xiao 0004, Lu Yuan 0001, Takuya Yoshioka, Michael Zeng 0001, Xuedong Huang 0001
AAAI8
2023 Target Sound Extraction with Variable Cross-Modality Clues
abstract
Automatic target sound extraction (TSE) is a machine learning approach to mimic the human auditory perception capability of attending to a sound source of interest from a mixture of sources. It often uses a model conditioned on a fixed form of target sound clues, such as a sound class label, which limits the ways in which users can interact with the model to specify the target sounds. To leverage variable number of clues cross modalities available in the inference phase, including a video, a sound event class, and a text caption, we propose a unified transformer-based TSE model architecture, where a multi-clue attention module integrates all the clues across the modalities. Since there is no off-the-shelf benchmark to evaluate our proposed approach, we build a dataset1based on public corpora, Audioset and AudioCaps. Experimental results for seen and unseen target-sound evaluation sets show that our proposed TSE model can effectively deal with a varying number of clues which improves the TSE performance and robustness against partially compromised clues.
Chenda Li, Yao Qian, Zhuo Chen 0006, Dongmei Wang, Takuya Yoshioka, Shujie Liu 0001, Yanmin Qian, Michael Zeng 0001
ICASSP2
2023 DATA2VEC-SG: Improving Self-Supervised Learning Representations for Speech Generation Tasks
abstract
Self-supervised learning has been successfully applied to various speech recognition and understanding tasks. However, for generative tasks such as speech enhancement and speech separation, most self-supervised speech representations did not show substantial improvements. To deal with this problem, in this paper, we propose data2vec-SG (Speech Generation), which is a teacher-student learning framework that addresses speech generation tasks. Our data2vec-SG introduces a reconstruction module into data2vec [1] and enforces the representations to contain not only the semantic information but also the acoustic knowledge to generate clean speech waveforms. Experimental results demonstrate that the proposed framework boosts the performance of various speech generation tasks including speech enhancement, speech separation, and packet loss concealment. Meanwhile, the learned representation is also capable of helping other downstream tasks, which is demonstrated by the good performance in the speech recognition task in both clean and noisy conditions.
Heming Wang, Yao Qian, Hemin Yang, Naoyuki Kanda, Takuya Yoshioka, Xiaofei Wang 0009, Shujie Liu 0001, Zhuo Chen 0006, DeLiang Wang, Michael Zeng 0001
ICASSP2
2023 Code-Switching Text Generation and Injection in Mandarin-English ASR
abstract
Code-switching speech refers to a means of expression by mixing two or more languages within a single utterance. Automatic Speech Recognition (ASR) with End-to-End (E2E) modeling for such speech can be a challenging task due to the lack of data. In this study, we investigate text generation and injection for improving the performance of an industry commonly-used streaming model, Transformer-Transducer (T-T), in Mandarin-English code-switching speech recognition. We first propose a strategy to generate codeswitching text data and then investigate injecting generated text into T-T model explicitly by Text-To-Speech (TTS) conversion or implicitly by tying speech and text latent spaces. Experimental results on the T-T model trained with a dataset containing 1,800 hours of real Mandarin-English code-switched speech show that our approaches to inject generated code-switching text significantly boost the performance of T-T models, i.e., 16% relative Token-based Error Rate (TER) reduction averaged on three evaluation sets, and the approach of tying speech and text latent spaces is superior to that of TTS conversion on the evaluation set which contains more homogeneous data with the training set.
Yuxuan Hu 0003, Yao Qian, Ma Jin, Linquan Liu, Shujie Liu 0001, Yu Shi 0001, Yanmin Qian, Edward Lin, Michael Zeng 0001
ICASSP3
2023 Adapting Multi-Lingual ASR Models for Handling Multiple Talkers
Chenda Li, Yao Qian, Zhuo Chen 0006, Naoyuki Kanda, Dongmei Wang, Takuya Yoshioka, Yanmin Qian, Michael Zeng 0001
INTERSPEECH2
2023 ComSL: A Composite Speech-Language Model for End-to-End Speech-to-Text Translation
abstract
Joint speech-language training is challenging due to the large demand for training data and GPU consumption, as well as the modality gap between speech and language. We present ComSL, a speech-language model built atop a composite architecture of public pre-trained speech-only and language-only models and optimized data-efficiently for spoken language tasks. Particularly, we propose to incorporate cross-modality learning into transfer learning and conduct them simultaneously for downstream tasks in a multi-task learning manner. Our approach has demonstrated effectiveness in end-to-end speech-to-text translation tasks, achieving a new state-of-the-art average BLEU score of 31.5 on the multilingual speech to English text translation task for 21 languages, as measured on the public CoVoST2 evaluation set.
Chenyang Le, Yao Qian, Shujie Liu 0001, Yanmin Qian, Michael Zeng 0001, Xuedong Huang 0001
NeurIPS2
2022 SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing
abstract
Junyi Ao, Rui Wang, Long Zhou, Chengyi Wang, Shuo Ren, Yu Wu, Shujie Liu, Tom Ko, Qing Li, Yu Zhang, Zhihua Wei, Yao Qian, Jinyu Li, Furu Wei. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Junyi Ao, Rui Wang 0073, Chengyi Wang 0002, Shuo Ren 0002, Yu Wu 0012, Shujie Liu 0001, Tom Ko, Qing Li 0001, Yu Zhang 0006, Zhihua Wei 0001, Yao Qian, Jinyu Li 0001, Furu Wei
ACL (1)12
2022 Large-Scale Self-Supervised Speech Representation Learning for Automatic Speaker Verification
abstract
The speech representations learned from large-scale unlabeled data have shown better generalizability than those from supervised learning and thus attract a lot of interest to be applied for various downstream tasks. In this paper, we explore the limits of speech representations learned by different self-supervised objectives and datasets for automatic speaker verification (ASV), especially with a well-recognized SOTA ASV model, ECAPA-TDNN [1], as a downstream model. The representations from all hidden layers of the pre-trained model are firstly averaged with learnable weights and then fed into the ECAPA-TDNN as input features. The experimental results on Voxceleb dataset show that the weighted average representation is significantly superior to FBank, a conventional handcrafted feature for ASV. Our best single system achieves 0.537%, 0.569%, and 1.180% equal error rate (EER) on the three official trials of VoxCeleb1, separately. Accordingly, the ensemble system with three pre-trained models can further improve the EER to 0.479%, 0.536% and 1.023%. Among the three evaluation trials, our best system outperforms the winner system [2] of the VoxCeleb Speaker Recognition Challenge 2021 (VoxSRC2021) on the VoxCeleb1-E trial.
Zhengyang Chen, Sanyuan Chen, Yu Wu 0012, Yao Qian, Chengyi Wang 0002, Shujie Liu 0001, Yanmin Qian, Michael Zeng 0001
ICASSP4
2022 Unispeech-Sat: Universal Speech Representation Learning With Speaker Aware Pre-Training
abstract
Self-supervised learning (SSL) is a long-standing goal for speech processing, since it utilizes large-scale unlabeled data and avoids extensive human labeling. Recent years have witnessed great successes in applying self-supervised learning in speech recognition, while limited exploration was attempted in applying SSL for modeling speaker characteristics. In this paper, we aim to improve the existing SSL framework for speaker representation learning. Two methods are introduced for enhancing the unsupervised speaker information extraction. First, we apply multi-task learning to the current SSL framework, where we integrate utterance-wise contrastive loss with the SSL objective function. Second, for better speaker discrimination, we propose an utterance mixing strategy for data augmentation, where additional overlapped utterances are created unsupervisely and incorporated during training. We integrate the proposed methods into the HuBERT framework. Experiment results on the SUPERB benchmark show that the proposed system achieves state-of-the-art performance in universal representation learning, especially for speaker identification oriented tasks. An ablation study is performed verifying the efficacy of each proposed method. Finally, we scale up the training dataset to 94 thousand hours of public audio data and achieve further performance improvement in all SUPERB tasks.
Sanyuan Chen, Yu Wu 0012, Chengyi Wang 0002, Zhengyang Chen, Zhuo Chen 0006, Shujie Liu 0001, Jian Wu 0027, Yao Qian, Furu Wei, Jinyu Li 0001, Xiangzhan Yu
ICASSP8
2022 Wav2vec-Switch: Contrastive Learning from Original-Noisy Speech Pairs for Robust Speech Recognition
abstract
The goal of self-supervised learning (SSL) for automatic speech recognition (ASR) is to learn good speech representations from a large amount of unlabeled speech for the downstream ASR task. However, most SSL frameworks do not consider noise robustness which is crucial for real-world applications. In this paper we propose wav2vec-Switch, a method to encode noise robustness into contextualized representations of speech via contrastive learning. Specifically, we feed original-noisy speech pairs simultaneously into the wav2vec 2.0 network. In addition to the existing contrastive learning task, we switch the quantized representations of the original and noisy speech as additional prediction targets of each other. By doing this, it enforces the network to have consistent predictions for the original and noisy speech, thus allows to learn contextualized representation with noise robustness. Our experiments on synthe-sized and real noisy data show the effectiveness of our method: it achieves 2.9–4.9% relative word error rate (WER) reduction on the synthesized noisy LibriSpeech data without deterioration on the original data, and 5.7% on CHiME-4 real 1-channel noisy data compared to a data augmentation baseline even with a strong language model for decoding. Our results on CHiME-4 can match or even surpass those with well-designed speech enhancement components.
Jinyu Li 0001, Heming Wang, Yao Qian, Chengyi Wang 0002, Yu Wu 0012
ICASSP4
2022 Improving Noise Robustness of Contrastive Speech Representation Learning with Speech Reconstruction
abstract
Noise robustness is essential for deploying automatic speech recognition (ASR) systems in real-world environments. One way to reduce the effect of noise interference is to employ a preprocessing module that conducts speech enhancement, and then feed the enhanced speech to an ASR backend. In this work, instead of suppressing background noise with a conventional cascaded pipeline, we employ a noise-robust representation learned by a refined self-supervised framework for noisy speech recognition. We propose to combine a reconstruction module with contrastive learning and perform multi-task continual pre-training on noisy data. The reconstruction module is used for auxiliary learning to improve the noise robustness of the learned representation and thus is not required during inference. Experiments demonstrate the effectiveness of our proposed method. Our model substantially reduces the word error rate (WER) for the synthesized noisy LibriSpeech test sets, and yields around 4.1/7.5% WER reduction on noisy clean/other test sets compared to data augmentation. For the real-world noisy speech from the CHiME-4 challenge (1-channel track), we have obtained the state of the art ASR performance without any denoising front-end. Moreover, we achieve comparable performance to the best supervised approach reported with only 16% of labeled data.
Heming Wang, Yao Qian, Xiaofei Wang 0009, Chengyi Wang 0002, Shujie Liu 0001, Takuya Yoshioka, Jinyu Li 0001, DeLiang Wang
ICASSP2
2022 Optimizing Alignment of Speech and Language Latent Spaces for End-To-End Speech Recognition and Understanding
abstract
The advances in attention-based encoder-decoder (AED) networks have brought great progress to end-to-end (E2E) automatic speech recognition (ASR). One way to further improve the performance of AED-based E2E ASR is to introduce an extra text encoder for leveraging extensive text data and thus capture more context-aware linguistic information. However, this approach brings a mismatch problem between the speech encoder and the text encoder due to the different units used for modeling. In this paper, we propose an embedding aligner and modality switch training to better align the speech and text latent spaces. The embedding aligner is a shared linear projection between text encoder and speech encoder trained by masked language modeling (MLM) loss and connectionist temporal classification (CTC), respectively. The modality switch training randomly swaps speech and text embeddings based on the forced alignment result to learn a joint representation space. Experimental results show that our proposed approach achieves a relative 14% to 19% word error rate (WER) reduction on Librispeech ASR task. We further verify its effectiveness on spoken language understanding (SLU), i.e., an absolute 2.5% to 2.8% F1 score improvement on SNIPS slot filling task.
Wei Wang 0010, Shuo Ren 0002, Yao Qian, Shujie Liu 0001, Yu Shi 0001, Yanmin Qian, Michael Zeng 0001
ICASSP3
2022 Improving Self-Supervised Learning for Speech Recognition with Intermediate Layer Supervision
abstract
Recently, pioneer work finds that self-supervised pre-training methods can improve multiple downstream speech tasks, because the model utilizes bottom layers to learn speaker-related information and top layers to encode content-related information. Since the network capacity is limited, we believe the speech recognition performance could be further improved if the model is dedicated to audio content information learning. To this end, we propose Intermediate Layer Supervision for Self-Supervised Learning (ILS-SSL), which forces the model to concentrate on content information as much as possible by adding an additional SSL loss on the intermediate layers. Experiments on LibriSpeech test-other set show that our method outperforms HuBERT significantly, which achieves a 23.5%/11.6% relative word error rate reduction in the w/o language model setting for Base/Large models. Detailed analysis shows the bottom layers of our model have a better correlation with phonetic units, which is consistent with our intuition and explains the success of our method for ASR. We will release our code and model at https://github.com/microsoft/UniSpeech.
Chengyi Wang 0002, Yu Wu 0012, Sanyuan Chen, Shujie Liu 0001, Jinyu Li 0001, Yao Qian, Zhenglu Yang
ICASSP6
2022 Pre-Training Transformer Decoder for End-to-End ASR Model with Unpaired Speech Data
abstract
This paper studies a novel pre-training technique with unpaired speech data, Speech2C, for encoder-decoder based automatic speech recognition (ASR).Within a multi-task learning framework, we introduce two pre-training tasks for the encoderdecoder network using acoustic units, i.e., pseudo codes, derived from an offline clustering model.One is to predict the pseudo codes via masked language modeling in encoder output, like HuBERT model, while the other lets the decoder learn to reconstruct pseudo codes autoregressively instead of generating textual scripts.In this way, the decoder learns to reconstruct original speech information with codes before learning to generate correct text.Comprehensive experiments on the LibriSpeech corpus show that the proposed Speech2C can relatively reduce the word error rate (WER) by 19.2% over the method without decoder pre-training, and also outperforms significantly the state-of-the-art wav2vec 2.0 and Hu-BERT on fine-tuning subsets of 10h and 100h.We release our code and model at https://github.com/microsoft/SpeechT5/tree/main/Speech2C.
Junyi Ao, Shujie Liu 0001, Haizhou Li 0001, Tom Ko, Li-Rong Dai 0001, Jinyu Li 0001, Yao Qian, Furu Wei
INTERSPEECH9
2022 A Comprehensive Study on Self-Supervised Distillation for Speaker Representation Learning
abstract
In real application scenarios, it is often challenging to obtain a large amount of labeled data for speaker representation learning due to speaker privacy concerns. Self-supervised learning with no labels has become a more and more promising way to solve it. Compared with contrastive learning, self-distilled approaches use only positive samples in the loss function and thus are more attractive. In this paper, we present a comprehensive study on self-distilled self-supervised speaker representation learning, especially on critical data augmentation. Our proposed strategy of audio perturbation augmentation has pushed the performance of the speaker representation to a new limit. The experimental results show that our model can achieve a new SoTA on Voxceleb 1 speaker verification evaluation benchmark (i.e., equal error rate (EER) 2.505%, 2.473%, and 4.791 % for trial Vox1-O, Vox1-E and Vox1-H, respectively), discarding any speaker labels in the training phase.
Zhengyang Chen, Yao Qian, Bing Han 0008, Yanmin Qian, Michael Zeng 0001
SLT2
2021 Speech-Language Pre-Training for End-to-End Spoken Language Understanding
abstract
End-to-end (E2E) spoken language understanding (SLU) can infer semantics directly from speech signal without cascading an automatic speech recognizer (ASR) with a natural language understanding (NLU) module. However, paired utterance recordings and corresponding semantics may not always be available or sufficient to train an E2E SLU model in a real production environment. In this paper, we propose to unify a well-optimized E2E ASR encoder (speech) and a pre-trained language model encoder (language) into a transformer decoder. The unified speech-language pre-trained model (SLP) is continually enhanced on limited labeled data from a target domain by using a conditional masked language model (MLM) objective, and thus can effectively generate a sequence of intent, slot type, and slot value for given input speech in the inference. The experimental results on two public corpora show that our approach to E2E SLU is superior to the conventional cascaded method. It also outperforms the present state-of-the-art approaches to E2E SLU with much less paired data.
Yao Qian, Ximo Bian, Yu Shi 0001, Naoyuki Kanda, Leo Shen, Michael Zeng 0001
ICASSP1
2021 UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled Data
abstract
In this paper, we propose a unified pre-training approach called UniSpeech to learn speech representations with both labeled and unlabeled data, in which supervised phonetic CTC learning and phonetically-aware contrastive self-supervised learning are conducted in a multi-task learning manner. The resultant representations can capture information more correlated with phonetic structures and improve the generalization across languages and domains. We evaluate the effectiveness of UniSpeech for cross-lingual representation learning on public CommonVoice corpus. The results show that UniSpeech outperforms self-supervised pretraining and supervised transfer learning for speech recognition by a maximum of 13.4% and 26.9% relative phone error rate reductions respectively (averaged over all testing languages). The transferability of UniSpeech is also verified on a domain-shift speech recognition task, i.e., a relative word error rate reduction of 6% against the previous approach.
Chengyi Wang 0002, Yu Wu 0012, Yao Qian, Ken'ichi Kumatani, Shujie Liu 0001, Furu Wei, Michael Zeng 0001, Xuedong Huang 0001
ICML3
2021 Automated Scoring of Spontaneous Speech from Young Learners of English Using Transformers
abstract
This study explores the use of Transformer-based models for the automated assessment of children's non-native spontaneous speech. Traditional approaches for this task have relied heavily on delivery features (e.g., fluency), whereas the goal of the current study is to build automated scoring models based solely on transcriptions in order to see how well they capture additional aspects of speaking proficiency (e.g., content appropriateness, vocabulary, and grammar) despite the high word error rate (WER) of automatic speech recognition (ASR) on children's non-native spontaneous speech. Transformer-based models are built using both manual transcriptions and ASR hypotheses, and versions of the models that incorporated the prompt text were investigated in order to more directly measure content appropriateness. Two baseline systems were used for comparison, including an attention-based Long Short-Term Memory (LSTM) Recurrent Neural Network (RNN) and a Support Vector Regressor (SVR) with manually engineered content-related features. Experimental results demonstrate the effectiveness of the Transformer-based models: the automated prompt-aware model using ASR hypotheses achieves a Pearson correlation coefficient (r) with holistic proficiency scores provided by human experts of 0.835, outperforming both the attention-based RNN-LSTM baseline (r = 0.791) and the SVR baseline (r = 0.767).
Keelan Evanini, Yao Qian, Matthew Mulholland
SLT3
2020 Discriminative Transfer Learning for Optimizing ASR and Semantic Labeling in Task-Oriented Spoken Dialog
Yao Qian, Yu Shi 0001, Michael Zeng 0001
INTERSPEECH1
2019 Native Language Identification from Raw Waveforms Using Deep Convolutional Neural Networks with Attentive Pooling
abstract
Automatic detection of an individual's native language (L1) based on speech data from their second language (L2) can be useful for informing a variety of speech applications such as automatic speech recognition (ASR), speaker recognition, voice biometrics, and computer assisted language learning (CALL). Previously proposed systems for native language identification from L2 acoustic signals rely on traditional feature extraction pipelines to extract relevant features such as mel-filterbanks, cepstral coefficients, i-vectors, etc. In this paper, we present a fully convolutional neural network approach that is trained end-to-end to predict the native language of the speaker directly from the raw waveforms, thereby removing the feature extraction step altogether. Experimental results using this approach on a database of 11 different L1s suggest that the learnable convolutional layers of our proposed attention-based end-to-end model extract meaningful features from raw waveforms. Further, the attentive pooling mechanism in our proposed network enables our model to focus on the most discriminative features leading to improvements over the conventional baseline.
Rutuja Ubale, Vikram Ramanarayanan, Yao Qian, Keelan Evanini, Chee Wee Leong, Chong Min Lee
ASRU3
2019 Using Very Deep Convolutional Neural Networks to Automatically Detect Plagiarized Spoken Responses
abstract
This study focuses on the automatic plagiarism detection in the context of high-stakes spoken language proficiency assessment, in which some test takers may attempt to game the test by memorizing prepared source materials before the test and then adapting them on-the-fly during the test to produce their spoken responses. When trying to identify such instances of plagiarism, experienced human raters attempt to find salient matching expressions that appear both in potential source materials and the test responses. This motivates an approach that visualizes a grid of lexical matches between a test response and a source and then applies state-of-the-art image recognition techniques to detect patterns of matching sequences. This study employs Inception networks-very deep convolutional neural networks-to build automatic detection models. The system achieves an F1-score of 79.6% on the class of plagiarized responses outperforming a baseline system based on word sequence matching (F1-score of 74.1%).
Keelan Evanini, Yao Qian, Klaus Zechner
ASRU3
2019 Neural Approaches to Automated Speech Scoring of Monologue and Dialogue Responses
abstract
We present Neural Network (NN) approaches to the automated assessment of non-native spontaneous speech in a monologic task and a simulated dialogic task. Three attention-based Bidirectional Long Short-Term Memory (BLSTM) Recurrent Neural Networks (RNN) are employed to learn three dimensions (i.e., delivery, language use, and content) of scoring rubrics for the spoken responses. The prompts or turn history information are encoded to low-dimensional vectors by either a BLSTM-RNN or an end-to-end memory network (MemN2N) and used as the conditions of the inputs of the NN for rating the subscore of content. The three subscores are fused together to generate a holistic score. The experimental results show that our approaches significantly outperform the conventional approaches to speech scoring and the correlations of automatically predicted scores with the reference human scores are higher than human-human agreement levels for both tasks.
Yao Qian, Patrick L. Lange, Keelan Evanini, Robert A. Pugh, Rutuja Ubale, Matthew Mulholland
ICASSP1
2019 Automated Estimation of Oral Reading Fluency During Summer Camp e-Book Reading with MyTurnToRead
Anastassia Loukina, Beata Beigman Klebanov, Patrick L. Lange, Yao Qian, Binod Gyawali, Nitin Madnani, Abhinav Misra, Klaus Zechner, John Sabatini 0001
INTERSPEECH4
2019 Automatic Detection of Off-Topic Spoken Responses Using Very Deep Convolutional Neural Networks
Su-Youn Yoon, Keelan Evanini, Klaus Zechner, Yao Qian
INTERSPEECH5
2019 Scoring Interactional Aspects of Human-Machine Dialog for Language Learning and Assessment using Text Features
abstract
While there has been much work in the language learning and assessment literature on human and automated scoring of essays and short constructed responses, there is little to no work examining text features for scoring of dialog data, particularly interactional aspects thereof, to assess conversational proficiency over and above constructed response skills.Our work bridges this gap by investigating both human and automated approaches towards scoring human-machine text dialog in the context of a real-world language learning application.We collected conversational data of human learners interacting with a cloud-based standards-compliant dialog system, triple-scored these data along multiple dimensions of conversational proficiency, and then analyzed the performance trends.We further examined two different approaches to automated scoring of such data and show that these approaches are able to perform at or above par with human agreement for a majority of dimensions of the scoring rubric.
Vikram Ramanarayanan, Matthew Mulholland, Yao Qian
SIGdial3
2018 End-to-End Neural Network Based Automated Speech Scoring
abstract
In recent years, machine learning models for automated speech scoring systems were mainly built using data-driven approaches with handcrafted features as one of the main components. However, the remarkable successes of deep learning (DL) technology in a variety of machine learning tasks has demonstrated its effectiveness in extracting features. Although there have been some efforts in utilizing DL technology for the automated speech scoring task, a thorough investigation of learning useful features is still missing. In this paper, we propose an end-to-end solution that consists of using deep neural network models to encode both lexical and acoustical cues to learn predictive features automatically. Experiments also confirm the effectiveness of our proposed solution compared to conventional methods based on handcrafted features.
Lei Chen 0004, Jidong Tao, Shabnam Ghaffarzadegan, Yao Qian
ICASSP4
2018 Improvements to an Automated Content Scoring System for Spoken CALL Responses: the ETS Submission to the Second Spoken CALL Shared Task
Keelan Evanini, Matthew Mulholland, Rutuja Ubale, Yao Qian, Robert A. Pugh, Vikram Ramanarayanan, Aoife Cahill
INTERSPEECH4
2018 A Prompt-Aware Neural Network Approach to Content-Based Scoring of Non-Native Spontaneous Speech
abstract
We present a neural network approach to the automated assessment of non-native spontaneous speech in a listen and speak task. An attention-based Long Short-Term Memory (LSTM) Recurrent Neural Network (RNN) is used to learn the relations (scoring rubrics) between the spoken responses and their assigned scores. Each prompt (listening material) is encoded as a vector in a low-dimensional space and then employed as a condition of the inputs of the attention LSTM-RNN. The experimental results show that our approach performs as well as the strong baseline of a Support Vector Regressor (SVR) using content-related features, i.e., a correlation of r = 0.806 with holistic proficiency scores provided by humans, without doing any feature engineering. The prompt-encoded vector improves the discrimination between the high-scoring sample and low-scoring sample, and it is more effective in grading responses to unseen prompts, which have no corresponding responses in the training set.
Yao Qian, Rutuja Ubale, Matthew Mulholland, Keelan Evanini
SLT1
2018 Exploring End-To-End Attention-Based Neural Networks For Native Language Identification
abstract
Automatic identification of speakers' native language (L1) based on their speech in a second language (L2) is a challenging research problem that can aid several spoken language technologies such as automatic speech recognition (ASR), speaker recognition, and voice biometrics in interactive voice applications. End-to-end learning, in which the features and the classification model are learned jointly in a single system, is an emerging field in the areas of speech recognition, speaker verification and spoken language understanding. In this paper, we present our study on attention-based end-to-end modeling for native language identification on a database of 11 different L1s. Using this methodology, we can determine the native language of the speaker directly from the raw acoustic features. Experimental results from our study show that our best end-to-end model can achieve promising results by capturing speech commonalities across L1s using an attention mechanism. In addition, fusion of proposed systems with the baseline system leads to significant performance improvements.
Rutuja Ubale, Yao Qian, Keelan Evanini
SLT2
2017 Improving native language (L1) identifation with better VAD and TDNN trained separately on native and non-native English corpora
abstract
Identifying a speaker's native language (L1), i.e., mother tongue, based upon non-native English (L2) speech input, is both challenging and useful for many human-machine voice interface applications, e.g., computer assisted language learning (CALL). In this paper, we improve our sub-phone TDNN based i-vector approach to L1 recognition with a more accurate TDNN-derived VAD and a highly discriminative classifier. Two TDNNs are separately trained on native and non-native English, LVCSR corpora, for contrasting their corresponding sub-phone posteriors and resultant supervectors. The derived i-vectors are then exploited for improving the performance further. Experimental results on a database of 25 L1s show a 3.1% identification rate improvement, from 78.7% to 81.8%, compared with a high performance baseline system which has already achieved the best published results on the 2016 ComParE corpus of only 11 L1s. The statistical analysis of the features used in our system provides useful findings, e.g. pronunciation similarity among the non-native English speakers with different L1s, for research on second-language (L2) learning and assessment.
Yao Qian, Keelan Evanini, Patrick L. Lange, Robert A. Pugh, Rutuja Ubale, Frank K. Soong
ASRU1
2017 Exploring ASR-free end-to-end modeling to improve spoken language understanding in a cloud-based dialog system
abstract
Spoken language understanding (SLU) in dialog systems is generally performed using a natural language understanding (NLU) model based on the hypotheses produced by an automatic speech recognition (ASR) system. However, when new spoken dialog applications are built from scratch in real user environments that often have sub-optimal audio characteristics, ASR performance can suffer due to factors such as the paucity of training data or a mismatch between the training and test data. To address this issue, this paper proposes an ASR-free, end-to-end (E2E) modeling approach to SLU for a cloud-based, modular spoken dialog system (SDS). We evaluate the effectiveness of our approach on crowdsourced data collected from non-native English speakers interacting with a conversational language learning application. Experimental results show that our approach is particularly promising in situations with low ASR accuracy. It can further improve the performance of a sophisticated CNN-based SLU system with more accurate ASR hypotheses by fusing the scores from E2E system, i.e., the overall accuracy of SLU is improved from 85.6% to 86.5%.
Yao Qian, Rutuja Ubale, Vikram Ramanarayanan, Patrick L. Lange, David Suendermann-Oeft, Keelan Evanini, Eugene Tsuprun
ASRU1
2017 Bidirectional LSTM-RNN for Improving Automated Assessment of Non-Native Children's Speech
Yao Qian, Keelan Evanini, Chong Min Lee, Matthew Mulholland
INTERSPEECH1
2017 Improving Sub-Phone Modeling for Better Native Language Identification with Non-Native English Speech
Yao Qian, Keelan Evanini, David Suendermann-Oeft, Robert A. Pugh, Patrick L. Lange, Hillary Molloy, Frank K. Soong
INTERSPEECH1
2016 Unsupervised speaker adaptation for DNN-based TTS synthesis
abstract
Multi-speaker TTS trained with a general DNN has outperformed individually modelled baseline [1]. Multi-speaker DNN takes advantages of larger amount of training data from multiple speakers to find robust transformations in the hidden layers and covers more speaker variability in the output regression layer. In this paper, we propose a new approach to unsupervised speaker adaptation with multi-speaker DNN. It takes advantage of shared hidden transformation to search for the labels of unlabelled acoustic frames and the found labels are used for speaker adaption. Experimental results show that the new approach of unsupervised adaptation can achieve comparable performance with supervised adaptation both objectively and subjectively. We further extend it to cross-lingual adaptation. It can remove non-native accent and improve the naturalness while keep the same speaker's characteristics.
Yuchen Fan 0001, Yao Qian, Frank K. Soong, Lei He 0005
ICASSP2
2016 Speaker and language factorization in DNN-based TTS synthesis
abstract
We have successfully proposed to use multi-speaker modelling in DNN-based TTS synthesis for improved voice quality with limited available data from a speaker. In this paper, we propose a new speaker and language factorized DNN, where speaker-specific layers are used for multi-speaker modelling, and shared layers and language-specific layers are employed for multi-language, linguistic feature transformation. Experimental results on a speech corpus of multiple speakers in both Mandarin and English show that the proposed factorized DNN can not only achieve a similar voice quality as that of a multi-speaker DNN, but also perform polyglot synthesis with a monolingual speaker's voice.
Yuchen Fan 0001, Yao Qian, Frank K. Soong, Lei He 0005
ICASSP2
2016 A comparison of ASR and human errors for transcription of non-native spontaneous speech
abstract
In this paper, we compare ASR and human transcriptions of non-native speech to investigate to what extent the accuracy and the patterns of errors of a modern ASR system match those of human listeners in the context of automated assessment of L2 English language proficiency. We obtained multiple naïve transcriptions of short fragments of non-native spontaneous speech with different proficiency levels using crowdsourcing and matched these against the output of an ASR system. We compare WER and recall at the fragment level and consider human-ASR agreement at the word level. We find that we are able to attain a commensurate level of transcription quality using ASR, but the patterns of errors between the two groups differ at the word level.
Matthew Mulholland, Melissa Lopez, Keelan Evanini, Anastassia Loukina, Yao Qian
ICASSP5
2016 Noise and Metadata Sensitive Bottleneck Features for Improving Speaker Recognition with Non-Native Speech Input
Yao Qian, Jidong Tao, David Suendermann-Oeft, Keelan Evanini, Alexei V. Ivanov, Vikram Ramanarayanan
INTERSPEECH1
2016 Self-Adaptive DNN for Improving Spoken Language Proficiency Assessment
Yao Qian, Keelan Evanini, David Suendermann-Oeft
INTERSPEECH1
2016 Learning Distributed Word Representations For Bidirectional LSTM Recurrent Neural Network
abstract
Peilu Wang, Yao Qian, Frank K. Soong, Lei He, Hai Zhao. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Peilu Wang, Yao Qian, Frank K. Soong, Lei He 0005
HLT-NAACL2
2016 Modeling F0 trajectories in hierarchically structured deep neural networks
Xiang Yin 0002, Yao Qian, Frank K. Soong, Lei He 0005, Zhen-Hua Ling, Li-Rong Dai 0001
Speech Commun.3
2015 Using bidirectional lstm recurrent neural networks to learn high-level abstractions of sequential features for automated scoring of non-native spontaneous speech
abstract
We introduce a new method to grade non-native spoken language tests automatically. Traditional automated response grading approaches use manually engineered time-aggregated features (such as mean length of pauses). We propose to incorporate general time-sequence features (such as pitch) which preserve more information than time-aggregated features and do not require human effort to design. We use a type of recurrent neural network to jointly optimize the learning of high level abstractions from time-sequence features with the time-aggregated features. We first automatically learn high level abstractions from time-sequence features with a Bidirectional Long Short Term Memory (BLSTM) and then combine the high level abstractions with time-aggregated features in a Multilayer Perceptron (MLP)/Linear Regression (LR). We optimize the BLSTM and the MLP/LR jointly. We find such models reach the best performance in terms of correlation with human raters. We also find that when there are limited time-aggregated features available, our model that incorporates time-sequence features improves performance drastically.
Zhou Yu 0005, Vikram Ramanarayanan, David Suendermann-Oeft, Klaus Zechner, Lei Chen 0004, Jidong Tao, Aliaksei Ivanou, Yao Qian
ASRU9
2015 Multi-speaker modeling and speaker adaptation for DNN-based TTS synthesis
abstract
In DNN-based TTS synthesis, DNNs hidden layers can be viewed as deep transformation for linguistic features and the output layers as representation of acoustic space to regress the transformed linguistic features to acoustic parameters. The deep-layered architectures of DNN can not only represent highly-complex transformation compactly, but also take advantage of huge amount of training data. In this paper, we propose an approach to model multiple speakers TTS with a general DNN, where the same hidden layers are shared among different speakers while the output layers are composed of speaker-dependent nodes explaining the target of each speaker. The experimental results show that our approach can significantly improve the quality of synthesized speech objectively and subjectively, comparing with speech synthesized from the individual, speaker-dependent DNN-based TTS. We further transfer the hidden layers for a new speaker with limited training data and the resultant synthesized speech of the new speaker can also achieve a good quality in term of naturalness and speaker similarity.
Yuchen Fan 0001, Yao Qian, Frank K. Soong, Lei He 0005
ICASSP2
2015 Word embedding for recurrent neural network based TTS synthesis
abstract
The current state of the art TTS synthesis can produce synthesized speech with highly decent quality if rich segmental and suprasegmental information are given. However, some suprasegmental features, e.g., Tone and Break (TOBI), are time consuming due to being manually labeled with a high inconsistency among different annotators. In this paper, we investigate the use of word embedding, which represents word with low dimensional continuous-valued vector and being assumed to carry a certain syntactic and semantic information, for bidirectional long short term memory (BLSTM), recurrent neural network (RNN) based TTS synthesis. Experimental results show that word embedding can significantly improve the performance of BLSTM-RNN based TTS synthesis without using features of TOBI and Part of Speech (POS).
Peilu Wang, Yao Qian, Frank K. Soong, Lei He 0005
ICASSP2
2015 Sequence generation error (SGE) minimization based deep neural networks training for text-to-speech synthesis
Yuchen Fan 0001, Yao Qian, Frank K. Soong, Lei He 0005
INTERSPEECH2
2015 Improved mispronunciation detection with deep neural network trained acoustic models and transfer learning based logistic regression classifiers
Wenping Hu, Yao Qian, Frank K. Soong
Speech Commun.2
2014 A DNN-based acoustic modeling of tonal language and its application to Mandarin pronunciation training
abstract
In this paper we investigate a Deep Neural Network (DNN) based approach to acoustic modeling of tonal language and assess its speech recognition performance with different features and modeling techniques. Mandarin Chinese, the most widely spoken tonal language, is chosen for testing the tone related ASR performance. Furthermore, the DNN-trained, tone-sensitive model is evaluated in automatic detection of mispronunciation among L2 Mandarin learners. The best DNN-HMM acoustic model of tonal syllable (initial and tonal final), trained with embedded F0 features, has shown improved ASR performance, when compared with the baseline DNN system of 39 MFCC features. The proposed system achieves better ASR performance than the baseline system, i.e., by 32% and 35% in relative tone error rate reduction and 20% and 23% in relative tonal syllable error rate reduction, for female and male speakers, respectively. In a speech database of L2 Mandarin learners (native speakers of European languages), 2% equal error rate reduction, from 27.5% to 25.5%, has been obtained with our DNN-HMM system in detecting mispronunciations, compared with the baseline system.
Wenping Hu, Yao Qian, Frank K. Soong
ICASSP2
2014 On the training aspects of Deep Neural Network (DNN) for parametric TTS synthesis
abstract
Deep Neural Network (DNN), which can model a long-span, intricate transform compactly with a deep-layered structure, has recently been investigated for parametric TTS synthesis with a fairly large corpus (33,000 utterances) [6]. In this paper, we examine DNN TTS synthesis with a moderate size corpus of 5 hours, which is more commonly used for parametric TTS training. DNN is used to map input text features into output acoustic features (LSP, F0 and V/U). Experimental results show that DNN can outperform the conventional HMM, which is trained in ML first and then refined by MGE. Both objective and subjective measures indicate that DNN can synthesize speech better than HMM-based baseline. The improvement is mainly on the prosody, i.e., the RMSE of natural and generated F0 trajectories by DNN is improved by 2 Hz. This benefit is likely from the key characteristics of DNN, which can exploit feature correlations, e.g., between F0 and spectrum, without using a more restricted, e.g. diagonal Gaussian probability family. Our experimental results also show: the layer-wise BP pre-training can drive weights to a better starting point than random initialization and result in a more effective DNN; state boundary info is important for training DNN to yield better synthesized speech; and a hyperbolic tangent activation function in DNN hidden layers yields faster convergence than a sigmoidal one.
Yao Qian, Yuchen Fan 0001, Wenping Hu, Frank K. Soong
ICASSP1
2014 TTS synthesis with bidirectional LSTM based recurrent neural networks
abstract
Feed-forward, Deep neural networks (DNN)-based text-tospeech (TTS) systems have been recently shown to outperform decision-tree clustered context-dependent HMM TTS systems [1, 4]. However, the long time span contextual effect in a speech utterance is still not easy to accommodate, due to the intrinsic, feed-forward nature in DNN-based modeling. Also, to synthesize a smooth speech trajectory, the dynamic features are commonly used to constrain speech parameter trajectory generation in HMM-based TTS [2]. In this paper, Recurrent Neural Networks (RNNs) with Bidirectional Long Short Term Memory (BLSTM) cells are adopted to capture the correlation or co-occurrence information between any two instants in a speech utterance for parametric TTS synthesis. Experimental results show that a hybrid system of DNN and BLSTM-RNN, i.e., lower hidden layers with a feed-forward structure which is cascaded with upper hidden layers with a bidirectional RNN structure of LSTM, can outperform either the conventional, decision tree-based HMM, or a DNN TTS system, both objectively and subjectively. The speech trajectory generated by the BLSTM-RNN TTS is fairly smooth and no dynamic constraints are needed.
Yuchen Fan 0001, Yao Qian, Fenglong Xie, Frank K. Soong
INTERSPEECH2
2014 Sequence error (SE) minimization training of neural network for voice conversion
abstract
Neural network (NN) based voice conversion, which employs a nonlinear function to map the features from a source to a target speaker, has been shown to outperform GMM-based voice conversion approach [4-7]. However, there are still limitations to be overcome in NN-based voice conversion, e.g. NN is trained on a Frame Error (FE) minimization criterion and the corresponding weights are adjusted to minimize the error squares over the whole source-target, stereo training data set. In this paper, we use the idea of sentence optimization based, minimum generation error (MGE) training in HMM-based TTS synthesis, and modify the FE minimization to Sequence Error (SE) minimization in NN training for voice conversion. The conversion error over a training sentence from a source speaker to a target speaker is minimized via a gradient descent-based, back propagation (BP) procedure. Experimental results show that the speech converted by the NN, which is first trained with frame error minimization and then refined with sequence error minimization, sounds subjectively better than the converted speech by NN trained with frame error minimization only. Scores on both naturalness and similarity to the target speaker are improved. Index Terms: voice conversion, neural network, pre-training, sequence error minimization
Fenglong Xie, Yao Qian, Yuchen Fan 0001, Frank K. Soong, Haifeng Li 0001
INTERSPEECH2
2014 Modeling DCT parameterized F0 trajectory at intonation phrase level with DNN or decision tree
Xiang Yin 0002, Yao Qian, Frank K. Soong, Lei He 0005, Zhen-Hua Ling, Li-Rong Dai 0001
INTERSPEECH3
2013 A fast table lookup based, statistical model driven non-uniform unit selection TTS
abstract
For multi-channel TTS applications, e.g. in a cloud service, it is highly desirable that high quality speech can be synthesized in low complexity. In this paper, we propose a fast table lookup based, statistical model driven approach to non-uniform unit selection TTS for that purpose. In TTS training, the voice font of all waveform segments is organized as a Gaussian kernel coded hash table and a table for storing quantized costs of all possible concatenation segment pairs. In synthesis, waveform segments with non-uniform lengths are first selected to construct a candidate lattice by looking up the Gaussian kernel coded hash table, and the best path is searched in the lattice by minimizing the accumulated concatenation scores, which are retrieved from the quantization table for possible concatenations. Experimental results show that the new approach can significantly reduce the search complexity while keep a high TTS voice quality.
Yao Qian, Frank K. Soong, Yundi Qian
ICASSP1
2013 A new DNN-based high quality pronunciation evaluation for computer-aided language learning (CALL)
Wenping Hu, Yao Qian, Frank K. Soong
INTERSPEECH2
2013 A Unified Trajectory Tiling Approach to High Quality Speech Rendering
abstract
It is technically challenging to make a machine talk as naturally as a human so as to facilitate “frictionless” interactions between machine and human. We propose a trajectory tiling-based approach to high-quality speech rendering, where speech parameter trajectories, extracted from natural, processed, or synthesized speech, are used to guide the search for the best sequence of waveform “tiles” stored in a pre-recorded speech database. We test the proposed unified algorithm in both Text-To-Speech (TTS) synthesis and cross-lingual voice transformation applications. Experimental results show that the proposed trajectory tiling approach can render speech which is both natural and highly intelligible. The perceived high quality of rendered speech is also confirmed in both objective and subjective evaluations.
Yao Qian, Frank K. Soong, Zhijie Yan
IEEE Trans. Speech Audio Process.1
2012 Turning a Monolingual Speaker into Multilingual for a Mixed-language TTS
Yao Qian, Frank K. Soong, Sheng Zhao 0002
INTERSPEECH2
2012 Tip tap tones: mobile microtraining of mandarin sounds
abstract
Learning a second language is hard, especially when the learner's brain must be retrained to identify sounds not present in his or her native language. It also requires regular practice, but many learners struggle to find the time and motivation. Our solution is to break down the challenge of mastering a foreign sound system into minute-long episodes of "microtraining" delivered through mobile gaming. We present the example of Tip Tap Tones - a mobile game with the purpose of helping learners acquire the tonal sound system of Mandarin Chinese. In a 3-week, 12-user study of this system, we found that an average of 71 minutes' gameplay significantly improved tone identification by around 25%, regardless of whether the underlying sounds had been used to train tone perception. Overall, results suggest that mobile microtraining is an efficient, effective, and enjoyable way to master the sounds of Mandarin Chinese, with applications to other languages and domains.
Darren Edge, Kai-Yin Cheng, Michael Whitney, Yao Qian, Zhijie Yan, Frank K. Soong
Mobile HCI4
2011 Improved F0 modeling and generation in voice conversion
abstract
F0 is an acoustic feature that varies largely from one speaker to an other. F0 is characterized by a discontinuity in the transition between voiced and unvoiced sounds that presents an obstacle to GMM modeling for use in voice conversion. A Multi-Space Distribution (MSD) [5] can be used to model unvoiced and voiced F0 regions in a linearly weighted mixture. However, the use of two incompatible probabilistic spaces, for example a continuous probability density for voiced observations, and a discrete probability for unvoiced observations, may result in an imprecise voiced/unvoiced (v/u) conversion in a maximum likelihood (ML) sense. In this paper we propose to use voicing strength, characterized by the normalized correlation coefficient magnitude, as calculated from F0 feature extraction, as an additional feature for improving F0 modeling and the v/u decision in the context of voice conversion. The proposed method was evaluated on male-to-female voice conversion tasks in both Mandarin and English. Objective tests showed that the approach is effective in reducing the Root Mean Square Error, while the results for subjective metrics including AB preference and ABX speaker similarity tests also showed gains.
Aki Kunikoshi, Yao Qian, Frank K. Soong, Nobuaki Minematsu
ICASSP2
2011 A frame mapping based HMM approach to cross-lingual voice transformation
abstract
Cross-lingual voice transformation is challenging when source language (L1) and target language (L2) are very different in corresponding phonetics and prosodies. We propose a frame mapping based HMM approach to this problem. The source speaker's speech data is first warped in frequency toward the target speaker by mapping corresponding formants of selected vowels. The parameter trajectories of the warped data are then "tiled" with the frames in target speaker's L2 data. The tiled new trajectories then form a simulated training set of target speaker in L1 and it is used to train an HMM TTS. With a bilingual (Mandarin and English) source speaker and a monolingual (English) target speaker, the frame mapping-based approach is capable of generating highly intelligible, good quality speech data in L1 (Mandarin), which sounds rather close to the target speaker. The good performance of the cross-lingual voice transformation is confirmed with speaker similarity, naturalness and intelligibility evaluations subjectively.
Yao Qian, Frank K. Soong
ICASSP1
2011 A New Phonetic Candidate Generator for Improving Search Query Efficiency
Yao Qian, Frank K. Soong
INTERSPEECH2
2011 Improved Prosody Generation by Maximizing Joint Probability of State and Longer Units
abstract
The current state-of-the-art hidden Markov model (HMM)-based text-to-speech (TTS) can produce highly intelligible, synthesized speech with decent segmental quality. However, its prosody, especially at phrase or sentence level, still tends to be bland. This blandness is partially due to the fact that the state-based HMM is inadequate in capturing global, hierarchical suprasegmental information in speech signals. In this paper, to improve the TTS prosody, longer units are first explicitly modeled with appropriate parametric distributions. The resultant models are then integrated with the state-based baseline models in generating better prosody by maximizing the joint probability. Experimental results in both Mandarin and English show consistent improvements over our baseline system with only state-based prosody model. The improvements are both objectively measurable and subjectively perceivable.
Yao Qian, Zhizheng Wu 0001, Boyang Gao, Frank K. Soong
IEEE Trans. Speech Audio Process.1
2010 RIch-context Unit Selection (RUS) approach to high quality TTS
abstract
This paper presents a Rich-context Unit Selection (RUS) approach to high quality speech synthesis. Based upon our previous work on rich context modeling, we use the corresponding parametric HMMs to represent waveform units and form a “sausage-like” lattice. A prune-and-search procedure is proposed, in which Kullback-Leibler divergence is adopted to select potential candidate units, and normalized cross-correlation is used as the final objective measure to search for the optimal unit path. The maximum cross-correlation criterion provides the optimal concatenation between successive units, in terms of spectral similarity, phase continuity and best connecting timing instants. Subjectively, both preference and MOS tests were conducted to compare RUS with our current Weight-table based Unit Selection (WUS) synthesis. Experimental results show that the voice quality of synthesized speech is significantly improved by RUS over the conventional WUS.
Zhijie Yan, Yao Qian, Frank K. Soong
ICASSP2
2010 Improved modeling for F0 generation and V/U decision in HMM-based TTS
abstract
The HMM-based TTS can produce a highly intelligible and decent quality voice. However, sometimes the synthesized speech exhibits perceptibly annoying glitches due to F0 extraction errors in the training data and voiced/unvoiced swapping errors in F0 generation. In the conventional MSD based F0 modeling [10], the dual but incompatible two probabilistic spaces, the continuous probability density for voiced observations or the discrete probability for unvoiced observations, prevent us from using likelihood based frame occupancy to alleviate the deteriorating effect of F0 extraction errors in training a more robust model for synthesis. In this paper, we propose a new approach to improved modeling the piece-wise continuous F0 trajectory and v/u decision for HMM-based TTS. Voicing strength, characterized by the normalized correlation coefficient magnitude calculated in F0 feature extraction, is used as an additional feature in F0 modeling and for v/u decision. Experimental results show the new approach to F0 modeling and generation outperforms MSD-HMM method and a newly proposed GTD-HMM method [9] significantly. The improvements are both objectively measurable and subjectively perceivable.
Frank K. Soong, Yao Qian, Zhijie Yan, Jielin Pan, Yonghong Yan 0002
ICASSP3
2010 An HMM trajectory tiling (HTT) approach to high quality TTS
Yao Qian, Zhijie Yan, Yi-Jian Wu, Frank K. Soong, Xin Zhuang, Shengyi Kong
INTERSPEECH1
2010 Formant-based frequency warping for improving speaker adaptation in HMM TTS
Xin Zhuang, Yao Qian, Frank K. Soong, Yi-Jian Wu
INTERSPEECH2
2009 State mapping for cross-language speaker adaptation in TTS
abstract
Cross-language speaker adaptation has many interesting applications, e.g. speech-to-speech translation. However, in cross-language speaker adaptation, a common phoneme set, assumed to be used by different speakers of the same language, does not exist any longer. Instead, a nearest neighbor based phoneme mapping from one language to the other has been adopted. In this study, we used our recently proposed sub-phonemic HMM state mapping for cross-language adaptations. The sub-phonemic HMM states, due to their phonetic segment nature, tend to be more sharable across different languages than phonemes. Kullback-Leibler divergence, an information-theoretic measure, is chosen here to measure the similarity between given states in different languages. Experimental results show that new state mapping outperforms the phoneme mapping baseline system in terms of three objective measures: log spectral distance, F0 adaptation error and F0 correlations. In comparing with intra-language adaptation, the cross-language result of the new algorithm is also fairly decent.
Yao Qian, Frank K. Soong
ICASSP3
2009 Improved prosody generation by maximizing joint likelihood of state and longer units
abstract
The current state-of-art HMM-bsed TTS can produce highly intelligible output speech and deliver a decent segmental quality. However, its prosody, especially at the phrase or sentence level, tends to be bland. The blandness of synthesized prosody is partially due to the fact that a state-based HMM is rather inadequate in modeling a global, hierarchical prosodic structure at a sentence or phrase level. In this study, the prosody of longer units are first modeled explicitly by appropriate parametric distributions. The resultant models are then integrated with the state-level baseline models to generate an optimal prosody by maximizing the joint likelihood of all, from state to longer, units. Experimental results in both Mandarin and English show consistent improvements over the state-based baseline system. The improvements are both objectively measurable and subjectively perceivable.
Yao Qian, Zhizheng Wu 0001, Frank K. Soong
ICASSP1
2009 A minimum v/u error approach to F0 generation in HMM-based TTS
Yao Qian, Frank K. Soong, Zhizheng Wu 0001
INTERSPEECH1
2009 Rich context modeling for high quality HMM-based TTS
abstract
This paper presents a rich context modeling approach to high quality HMM-based speech synthesis. We first analyze the over-smoothing problem in conventional decision tree tyingbased HMM, and then propose to model the training speech tokens with rich context models. Special training procedure is adopted for reliable estimation of the rich context model parameters. In synthesis, a search algorithm following a contextbased pre-selection is performed to determine the optimal rich context model sequence which generates natural and crisp output speech. Experimental results show that spectral envelopes synthesized by the rich context models are with crisper formant structures and evolve with richer details than those obtained by the conventional models. The speech quality improvement is also perceived by listeners in a subjective preference test, in which 76% of the sentences synthesized using rich context modeling are preferred. Index Terms: HMM-based TTS, rich context modeling
Zhijie Yan, Yao Qian, Frank K. Soong
INTERSPEECH2
2009 A Multi-Space Distribution (MSD) and two-stream tone modeling approach to Mandarin speech recognition
Yao Qian, Frank K. Soong
Speech Commun.1
2009 A Cross-Language State Sharing and Mapping Approach to Bilingual (Mandarin-English) TTS
abstract
We propose a hidden Markov model (HMM)-based bilingual (Mandarin and English) text-to-speech (TTS) system to synthesize natural speech for given bilingual text. A simple baseline system consisting of two independent monolingual HMM synthesizers is built first from corresponding Mandarin and English data recorded by a bilingual speaker. A new, mixed language TTS is then constructed by asking language-independent and language-specific questions for sharing HMM states across the two languages in decision-tree based clustering. By sharing states, the new system has a smaller footprint than the baseline system. Speech synthesized by the new system sounds very similar to the baseline for non-mixed, Mandarin or English, monolingual sentences but much better for mixed-language sentences. This higher quality of mixed-language output is confirmed by a preference score, 60.2% to 39.8%, in a subjective listening test. A cross-language state mapping algorithm is further proposed for cross-language synthesis when only monolingual (English) recorded data from a source language speaker is available. Mandarin speech is then synthesized with the HMM model parameters in the nearest neighbor leaf nodes of the English decision tree. The nearest neighbor is measured with the Kullback-Leibler divergence (KLD) and mappings between leaf nodes in the decision trees of the source and target languages are established via the speech data recorded by a different, bilingual speaker. High voice (speaker) similarity is preserved in the synthesized target language sentences by using the recording of a source language from a monolingual speaker. Perceptual test results conducted on synthesized Mandarin speech show 1) high intelligibility which is confirmed by a Chinese character transcription accuracy of 92.1% and 2) decent speech quality with an average MOS score of 3.1.
Yao Qian, Frank K. Soong
IEEE Trans. Speech Audio Process.1
2008 A cross-language state mapping approach to bilingual (Mandarin-English) TTS
abstract
We propose a cross-language state mapping approach to HMM-based bilingual TTS. Two language-dependent decision trees are built first with a bilingual speech database recorded by a single speaker. A state mapping for every leaf node in the decision tree of a target language is created by finding the nearest leaf node in the tree of a source language. Kullback-Leibler divergence between two distributions is used to find the nearest leaf node. To synthesize target language speech by a monolingual, (source language) speaker's voice, we find HMM parameters trained by the monolingual (source language) speaker in the mapped leaf nodes. Similar mappings can be constructed by reversing the source and target languages. With these bi-directional cross-lingual mappings, we can synthesize bilingual or mixed-code speech by HMMs trained by any monolingual speaker. High voice (speaker) similarity is preserved in synthesized speech of the target language. Two perceptual tests on synthesized Mandarin speech confirms high intelligibility with a Chinese character transcription accuracy of 92.1% and an MOS score of 3.08.
Yao Qian, Frank K. Soong, Gongshen Liu
ICASSP2
2008 Duration refinement by jointly optimizing state and longer unit likelihood
Boyang Gao, Yao Qian, Zhizheng Wu 0001, Frank K. Soong
INTERSPEECH2
2008 Generating natural F0 trajectory with additive trees
Yao Qian, Frank K. Soong
INTERSPEECH1
2008 A real-time text to audio-visual speech synthesis system
Xiaojun Qian, Yao Qian, Frank K. Soong
INTERSPEECH4
2008 Prosody for Mandarin speech recognition: a comparative study of read and spontaneous speech
Yu Ting Yeung, Yao Qian, Tan Lee, Frank K. Soong
INTERSPEECH2
2008 Tone-enhanced generalized character posterior probability (GCPP) for Cantonese LVCSR
Yao Qian, Frank K. Soong, Tan Lee
Comput. Speech Lang.1
2007 Robust F0 modeling for Mandarin speech recognition in noise
Sheng Qiang, Yao Qian, Frank K. Soong, Congfu Xu
INTERSPEECH2
2006 Tone-Enhanced Generalized Character Posterior Probability (GCPP) for Cantonese LVCSR
abstract
Tone-enhanced, generalized character posterior probability (GCPP), a generalized form of posterior probability at subword (Chinese character) level, is proposed as a rescoring metric for improving Cantonese LVCSR performance. The search network is constructed first by converting the original word graph to a restructured word graph, then a character graph and finally, a character confusion network (CCN). Based upon GCPP enhanced with tone information, the character error rate (CER) is minimized or the GCPP product is maximized over a chosen graph. Experimental results show that the tone enhanced GCPP can improve character error rate by up to 15.1%, relatively
Yao Qian, Frank K. Soong, Tan Lee
ICASSP (1)1
2006 A multi-space distribution (MSD) approach to speech recognition of tonal languages
Huanliang Wang, Yao Qian, Frank K. Soong, Jian-Lai Zhou, Jiqing Han 0001
INTERSPEECH2
2004 Tone information as a confidence measure for improving Cantonese LVCSR
Yao Qian, Tan Lee, Frank K. Soong
INTERSPEECH1
2004 Analysis and modeling of F0 contours for cantonese text-to-speech
abstract
For the generation of highly natural synthetic speech, the control of prosody is of primary importance. The fundamental frequency (F0) is one of the most important components of speech prosody. This research investigates the variation of F0 in continuous Cantonese speech, with the goal of establishing an effective mechanism of prosody control in Cantonese text-to-speech (TTS) applications. Cantonese is a commonly used Chinese dialect that is well known for being rich in tones. This article describes a simple yet effective approach to the analysis and modeling of F0. The surface F0 contour of a continuous Cantonese utterance is considered to be the combination of a global component--phrase-level intonation curve, and local components--syllable-level tone contoursA novel method of F0 normalization is proposed to separate the local components from the global one. As a result, the variation in tone contours is greatly reduced. Statistical analysis is performed for the phrase curves and context-dependent tone contours that are extracted from a large corpus of 1,200 utterances. Specifically, the analysis is focused on co-articulated tone contours for disyllabic words, cross-word contours, and phrase-initial tone contours. Based on the results of the analysis, a template-based model for F0 generation is established and integrated with a Cantonese TTS system. Subjective listening tests show that the proposed model significantly improves the naturalness of the output speech.
Tan Lee, Yao Qian
ACM Trans. Asian Lang. Inf. Process.3
2003 Overlapped di-tone modeling for tone recognition in continuous Cantonese speech
abstract
This paper presents a novel approach to tone recognition in continuous Cantonese speech based on overlapped di-tone Gaussian mixture models (ODGMM). The ODGMM is designed with special consideration on the fact that Cantonese tone identification relies more on the relative pitch level than on the pitch contour. A di-tone unit covers a group of two consecutive tone occurrences. The tone sequence carried by a Cantonese utterance can be considered as the connection of such di-tone units. Adjacent di-tone units overlap with each other by exactly one tone. For each di-tone unit, a GMM is trained with a 10-dimensional feature vector that characterizes the F0 movement within the unit. In particular, the di-tone models capture the relative deviation between the F0 levels of the two tones. Viterbi decoding algorithm is adopted to search for the optimal tone sequence, under the phonological constraints on syllable-tone combination. Experimental results show the ODGMM approach significantly outperforms the previously proposed methods for tone recognition in continuous Cantonese speech.
Yao Qian, Tan Lee
INTERSPEECH1
2002 Assigning phrase accent to Chinese Text-to-Speech system
abstract
This paper mainly solves the problem of phrase accent placement for Chinese Text-to-Speech (TIS) systems. Part-of-speech, prosodic boundary, word position, word unigram score and TF-IDF weight are employed in two speech corpus training by CART, which achieved the 80.01% and 77.15% of accuracy on testing data respectively. A perceptual experiment is designed to evaluate the phrase accent assignment. The result shows that it is important to assign phrase accent for Chinese ITS systems from the perceptual point of view, and that the performance of our automatic method is close to the manually annotation
Yao Qian
ICASSP1
2001 Segmenting unrestricted Chinese text into prosodic words instead of lexical words
abstract
This paper stresses the importance of converting a string of lexical words to that of prosodic words in text-to-speech (TTS) systems by presenting the surface differences and perceptual differences between them. A statistical rule based method and a classification and regression tree (CART) based method are proposed as solutions. Though ComplicatedSet based CART method performs the best, the achievement is obtained at the cost of heavy computation workloads needed by a parser. Statistical rule based method results in higher recall but lower precision, comparing to SimpleSet CART method. It is very difficult to tell which is better, since we don't know which affects naturalness more, precision or recall. Both of them require only lexicon word segmentation and part of speech (POS) tagging in the preprocessing stage, and are easily realized in TTS systems. Results of the preference test discloses that significant improvements on naturalness are perceived when lexical word strings are converted into prosodic word strings by our approach.
Yao Qian, Min Chu, Hu Peng
ICASSP1