VLDB 2026 Research / reviewers in the wild / expert
Xu Li 0015
dblp:25/3528-15
· DBLP profile ↗
34ranked-venue papers
8as first author
21since 2021 · last 2025
0000-0003-2954-3271ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 8 first-author · 19 since 2021Artificial intelligence and machine learning · 17 · 6 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fewer Hallucinations, More Verification: A Three-Stage LLM-Based Framework for ASR Error CorrectionabstractAutomatic Speech Recognition (ASR) error correction aims to correct recognition errors while preserving accurate text. Although traditional approaches demonstrate moderate effectiveness, LLMs offer a paradigm that eliminates the need for training and labeled data. However, directly using LLMs will encounter hallucinations problem, which may lead to the modification of the correct text. To address this problem, we propose the Reliable LLM Correction Framework (RLLMCF), which consists of three stages: (1) error pre-detection, (2) chain-of-thought sub-tasks iterative correction, and (3) reasoning process verification. The advantage of our method is that it does not require additional information or fine-tuning of the model, and ensures the correctness of the LLM correction under multipass programming. Experiments on AISHELL-1, AISHELL-2, and Librispeech show that the GPT-4o model enhanced by our framework achieves $21 \%, 11 \%, 9 \%$, and $11.4 \%$ relative reductions in CER/WER. Yangui Fang, Baixu Cheng, Xu Li 0015, Yu Xi, Chengwei Zhang 0002, Guohui Zhong |
ASRU | 4 |
| 2025 | Low-Resource Domain Adaptation for Speech LLMs via Text-Only Fine-TuningabstractRecent advances in automatic speech recognition (ASR) have combined speech encoders with large language models (LLMs) through projection, forming Speech LLMs with strong performance. However, adapting them to new domains remains challenging, especially in low-resource settings where paired speech-text data is scarce. We propose a text-only fine-tuning strategy for Speech LLMs using unpaired target-domain text without requiring additional audio. To preserve speech-text alignment, we introduce a real-time evaluation mechanism during fine-tuning. This enables effective domain adaptation while maintaining source-domain performance. Experiments on LibriSpeech, SlideSpeech, and Medical datasets show that our method achieves competitive recognition performance, with minimal degradation compared to full audio-text fine-tuning. It also improves generalization to new domains without catastrophic forgetting, highlighting the potential of text-only fine-tuning for low-resource domain adaptation of ASR. Yangui Fang, Xu Li 0015, Yu Xi, Chengwei Zhang 0002, Guohui Zhong, Kai Yu 0004 |
ASRU | 3 |
| 2025 | Language-Queried Target Sound Extraction Without Parallel Training DataabstractLanguage-queried target sound extraction (TSE) aims to extract specific sounds from mixtures based on language queries. Traditional fully-supervised training schemes require extensively annotated parallel audio-text data, which are labor-intensive. We introduce a parallel-data-free training scheme, requiring only unlabelled audio clips for TSE model training by utilizing the contrastive language-audio pre-trained model (CLAP). In a vanilla parallel-data-free training stage, target audio is encoded using the pre-trained CLAP audio encoder to form a condition embedding, while during testing, user language queries are encoded by CLAP text encoder as the condition embedding. This vanilla approach assumes perfect alignment between text and audio embeddings, which is unrealistic. Two major challenges arise from training-testing mismatch: the persistent modality gap between text and audio and the risk of overfitting due to the exposure of rich acoustic details in target audio embedding during training. To address this, we propose a retrieval-augmented strategy. Specifically, we create an embedding cache using audio captions generated by a large language model (LLM). During training, target audio embeddings retrieve text embeddings from this cache to use as condition embeddings, ensuring consistent modalities between training and testing and eliminating information leakage. Extensive experiment results show that our retrieval-augmented approach achieves consistent and notable performance improvements over existing state-of-the-art with better generalizability. Xu Li 0015, Yukai Li, Mingjie Shao, Qiuqiang Kong |
ICASSP | 3 |
| 2025 | NTC-KWS: Noise-aware CTC for Robust Keyword SpottingabstractIn recent years, there has been a growing interest in designing small-footprint yet effective Connectionist Temporal Classification based keyword spotting (CTC-KWS) systems. They are typically deployed on low-resource computing platforms, where limitations on model size and computational capacity create bottlenecks under complicated acoustic scenarios. Such constraints often result in overfitting and confusion between keywords and background noise, leading to high false alarms. To address these issues, we propose a noise-aware CTC-based KWS (NTC-KWS) framework designed to enhance model robustness in noisy environments, particularly under extremely low signal-to-noise ratios. Our approach introduces two additional noise-modeling wildcard arcs into the training and decoding processes based on weighted finite state transducer (WFST) graphs: self-loop arcs to address noise insertion errors and bypass arcs to handle masking and interference caused by excessive noise. Experiments on clean and noisy Hey Snips show that NTC-KWS outperforms state-of-the-art (SOTA) end-to-end systems and CTC-KWS baselines across various acoustic conditions, with particularly strong performance in low SNR scenarios. Yu Xi, Xu Li 0015, Wen Ding 0005, Kai Yu 0004 |
ICASSP | 5 |
| 2024 | EA-VTR: Event-Aware Video-Text Retrieval
Zongyang Ma, Ziqi Zhang 0010, Zhongang Qi, Chunfeng Yuan, Bing Li 0001, Yingmin Luo, Xu Li 0015, Xiaojuan Qi 0001, Ying Shan, Weiming Hu 0004 |
ECCV (52) | 8 |
| 2024 | Humtrans: A Novel Open-Source Dataset for Humming Melody Transcription and BeyondabstractThis paper introduces the HumTrans dataset, which is publicly available and primarily designed for humming melody transcription. The dataset can also serve as a foundation for downstream tasks such as humming melody based music generation. It consists of 500 musical compositions of different genres and languages, with each composition divided into multiple segments. In total, the dataset comprises 1000 music segments. To collect this humming dataset, we employed 10 college students, all of whom are either music majors or proficient in playing at least one musical instrument. Each of them hummed every segment twice using the web recording interface provided by our designed website1. The humming recordings were sampled at a frequency of 44,100 Hz. During the humming session, the main interface provides a musical score for students to reference, with the melody audio playing simultaneously to aid in capturing both melody and rhythm. The dataset encompasses approximately 56.22 hours of audio, making it the largest known humming dataset to date. The dataset will be released on Hugging Face2, and we will provide a GitHub repository containing baseline results and evaluation codes3. Shansong Liu, Xu Li 0015, Ying Shan |
ICASSP | 2 |
| 2024 | Neural Concatenative Singing Voice Conversion: Rethinking Concatenation-Based Approach for One-Shot Singing Voice ConversionabstractAny-to-any singing voice conversion (SVC) is confronted with the challenge of "timbre leakage" issue caused by inadequate disentanglement between the content and the speaker timbre. To address this issue, this study introduces NeuCoSVC, a novel neural concatenative SVC framework. It consists of a self-supervised learning (SSL) representation extractor, a neural harmonic signal generator, and a waveform synthesizer. The SSL extractor condenses audio into fixed-dimensional SSL features, while the harmonic signal generator leverages linear time-varying filters to produce both raw and filtered harmonic signals for pitch information. The synthesizer reconstructs waveforms using SSL features, harmonic signals, and loudness information. During inference, voice conversion is performed by substituting source SSL features with their nearest counterparts from a matching pool which comprises SSL features extracted from the reference audio, while preserving raw harmonic signals and loudness from the source audio. By directly utilizing SSL features from the reference audio, the proposed framework effectively resolves the "timbre leakage" issue caused by previous disentanglement-based approaches. Experimental results demonstrate that the proposed NeuCoSVC system outperforms the disentanglement-based speaker embedding approach in one-shot SVC across intra-language, cross-language, and cross-domain evaluations. Binzhu Sha, Xu Li 0015, Zhiyong Wu 0001, Ying Shan, Helen M. Meng |
ICASSP | 2 |
| 2024 | AFL-Net: Integrating Audio, Facial, and Lip Modalities with a Two-step Cross-attention for Robust Speaker Diarization in the Wild
Yongkang Yin, Xu Li 0015, Ying Shan, Yuexian Zou |
INTERSPEECH | 2 |
| 2024 | CLAPSep: Leveraging Contrastive Pre-Trained Model for Multi-Modal Query-Conditioned Target Sound ExtractionabstractUniversal sound separation (USS) aims to extract arbitrary types of sounds from real-world recordings. This can be achieved by language-queried target sound extraction (TSE), which typically consists of two components: a query network that converts user queries into conditional embeddings, and a separation network that extracts the target sound accordingly. Existing methods commonly train models from scratch. As a consequence, substantial data and computational resources are required to make the randomly initialized model comprehend sound events and perform separation accordingly. In this paper, we propose to integrate pre-trained models into TSE models to address the above issue. To be specific, we tailor and adapt the powerful contrastive language-audio pre-trained model (CLAP) for USS, denoted as CLAPSep. CLAPSep also accepts flexible user inputs, taking both positive and negative user prompts of uni- and/or multi-modalities for target sound extraction. These key features of CLAPSep can not only enhance the extraction performance but also improve the versatility of its application. We provide extensive experiments on 5 diverse datasets to demonstrate the superior performance and zero- and few-shot generalizability of our proposed CLAPSep with fast training convergence, surpassing previous methods by a significant margin. Full codes and some audio examples are released for reproduction and evaluation. Xu Li 0015, Mingjie Shao, Xixin Wu |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Covariance Regularization for Probabilistic Linear Discriminant AnalysisabstractProbabilistic linear discriminant analysis (PLDA) is commonly used in speaker verification systems to score the similarity of speaker embeddings. Recent studies improved the performance of PLDA in domain-matched conditions by diagonalizing its covariance. We suspect such a brutal pruning approach could eliminate its capacity in modeling dimension correlation of speaker embeddings, leading to inadequate performance with domain adaptation. This paper explores two alternative covariance regularization approaches, namely, interpolated PLDA and sparse PLDA, to tackle the problem. The interpolated PLDA incorporates the prior knowledge from cosine scoring to interpolate the covariance of PLDA. The sparse PLDA introduces a sparsity penalty to update the covariance. Experimental results demonstrate that both approaches outperform diagonal regularization noticeably with domain adaptation. In addition, in-domain data can be significantly reduced when training sparse PLDA for domain adaptation. Mingjie Shao, Xuanji He, Xu Li 0015, Tan Lee, Guanglu Wan |
ICASSP | 4 |
| 2023 | Enhancing the Vocal Range of Single-Speaker Singing Voice Synthesis with Melody-Unsupervised Pre-TrainingabstractThe single-speaker singing voice synthesis (SVS) usually underperforms at pitch values that are out of the singer's vocal range or associated with limited training samples. Based on our previous work, this work proposes a melody-unsupervised multi-speaker pretraining method conducted on a multi-singer dataset to enhance the vocal range of the single-speaker, while not degrading the timbre similarity. This pre-training method can be deployed to a large-scale multi-singer dataset, which only contains audio-and-lyrics pairs without phonemic timing information and pitch annotation. Specifically, in the pre-training step, we design a phoneme predictor to produce the frame-level phoneme probability vectors as the phonemic timing information and a speaker encoder to model the timbre variations of different singers, and directly estimate the frame-level f0 values from the audio to provide the pitch information. These pre-trained model parameters are delivered into the fine-tuning step as prior knowledge to enhance the single speaker's vocal range. Moreover, this work also contributes to improving the sound quality and rhythm naturalness of the synthesized singing voices. It is the first to introduce a differentiable duration regulator to improve the rhythm naturalness of the synthesized voice, and a bi-directional flow model to improve the sound quality. Experimental results verify that the proposed SVS system outperforms the baseline on both sound quality and naturalness. Shaohuan Zhou, Xu Li 0015, Zhiyong Wu 0001, Ying Shan, Helen M. Meng |
ICASSP | 2 |
| 2023 | Prosody Modeling with 3D Visual Information for Expressive Video Dubbing
Shansong Liu, Xu Li 0015, Haozhe Wu, Zhiyong Wu 0001, Ying Shan, Jia Jia 0001 |
INTERSPEECH | 3 |
| 2022 | Characterizing the Adversarial Vulnerability of Speech self-Supervised LearningabstractA leaderboard named Speech processing Universal PERformance Benchmark (SUPERB), which aims at benchmarking the performance of a shared self-supervised learning (SSL) speech model across various downstream speech tasks with minimal modification of architectures and a small amount of data, has fueled the research for speech representation learning. The SUPERB demonstrates speech SSL upstream models improve the performance of various downstream tasks through just minimal adaptation. As the paradigm of the self-supervised learning upstream model followed by downstream tasks arouses more attention in the speech community, characterizing the adversarial robustness of such paradigm is of high priority. In this paper, we make the first attempt to investigate the adversarial vulnerability of such paradigm under the attacks from both zero-knowledge adversaries and limited-knowledge adversaries. The experimental results illustrate that the paradigm proposed by SUPERB is seriously vulnerable to limited-knowledge adversaries, and the attacks generated by zero-knowledge adversaries are with transferability. The XAB test verifies the imperceptibility of crafted adversarial attacks. Xu Li 0015, Xixin Wu, Hung-yi Lee, Helen M. Meng |
ICASSP | 3 |
| 2022 | A Hierarchical Speaker Representation Framework for One-shot Singing Voice ConversionabstractTypically, singing voice conversion (SVC) depends on an embedding vector, extracted from either a speaker lookup table (LUT) or a speaker recognition network (SRN), to model speaker identity.However, singing contains more expressive speaker characteristics than conversational speech.It is suspected that a single embedding vector may only capture averaged and coarse-grained speaker characteristics, which is insufficient for the SVC task.To this end, this work proposes a novel hierarchical speaker representation framework for SVC, which can capture fine-grained speaker characteristics at different granularity.It consists of an up-sampling stream and three down-sampling streams.The up-sampling stream transforms the linguistic features into audio samples, while one downsampling stream of the three operates in the reverse direction.It is expected that the temporal statistics of each down-sampling block can represent speaker characteristics at different granularity, which will be engaged in the up-sampling blocks to enhance the speaker modeling.Experiment results verify that the proposed method outperforms both the LUT and SRN based SVC systems.Moreover, the proposed system supports the one-shot SVC with only a few seconds of reference audio. Xu Li 0015, Shansong Liu, Ying Shan |
INTERSPEECH | 1 |
| 2022 | Spoofing-Aware Speaker Verification by Multi-Level FusionabstractRecently, many novel techniques have been introduced to deal with spoofing attacks, and achieve promising countermeasure (CM) performances.However, these works only take the standalone CM models into account.Nowadays, a spoofing aware speaker verification (SASV) challenge which aims to facilitate the research of integrated CM and ASV models, arguing that jointly optimizing CM and ASV models will lead to better performance, is taking place.In this paper, we propose a novel multi-model and multi-level fusion strategy to tackle the SASV task.Compared with purely scoring fusion and embedding fusion methods, this framework first utilizes embeddings from CM models, propagating CM embeddings into a CM block to obtain a CM score.In the second-level fusion, the CM score and ASV scores directly from ASV systems will be concatenated into a prediction block for the final decision.As a result, the best single fusion system has achieved the SASV-EER of 0.97% on the evaluation set.Then by ensembling the top-5 fusion systems, the final SASV-EER reached 0.89%. Lingwei Meng, Jiawen Kang 0002, Jinchao Li, Xu Li 0015, Xixin Wu, Hung-yi Lee, Helen M. Meng |
INTERSPEECH | 5 |
| 2022 | Improving the Adversarial Robustness for Speaker Verification by Self-Supervised LearningabstractPrevious works have shown that automatic speaker verification (ASV) is seriously vulnerable to malicious spoofing attacks, such as replay, synthetic speech, and recently emerged adversarial attacks. Great efforts have been dedicated to defending ASV against replay and synthetic speech; however, only a few approaches have been explored to deal with adversarial attacks. All the existing approaches to tackle adversarial attacks for ASV require the knowledge for adversarial samples generation, but it is impractical for defenders to know the exact attack algorithms that are applied by the in-the-wild attackers. This work is among the first to perform adversarial defense for ASV without knowing the specific attack algorithms. Inspired by self-supervised learning models (SSLMs) that possess the merits of alleviating the superficial noise in the inputs and reconstructing clean samples from the interrupted ones, this work regards adversarial perturbations as one kind of noise and conducts adversarial defense for ASV by SSLMs. Specifically, we propose to perform adversarial defense from two perspectives: 1) adversarial perturbation purification and 2) adversarial perturbation detection. The purification module aims at alleviating the adversarial perturbations in the samples and pulling the contaminated adversarial inputs back towards the decision boundary. Experimental results show that our proposed purification module effectively counters adversarial attacks and outperforms traditional filters from both alleviating the adversarial noise and maintaining the performance of genuine samples. The detection module aims at detecting adversarial samples from genuine ones based on the statistical properties of ASV scores derived by a unique ASV integrating with different number of SSLMs. Experimental results show that our detection module helps shield the ASV by detecting adversarial samples. Both purification and detection methods are helpful for defending against different kinds of attack algorithms. Moreover, since there is no common metric for evaluating the ASV performance under adversarial attacks, this work also formalizes evaluation metrics for adversarial defense considering both purification and detection based approaches into account. We sincerely encourage future works to benchmark their approaches based on the proposed evaluation framework. Xu Li 0015, Andy T. Liu, Zhiyong Wu 0001, Helen M. Meng, Hung-yi Lee |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Replay and Synthetic Speech Detection with Res2Net ArchitectureabstractExisting approaches for replay and synthetic speech detection still lack generalizability to unseen spoofing attacks. This work proposes to leverage a novel model structure, so-called Res2Net, to improve the anti-spoofing countermeasure’s generalizability. Res2Net mainly modifies the ResNet block to enable multiple feature scales. Specifically, it splits the feature maps within one block into multiple channel groups and designs a residual-like connection across different channel groups. Such connection increases the possible receptive fields, resulting in multiple feature scales. This multiple scaling mechanism significantly improves the countermeasure’s generalizability to unseen spoofing attacks. It also decreases the model size compared to ResNet-based models. Experimental results show that the Res2Net model consistently outperforms ResNet34 and ResNet50 by a large margin in both physical access (PA) and logical access (LA) of the ASVspoof 2019 corpus. Moreover, integration with the squeeze-and-excitation (SE) block can further enhance performance. For feature engineering, we investigate the gen-eralizability of Res2Net combined with different acoustic features, and observe that the constant-Q transform (CQT) achieves the most promising performance in both PA and LA scenarios. Our best single system outperforms other state-of-the-art single systems in both PA and LA of the ASVspoof 2019 corpus. Xu Li 0015, Na Li 0012, Chao Weng, Xunying Liu, Dan Su 0002, Dong Yu 0001, Helen M. Meng |
ICASSP | 1 |
| 2021 | Adversarial Defense for Automatic Speaker Verification by Cascaded Self-Supervised Learning ModelsabstractAutomatic speaker verification (ASV) is one of the core technologies in biometric identification. With the ubiquitous usage of ASV systems in safety-critical applications, more and more malicious attackers attempt to launch adversarial attacks at ASV systems. In the midst of the arms race between attack and defense in ASV, how to effectively improve the robustness of ASV against adversarial attacks remains an open question. We note that the self-supervised learning models possess the ability to mitigate superficial perturbations in the input after pretraining. Hence, with the goal of effective defense in ASV against adversarial attacks, we propose a standard and attack-agnostic method based on cascaded self-supervised learning models to purify the adversarial perturbations. Experimental results demonstrate that the proposed method achieves effective defense performance and can successfully counter adversarial attacks in scenarios where attackers may either be aware or unaware of the self-supervised learning models. Xu Li 0015, Andy T. Liu, Zhiyong Wu 0001, Helen M. Meng, Hung-yi Lee |
ICASSP | 2 |
| 2021 | Channel-Wise Gated Res2Net: Towards Robust Detection of Synthetic Speech AttacksabstractExisting approaches for anti-spoofing in automatic speaker verification (ASV) still lack generalizability to unseen attacks.The Res2Net approach designs a residual-like connection between feature groups within one block, which increases the possible receptive fields and improves the system's detection generalizability.However, such a residual-like connection is performed by a direct addition between feature groups without channelwise priority.We argue that the information across channels may not contribute to spoofing cues equally, and the less relevant channels are expected to be suppressed before adding onto the next feature group, so that the system can generalize better to unseen attacks.This argument motivates the current work that presents a novel, channel-wise gated Res2Net (CG-Res2Net), which modifies Res2Net to enable a channel-wise gating mechanism in the connection between feature groups.This gating mechanism dynamically selects channel-wise features based on the input, to suppress the less relevant channels and enhance the detection generalizability.Three gating mechanisms with different structures are proposed and integrated into Res2Net.Experimental results conducted on ASVspoof 2019 logical access (LA) demonstrate that the proposed CG-Res2Net significantly outperforms Res2Net on both the overall LA evaluation set and individual difficult unseen attacks, which also outperforms other state-of-the-art single systems, depicting the effectiveness of our method. Xu Li 0015, Xixin Wu, Xunying Liu, Helen M. Meng |
Interspeech | 1 |
| 2021 | VAENAR-TTS: Variational Auto-Encoder Based Non-AutoRegressive Text-to-Speech SynthesisabstractThis paper describes a variational auto-encoder based nonautoregressive text-to-speech (VAENAR-TTS) model.The autoregressive TTS (AR-TTS) models based on the sequenceto-sequence architecture can generate high-quality speech, but their sequential decoding process can be time-consuming.Recently, non-autoregressive TTS (NAR-TTS) models have been shown to be more efficient with the parallel decoding process.However, these NAR-TTS models rely on phoneme-level durations to generate a hard alignment between the text and the spectrogram.Obtaining duration labels, either through forced alignment or knowledge distillation, is cumbersome.Furthermore, hard alignment based on phoneme expansion can degrade the naturalness of the synthesized speech.In contrast, the proposed model of VAENAR-TTS is an end-to-end approach that does not require phoneme-level durations.The VAENAR-TTS model does not contain recurrent structures and is completely non-autoregressive in both the training and inference phases.Based on the VAE architecture, the alignment information is encoded in the latent variable, and attention-based soft alignment between the text and the latent variable is used in the decoder to reconstruct the spectrogram.Experiments show that VAENAR-TTS achieves state-of-the-art synthesis quality, while the synthesis speed is comparable with other NAR-TTS models. Zhiyong Wu 0001, Xixin Wu, Xu Li 0015, Shiyin Kang, Xunying Liu, Helen M. Meng |
Interspeech | 4 |
| 2021 | Pairing Weak with Strong: Twin Models for Defending Against Adversarial Attack on Speaker Verification
Xu Li 0015, Tan Lee |
Interspeech | 2 |
| 2020 | Adversarial Attacks on GMM I-Vector Based Speaker Verification SystemsabstractThis work investigates the vulnerability of Gaussian Mixture Model (GMM) i-vector based speaker verification systems to adversarial attacks, and the transferability of adversarial samples crafted from GMM i-vector based systems to x-vector based systems. In detail, we formulate the GMM i-vector system as a scoring function of enrollment and testing utterance pairs. Then we leverage the fast gradient sign method (FGSM) to optimize testing utterances for adversarial samples generation. These adversarial samples are used to attack both GMM i-vector and x-vector systems. We measure the system vulnerability by the degradation of equal error rate and false acceptance rate. Experiment results show that GMM i-vector systems are seriously vulnerable to adversarial attacks, and the crafted adversarial samples are proved to be transferable and pose threats to neural network speaker embedding based systems (e.g. x-vector systems). Xu Li 0015, Jinghua Zhong, Xixin Wu, Jianwei Yu 0001, Xunying Liu, Helen M. Meng |
ICASSP | 1 |
| 2020 | Investigating Robustness of Adversarial Samples Detection for Automatic Speaker VerificationabstractRecently adversarial attacks on automatic speaker verification (ASV) systems attracted widespread attention as they pose severe threats to ASV systems.However, methods to defend against such attacks are limited.Existing approaches mainly focus on retraining ASV systems with adversarial data augmentation.Also, countermeasure robustness against different attack settings are insufficiently investigated.Orthogonal to prior approaches, this work proposes to defend ASV systems against adversarial attacks with a separate detection network, rather than augmenting adversarial data into ASV training.A VGG-like binary classification detector is introduced and demonstrated to be effective on detecting adversarial samples.To investigate detector robustness in a realistic defense scenario where unseen attack settings may exist, we analyze various kinds of unseen attack settings' impact and observe that the detector is robust (6.27%EER det degradation in the worst case) against unseen substitute ASV systems, but it has weak robustness (50.37%EER det degradation in the worst case) against unseen perturbation methods.The weak robustness against unseen perturbation methods shows a direction for developing stronger countermeasures. Xu Li 0015, Na Li 0012, Jinghua Zhong, Xixin Wu, Xunying Liu, Dan Su 0002, Dong Yu 0001, Helen M. Meng |
INTERSPEECH | 1 |
| 2019 | End-to-end Code-switched TTS with Mix of Monolingual RecordingsabstractState-of-the-art text-to-speech (TTS) synthesis models can produce monolingual speech with high intelligibility and naturalness. However, when the models are applied to synthesize code-switched (CS) speech, the performance declines seriously. Conventionally, developing a CS TTS system requires multilingual data to incorporate language-specific and cross-lingual knowledge. Recently, end-to-end (E2E) architecture has achieved satisfactory results in monolingual TTS. The architecture enables the training from one end of alphabetic text input to the other end of acoustic feature output. In this paper, we explore the use of E2E framework for CS TTS, using a combination of Mandarin and English monolingual speech corpus uttered by two female speakers. To handle alphabetic input from different languages, we explore two kinds of encoders: (1) shared multilingual encoder with explicit language embedding (LDE); (2) separated monolingual encoder (SPE) for each language. The two systems use identical decoder architecture, where a discriminative code is incorporated to enable the model to generate speech in one speaker's voice consistently. Experiments confirm the effectiveness of the proposed modifications on the E2E TTS framework in terms of quality and speaker similarity of the generated speech. Moreover, our proposed systems can generate controllable foreign-accented speech at character-level using only mixture of monolingual training data. Yuewen Cao, Xixin Wu, Songxiang Liu, Jianwei Yu 0001, Xu Li 0015, Zhiyong Wu 0001, Xunying Liu, Helen M. Meng |
ICASSP | 5 |
| 2019 | Speech Emotion Recognition Using Capsule NetworksabstractSpeech emotion recognition (SER) is a fundamental step towards fluent human-machine interaction. One challenging problem in SER is obtaining utterance-level feature representation for classification. Recent works on SER have made significant progress by using spectrogram features and introducing neural network methods, e.g., convolutional neural networks (CNNs). However the fundamental problem of CNNs is that the spatial information in spectrograms is not captured, which are basically position and relationship information of low-level features like pitch and formant frequencies. This paper presents a novel architecture based on the capsule networks (CapsNets) for SER. The proposed system can take into account the spatial relationship of speech features in spectrograms, and provide an effective pooling method for obtaining utterance global features. We also introduce a recurrent connection to CapsNets to improve the model's time sensitivity. We compare the proposed model to previous published results based on combined CNN-long short-term memory (CNN-LSTM) models on the benchmark corpus IEMOCAP over four emotions, i.e., neutral, angry, happy and sad. Experimental results show that our model achieves better results than the baseline system on weighted accuracy (WA) (72.73% vs. 68.8%) and un-weighted accuracy (UA) (59.71% vs. 59.4%), which demonstrates the effectiveness of CapsNets for SER. Xixin Wu, Songxiang Liu, Yuewen Cao, Xu Li 0015, Jianwei Yu 0001, Dongyang Dai, Xi Ma, Shoukang Hu, Zhiyong Wu 0001, Xunying Liu, Helen M. Meng |
ICASSP | 4 |
| 2019 | Comparative Study of Parametric and Representation Uncertainty Modeling for Recurrent Neural Network Language Models
Jianwei Yu 0001, Max W. Y. Lam, Shoukang Hu, Xixin Wu, Xu Li 0015, Yuewen Cao, Xunying Liu, Helen M. Meng |
INTERSPEECH | 5 |
| 2018 | Unsupervised Discovery of an Extended Phoneme Set in L2 English Speech for Mispronunciation Detection and DiagnosisabstractSecond language (L2) speech is often labelled with the native, phoneme categories. Hence, we often observe segments for which it is difficult, if not impossible, to decide on a categorical phoneme label. We refer to these segments as “non-categorical” phoneme units. Existing approaches to mispronunciation detection and diagnosis (MDD) mostly focus on categorical phoneme errors, where one native phoneme is substituted for another. However, noncategorical errors are not considered. To better represent L2 speech for improved MDD, this work aims to discover an Extended Phoneme Set in L2 speech (L2-EPS) which includes not only the categorical phonemes based on the native set, but also non-categorical phoneme units. We apply an optimized k-means algorithm to cluster phoneme-based phonemic posterior-grams (PPGs), which are generated through an acoustic-phonemic model (APM). Then we find the L2-EPS based on analysis of the clusters obtained. We verified experimentally that the non-categorical phonemes in L2-EPS can extend the native phoneme categories to better describe L2 speech. Hence L2-EPS can enrich the existing approaches to MDD for better performance. Shaoguang Mao, Xu Li 0015, Kun Li 0003, Zhiyong Wu 0001, Xunying Liu, Helen M. Meng |
ICASSP | 2 |
| 2018 | Applying Multitask Learning to Acoustic-Phonemic Model for Mispronunciation Detection and Diagnosis in L2 English SpeechabstractFor mispronunciation detection and diagnosis (MDD), nowadays approaches generally treat the phonemes in correct and mispronunciations as the same despite the fact they may actually carry different characteristics. Furthermore, serious data imbalance issue between correct and mispronunciation in dataset further influences the performances. To address these problems, this paper investigates the use of multi-task (MT) learning technique to enhance the acoustic-phonemic model (APM) for MDD. The phonemes in correct and mispronunciations are processed separately but in multi-task manner considering both correct and mispronunciation recognition tasks. A feature representation module is further proposed to improve performance. Compared with baseline APM, the proposed MT-APM, R-MT-APM achieve better performance not only in Precision, Recall and F-Measure, but also in mispronunciation detection and diagnosis accuracies. With feature representation module, R-MT-APM achieves the highest mispronunciation detection accuracy. Shaoguang Mao, Zhiyong Wu 0001, Runnan Li, Xu Li 0015, Helen M. Meng, Lianhong Cai |
ICASSP | 4 |
| 2018 | Integrating Articulatory Features into Acoustic-Phonemic Model for Mispronunciation Detection and Diagnosis in L2 English SpeechabstractThis paper proposes novel approaches to mispronunciation detection and diagnosis (MDD) on second-language (L2) learners' speech with articulatory features. Here, articulatory features are the positions of articulators when pronouncing phonemes and reflect the pronunciation mechanisms of each phoneme. The use of articulatory features in MDD is helpful in distinguishing phonemes. Three models with articulatory features are proposed based on acoustic-phonemic model (APM): 1) articulatory-acoustic-phonemic model (AAPM) that embeds articulatory features directly into input features; 2) AAPM with feature representation (R-AAPM) to represent original input features with articulatory features; and 3) articulatory multi-task acoustic-phonemic model (A-MT-APM) where phoneme recognizer and articulatory feature classifiers are trained simultaneously in multi-task manner. Compared with baseline phoneme-based APM, proposed approaches perform better in mispronunciation detection and diagnosis measured with Precision, Recall and F1-Measure metrics. Specifically, the A-MT-APM approach gains 5.6% and 7.0% improvement in F1-Measure and diagnostic accuracy respectively. The contributions include: 1) introducing the articulatory features to MDD in deep learning framework; 2) investigating several model architectures for better exploiting articulatory features. Shaoguang Mao, Zhiyong Wu 0001, Xu Li 0015, Runnan Li, Xixin Wu, Helen M. Meng |
ICME | 3 |
| 2018 | Unsupervised Discovery of Non-native Phonetic Patterns in L2 English Speech for Mispronunciation Detection and Diagnosis
Xu Li 0015, Shaoguang Mao, Xixin Wu, Kun Li 0003, Xunying Liu, Helen M. Meng |
INTERSPEECH | 1 |
| 2018 | Automatic lexical stress and pitch accent detection for L2 English speech using multi-distribution deep neural networks
Kun Li 0003, Shaoguang Mao, Xu Li 0015, Zhiyong Wu 0001, Helen M. Meng |
Speech Commun. | 3 |
| 2016 | Low level descriptors based DBLSTM bottleneck feature for speech driven talking avatarabstractSpeech is bimodal in nature. There are close correlations between the acoustic speech signals and the visual gestures such as lip movements, facial expressions and head motions. For speech driven talking avatar, how to derive more representative acoustic features from which to predict more accurate and realistic visual gestures still remains the research problem. Inspired by the promising performance of low level descriptors (LLD) in speech emotion recognition, in this work, we investigate the usage of LLD feature for the task of speech driven talking avatar. Furthermore, visual gestures also demonstrate correlations with not only context information of past or future acoustic features (e.g. anticipatory co-articulation phenomena) but also textual information (e.g. textual hints for lip movement). To incorporate such information, we also propose to use deep bidirectional long short-term memory (DBLSTM) as the bottleneck feature extractor, which can combine LLD feature with contextual information. Experimental results indicate that the proposed LLD based DBLSTM bottleneck feature outperforms the conventional spectrum related features for the task of speech driven talking avatar, and more sophisticated contextual information can further improve the performance. Xinyu Lan, Xu Li 0015, Yishuang Ning, Zhiyong Wu 0001, Helen M. Meng, Jia Jia 0001, Lianhong Cai |
ICASSP | 2 |
| 2016 | Phoneme Embedding and its Application to Speech Driven Talking Avatar Synthesis
Xu Li 0015, Zhiyong Wu 0001, Helen M. Meng, Jia Jia 0001, Xiaoyan Lou, Lianhong Cai |
INTERSPEECH | 1 |
| 2016 | Expressive Speech Driven Talking Avatar Synthesis with DBLSTM Using Limited Amount of Emotional Bimodal Data
Xu Li 0015, Zhiyong Wu 0001, Helen M. Meng, Jia Jia 0001, Xiaoyan Lou, Lianhong Cai |
INTERSPEECH | 1 |