VLDB 2026 Research / reviewers in the wild / expert
Sheng Li 0010
dblp:23/3439-10
· DBLP profile ↗
80ranked-venue papers
16as first author
51since 2021 · last 2025
0000-0001-7636-3797ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 61 · 12 first-author · 36 since 2021Artificial intelligence and machine learning · 46 · 13 first-author · 29 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language ModelsabstractZhen Wan, Chao-Han Huck Yang, Yahan Yu, Jinchuan Tian, Sheng Li, Ke Hu, Zhehuai Chen, Shinji Watanabe, Fei Cheng, Chenhui Chu, Sadao Kurohashi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Chao-Han Huck Yang, Yahan Yu, Jinchuan Tian, Sheng Li 0010, Zhehuai Chen, Shinji Watanabe 0001, Fei Cheng 0002, Chenhui Chu, Sadao Kurohashi |
ACL (1) | 5 |
| 2025 | CoVoGER: A Multilingual Multitask Benchmark for Speech-to-text Generative Error Correction with Large Language ModelsabstractLarge language models (LLMs) can rewrite the N -best hypotheses from a speech-to-text model, often fixing recognition or translation errors that traditional rescoring cannot.Yet research on generative error correction (GER) has been focusing on monolingual automatic speech recognition (ASR), leaving its multilingual and multitask potential underexplored.We introduce CoVoGER, a benchmark for GER that covers both ASR and speech-to-text translation (ST) across 15 languages and 28 language pairs.CoVoGER is constructed by decoding Common Voice 20.0 and CoVoST-2 with Whisper of three model sizes and Seam-lessM4T of two model sizes, providing 5-best lists obtained via a mixture of beam search and temperature sampling.We evaluated various instruction-tuned LLMs, including commercial models in zero-shot mode and open-sourced models with LoRA fine-tuning, and found that the mixture decoding strategy yields the best GER performance in most settings.CoVoGER will be released to promote research on reliable language-universal speech-to-text GER. Zhengdong Yang, Sheng Li 0010, Chao-Han Huck Yang, Chenhui Chu |
EMNLP | 3 |
| 2025 | Joint Automatic Speech Recognition And Structure Learning For Better Speech UnderstandingabstractSpoken language understanding (SLU) is a structure prediction task in the field of speech. Recently, many works on SLU that treat it as a sequence-to-sequence task have achieved great success. However, This method is not suitable for simultaneous speech recognition and understanding. In this paper, we propose a joint speech recognition and structure learning framework (JSRSL), an end-to-end SLU model based on span, which can accurately transcribe speech and extract structured content simultaneously. We conduct experiments on name entity recognition and intent classification using the Chinese dataset AISHELL-NER and the English dataset SLURP. The results show that our proposed method not only outperforms the traditional sequence-to-sequence method in both transcription and extraction capabilities but also achieves state-of-the-art performance on the two datasets. Jiliang Hu 0001, Zuchao Li, Mengjia Shen, Haojun Ai, Sheng Li 0010 |
ICASSP | 5 |
| 2025 | Similarity-based Accent Recognition with Continuous and Discrete Self-supervised Speech RepresentationsabstractThe primary challenge in accent recognition lies in data scarcity due to the high diversity of accents, which make the collection of large-scale training data for each accent almost impossible in practice. To overcome this challenge, we propose a simple solution that leverages both continuous and discrete feature representations from pretrained speech self-supervised learning (SSL) models. Our model is simplified to a linear projection layer and a set of trainable accent class embeddings. Cosine similarity between the accent embeddings and the latent features of an audio sample is used to predict its accent class. This approach enables the model to access features that contain rich accent-related information while reducing the risk of model overfitting. Our method provides a practical and efficient way to tackle accent recognition, especially in low-resource scenarios. Experimental results on English accent recognition show that our best model achieves an accuracy of 84.0% on the AESRC 2020 dataset and an Unweighted Average Recall (UAR) of 50.0% on the VCTK corpus, setting new state-of-the-art results on both datasets. Jun-You Wang, Sheng Li 0010, Li-An Lu, Sydney Chia-Chun Kao, Jyh-Shing Roger Jang |
ICASSP | 2 |
| 2025 | Extending Whisper for Emotion Prediction Using Word-level Pseudo LabelsabstractThis paper extends Whisper’s automatic speech recognition (ASR) capabilities to perform speech-based emotion recognition (SER) by incorporating word-level emotion classification alongside ASR output. We generate four emotion pseudo-labels (neutral, happy, sad, angry) for each word using a pretrained frame-level SER model, and Whisper is fine-tuned for joint ASR and emotion classification at the word level. Sentence-level emotion labels are masked during training to encourage the transformer to use the ASR output for word-level emotion prediction. During inference, word-level predictions are combined with sentence-level predictions through majority voting to generate the final sentence-level label. When evaluated on the IEMOCAP dataset, our method maintains Whisper’s ASR word error rate while improving the SER weighted accuracy from 74.4% to 76.4% and the unweighted average recall from 77.1% to 79.0%. Kwok Chin Yuen, Sheng Li 0010, Jia Qi Yip, Chenhui Chu, Tatsuya Kawahara, Chng Eng Siong |
ICASSP | 2 |
| 2025 | Language-Aware Prompt Tuning for Parameter-Efficient Seamless Language Expansion in Multilingual ASRabstractRecent advancements in multilingual automatic speech recognition (ASR) have been driven by large-scale end-to-end models like Whisper. However, challenges such as language interference and expanding to unseen languages (language expansion) without degrading performance persist. This paper addresses these with three contributions: 1) Entire Soft Prompt Tuning (Entire SPT), which applies soft prompts to both the encoder and decoder, enhancing feature extraction and decoding; 2) Language-Aware Prompt Tuning (LAPT), which leverages cross-lingual similarities to encode shared and language-specific features using lightweight prompt matrices; 3) SPT-Whisper, a toolkit that integrates SPT into Whisper and enables efficient continual learning. Experiments across three languages from FLEURS demonstrate that Entire SPT and LAPT outperform Decoder SPT by 5.0% and 16.0% in language expansion tasks, respectively, providing an efficient solution for dynamic, multilingual ASR models with minimal computational overhead. Sheng Li 0010, Hao Huang 0009, Ayiduosi Tuohan, Yizhou Peng |
INTERSPEECH | 2 |
| 2025 | Adapting Whisper for Parameter-efficient Code-Switching Speech Recognition via Soft Prompt TuningabstractLarge-scale multilingual ASR models like Whisper excel in high-resource settings but face challenges in low-resource scenarios, such as rare languages and code-switching (CS), due to computational costs and catastrophic forgetting. We explore Soft Prompt Tuning (SPT), a parameter-efficient method to enhance CS ASR while preserving prior knowledge. We evaluate two strategies: (1) full fine-tuning (FFT) of both soft prompts and the entire Whisper model, demonstrating improved cross-lingual capabilities compared to traditional methods, and (2) adhering to SPT's original design by freezing model parameters and only training soft prompts. Additionally, we introduce SPT4ASR, a combination of different SPT variants. Experiments on the SEAME and ASRU2019 datasets show that deep prompt tuning is the most effective SPT approach, and our SPT4ASR methods achieve further error reductions in CS ASR, maintaining parameter efficiency similar to LoRA, without degrading performance on existing languages. Yizhou Peng, Hao Huang 0009, Sheng Li 0010 |
INTERSPEECH | 4 |
| 2025 | Simple and Effective Content Encoder for Singing Voice Conversion via SSL-Embedding Dimension Reduction
Wangjin Zhou, Tianjiao Du, Chenglin Xu, Sheng Li 0010, Yi Zhao 0006, Tatsuya Kawahara |
INTERSPEECH | 4 |
| 2025 | Towards Emotion Co-regulation with LLM-powered Socially Assistive Robots: Integrating LLM Prompts and Robotic Behaviors to Support Parent-Neurodivergent Child DyadsabstractSocially Assistive Robotics (SAR) has shown promise in supporting emotion regulation for neurodivergent children. Recently, there has been increasing interest in leveraging advanced technologies to assist parents in co-regulating emotions with their children. However, limited research has explored the integration of large language models (LLMs) with SAR to facilitate emotion co-regulation between parents and children with neurodevelopmental disorders. To address this gap, we developed an LLM-powered social robot by deploying a speech communication module on the MiRo-E robotic platform. This supervised autonomous system integrates LLM prompts and robotic behaviors to deliver tailored interventions for both parents and neurodivergent children. Pilot tests were conducted with two parent-child dyads, followed by a qualitative analysis. The findings reveal MiRo-E’s positive impacts on interaction dynamics and its potential to facilitate emotion regulation, along with identified design and technical challenges. Based on these insights, we provide design implications to advance the future development of LLM-powered SAR for mental health applications. Jing Li 0133, Felix Schijve, Sheng Li 0010, Yuye Yang, Jun Hu 0001, Emilia I. Barakova |
IROS | 3 |
| 2025 | Empowering Māori Automatic Speech Recognition through EMD-Based Augmentation
Chengxi Lei, Sheng Li 0010, Satwinder Singh, Feng Hou, Huia Jahnke, Ruili Wang 0001 |
PRICAI | 2 |
| 2025 | LatentSpeech: Latent Diffusion for Text-To-Speech GenerationabstractText-To-Speech (TTS) generation plays a crucial role in human-robot interaction by allowing robots to communicate naturally with humans. Researchers have developed various TTS models to enhance speech generation. More recently, diffusion models have emerged as a powerful generative framework, achieving state-of-the-art performance in tasks such as image and video generation. However, their application in TTS has been limited by its slow inference speeds due to their iterative denoising process. Previous work has applied diffusion models to Mel-Spectrograms with an additional vocoder to convert them into waveforms. To address these limitations, we propose LatentSpeech, a novel diffusion-based TTS framework that operates directly in a latent space. This space is significantly more compact and information-rich than raw Mel-Spectrograms. Furthermore, we introduce an alternative latent space of Pseudo-Quadrature Mirror Filters (PQMF), which decomposes speech into multiple subbands. By leveraging PQMF’s near-perfect waveform reconstruction capability, LatentSpeech eliminates the need for a separate vocoder and reduces both model size and inference time. Our PQMF-based LatentSpeech model reduces inference time by 45% and model size by 77% compared to Mel-Spectrogram diffusion models. On benchmark datasets, it achieves 25% lower WER and 58% higher MOS using the same training data. These results highlight LatentSpeech as an efficient, high-quality TTS solution for real-time and human-robot interaction. Code and models are available here. Haowei Lou, Hye-Young Paik, Pari Delir Haghighi, Sheng Li 0010, Wen Hu 0001, Lina Yao 0001 |
RO-MAN | 4 |
| 2024 | Enhancing Privacy of Spatiotemporal Federated Learning Against Gradient Inversion Attacks
Lele Zheng, Yang Cao 0011, Renhe Jiang, Kenjiro Taura, Yulong Shen 0001, Sheng Li 0010, Masatoshi Yoshikawa |
DASFAA (1) | 6 |
| 2024 | Enhancing Realism in 3D Facial Animation Using Conformer-Based Generation and Automated Post-ProcessingabstractRecent progress has propelled the development of realistic talking-face videos for avatars. Yet, animating 3D cartoon avatars remains intricate due to the imprecise nature of facial-driven data. This often manifests as inconsistent mouth configurations and rigid facial expressions, curbing the animation’s realism. Addressing these issues, we introduce a conformer-based framework that derives expression coefficients directly from phonemes, thereby elevating prediction precision and minimizing manual oversight. Furthermore, by harnessing a pre-trained emotion blending module coupled with the keyframe of the target emotional character, we employ a zero-shot adaptation technique. This serves to amplify emotional expressions and bolster the authenticity of lip dynamics. Our methodology adeptly registers nuanced expression shifts in avatars, leading to remarkably lifelike animations, as substantiated by our experimental findings. Yi Zhao 0006, Chunyu Qiang, Hao Li 0078, Yulan Hu, Wangjin Zhou, Sheng Li 0010 |
ICASSP | 6 |
| 2024 | MOS-FAD: Improving Fake Audio Detection Via Automatic Mean Opinion Score PredictionabstractIEEE Automatic Mean Opinion Score (MOS) prediction is employed to evaluate the quality of synthetic speech. This study extends the application of predicted MOS to the task of Fake Audio Detection (FAD) as we expect that MOS can be used to assess how close synthesized speech is to the natural human voice. We propose MOS-FAD, where MOS can be leveraged at two key points in FAD: training data selection and model fusion. In training data selection, we demonstrate that MOS enables effective filtering of samples from unbalanced datasets. In the model fusion, our results demonstrate that incorporating MOS as a gating mechanism in FAD model fusion enhances overall performance. Wangjin Zhou, Zhengdong Yang, Chenhui Chu, Sheng Li 0010, Raj Dabre, Yi Zhao 0006, Tatsuya Kawahara |
ICASSP | 4 |
| 2024 | Investigating ASR Error Correction with Large Language Model and Multilingual 1-best Hypotheses
Sheng Li 0010, Chen Chen 0075, Kwok Chin Yuen, Chenhui Chu, Chng Eng Siong, Hisashi Kawai |
INTERSPEECH | 1 |
| 2024 | Reproducibility Companion Paper: Stable Diffusion for Content-Style Disentanglement in Art AnalysisabstractIn this companion paper, we provide the artifacts of the GOYA model for disentangling content and style in art paintings, as presented at ICMR2023. The scripts are written in Python. Yankun Wu, Yuta Nakashima, Noa Garcia, Sheng Li 0010, Zhaoyang Zeng |
ICMR | 4 |
| 2024 | Investigating Effective Speaker Property Privacy Protection in Federated Learning for Speech Emotion Recognition
Sheng Li 0010, Yang Cao 0011, Zhao Ren, Tanja Schultz |
MMAsia | 2 |
| 2024 | Robust voice activity detection using an auditory-inspired masked modulation encoder based convolutional attention network
Longbiao Wang, Meng Ge, Masashi Unoki, Sheng Li 0010, Jianwu Dang 0001 |
Speech Commun. | 5 |
| 2023 | LE-SSL-MOS: Self-Supervised Learning MOS Prediction with Listener EnhancementabstractRecently, researchers have shown an increasing interest in automatically predicting the subjective evaluation for speech synthesis systems. This prediction is a challenging task, especially on the out-of-domain test set. In this paper, we proposed a novel fusion model for MOS prediction that combines both supervised and unsupervised approaches. In the supervised aspect, we developed a SSL-based predictor called LE-SSL-MOS. The LE-SSL-MOS utilizes pre-trained self-supervised learning models and further improves prediction accuracy by utilizing the opinion scores of each utterance in the listener enhancement branch. In the unsupervised aspect, two steps are contained: one is that we fine-tuned unit language model (ULM) using highly-intelligible domain data to improve the correlation of an unsupervised metric SpeechLMScore. Another is that we utilized ASR confidence as a new metric with the help of ensemble learning. To the best of our knowledge, this is the first architecture that fuses supervised and unsupervised methods for MOS prediction.With these approaches, our experimental results on the VoiceMOS Challenge 2023 show that LE-SSL-MOS performs better than the baseline. Our fusion system achieved an absolute improvement of 13 % over LE-SSL-MOS on the noisy and enhanced speech track. And our system ranked 1st and 2 nd respectively in the French speech synthesis track and the noisy and enhanced speech track of the challenge. Zili Qi, Xinhui Hu, Wangjin Zhou, Sheng Li 0010, Xinkang Xu |
ASRU | 4 |
| 2023 | FedCPC: An Effective Federated Contrastive Learning Method for Privacy Preserving Early-Stage Alzheimers Speech DetectionabstractThe early-stage Alzheimer’s disease (AD) detection has been considered an important field of medical studies. Like traditional machine learning methods, speech-based automatic detection also suffers from data privacy risks because the data of specific patients are exclusive to each medical institution. A common practice is to use federated learning to protect the patients’ data privacy. However, its distributed learning process also causes performance reduction. To alleviate this problem while protecting user privacy, we propose a federated contrastive pre-training (FedCPC) performed before federated training for AD speech detection, which can learn a better representation from raw data and enables different clients to share data in the pre-training and training stages. Experimental results demonstrate that the proposed methods can achieve satisfactory performance while preserving data privacy. Wenqing Wei, Zhengdong Yang, Yuan Gao 0040, Jiyi Li, Chenhui Chu, Shogo Okada, Sheng Li 0010 |
ASRU | 7 |
| 2023 | Correction while Recognition: Combining Pretrained Language Model for Taiwan-Accented Speech Recognition
Sheng Li 0010, Jiyi Li |
ICANN (7) | 1 |
| 2023 | Domain and Language Adaptation Using Heterogeneous Datasets for Wav2vec2.0-Based Speech Recognition of Low-Resource LanguageabstractWe address the effective finetuning of a large-scale pretrained model for automatic speech recognition (ASR) of lowresource languages with only a one-hour matched dataset. The finetuning is composed of domain adaptation and language adaptation, and they are conducted by using heterogeneous datasets, which are matched with either domain or language. For effective adaptation, we incorporate auxiliary tasks of domain identification and language identification with multi-task learning. Moreover, the embedding result of the auxiliary tasks is fused to the encoder output of the pretrained model for ASR. Experimental evaluations on the Khmer ASR using the corpus of ECCC (the Extraordinary Chambers in the Courts of Cambodia) demonstrate that first conducting domain adaption and then language adaption is effective. In addition, multi-tasking with domain identification and fusing the domain ID embedding gives the best performance, which is a CER improvement of 6.47% absolute from the baseline finetuning method. Soky Kak, Sheng Li 0010, Chenhui Chu, Tatsuya Kawahara |
ICASSP | 2 |
| 2023 | Hierarchical Softmax for End-To-End Low-Resource Multilingual Speech RecognitionabstractLow-resource speech recognition has been long-suffering from insufficient training data. In this paper, we propose an approach that leverages neighboring languages to improve low-resource scenario performance, founded on the hypothesis that similar linguistic units in neighboring languages exhibit comparable term frequency distributions, which enables us to construct a Huffman tree for performing multilingual hierarchical Softmax decoding. This hierarchical structure enables cross-lingual knowledge sharing among similar tokens, thereby enhancing low-resource training outcomes. Empirical analyses demonstrate that our method is effective in improving the accuracy and efficiency of low-resource speech recognition. Qianying Liu, Zhuo Gong, Zhengdong Yang, Sheng Li 0010, Chenchen Ding, Nobuaki Minematsu, Hao Huang 0009, Fei Cheng 0002, Chenhui Chu, Sadao Kurohashi |
ICASSP | 5 |
| 2023 | General or Specific? Investigating Effective Privacy Protection in Federated Learning for Speech Emotion RecognitionabstractFederated Learning (FL) is considered a new paradigm of privacy-preserving machine learning since the server trains a machine learning model in a distributed way without collecting clients’ raw data but only local models. However, recent studies show that FL suffers inference attacks. Sensitive information can still be inferred from the shared local models. In this work, we investigate the effectiveness of existing rigorous privacy-enhancing techniques, i.e., user-level differential privacy (UDP) and Voice-Indistinguishability (Voice-Ind), for enhancing FL in the scenario of Speech Emotion Recognition (SER), against gender inference attacks. UDP is a general-purpose privacy notion, whereas Voice-Ind is proposed for protecting voiceprint. In addition, we propose a new privacy notion Gender-Indistinguishability (Gender-Ind), which is specifically designed for protecting gender information in speech data, and test its privacy-utility tradeoff compared with the above two privacy notions. The experiments reveal that our specifically designed privacy notion, Gender-Ind, can achieve better utility while preventing the same level of attacks. This finding sheds some light on how to design privacy protection methods in speech data processing. Yang Cao 0011, Sheng Li 0010, Masatoshi Yoshikawa |
ICASSP | 3 |
| 2023 | Speakeraugment: Data Augmentation for Generalizable Source Separation via Speaker Parameter ManipulationabstractExisting speech separation models based on deep learning typically generalize poorly due to domain mismatch. In this paper, we propose SpeakerAugment (SA), a data augmentation method for generalizable speech separation that aims to increase the diversity of speaker identity in training data, to mitigate speaker mismatch of domain mismatch. The SA consists of two sub-policies: (1) SA-Vocoder, which uses a vocoder to manipulate pitch and formants parameters of speakers. (2) SA-Spectrum, which directly performs pitch-shift and time-stretch on the spectrum of each speech signal. The SA is simple and effective. Experimental results show that using SA can significantly improve the generalization ability of models, especially for: 1) The training set with fewer speakers, e.g., WSJ0-2mix, or 2) The target test set with complex linguistic conditions, e.g., the TIMIT based test set. Moreover, as a data augmentation method, SA has good potential to be applicable to other speech related tasks. We validate this by applying SA in speech recognition, and experimental results show that the generalization ability is also improved. Hao Huang 0009, Ying Hu 0005, Sheng Li 0010 |
ICASSP | 5 |
| 2023 | Speech-Text Based Multi-Modal Training with Bidirectional Attention for Improved Speech RecognitionabstractTo let the state-of-the-art end-to-end ASR model enjoy data efficiency, as well as much more unpaired text data by multi-modal training, one needs to address two problems: 1) the synchronicity of feature sampling rates between speech and language (aka text data); 2) the homogeneity of the learned representations from two encoders. In this paper we propose to employ a novel bidirectional attention mechanism (BiAM) to jointly learn both ASR encoder (bottom layers) and text encoder with a multi-modal learning method. The BiAM is to facilitate feature sampling rate exchange, realizing the quality of the transformed features for the one kind to be measured in another space, with diversified objective functions. As a result, the speech representations are enriched with more linguistic information, while the representations generated by the text encoder are more similar to corresponding speech ones, and therefore the shared ASR models are more amenable for unpaired text data pretraining. To validate the efficacy of the proposed method, we perform two categories of experiments with or without extra unpaired text data. Experimental results on Librispeech corpus show it can achieve up to 6.15% word error rate reduction (WERR) with only paired data learning, while 9.23% WERR when more unpaired text data is employed1. Haihua Xu 0001, Hao Huang 0009, Chng Eng Siong, Sheng Li 0010 |
ICASSP | 5 |
| 2023 | Reprogramming Self-supervised Learning-based Speech Representations for Speaker AnonymizationabstractCurrent speaker anonymization methods, especially with self-supervised learning (SSL) models, require massive computational resources when hiding speaker identity. This paper proposes an effective and parameter-efficient speaker anonymization method based on recent End-to-End model reprogramming technology. To improve the anonymization performance, we first extract speaker representation from large SSL models as the speaker identifies. To hide the speaker’s identity, we reprogram the speaker representation by adapting the speaker to a pseudo domain. Extensive experiments are carried out on the VoicePrivacy Challenge (VPC) 2022 datasets to demonstrate the effectiveness of our proposed parameter-efficient learning anonymization methods. Additionally, while achieving comparable performance with the VPC 2022 strong baseline 1.b, our approach also consumes less computational resources during anonymization. Sheng Li 0010, Jiyi Li, Hao Huang 0009, Yang Cao 0011, Liang He 0003 |
MMAsia | 2 |
| 2023 | GhostVec: A New Threat to Speaker Privacy of End-to-End Speech Recognition SystemabstractSpeaker adaptation systems face privacy concerns, for such systems are trained on private datasets and often overfitting. This paper demonstrates that an attacker can extract speaker information by querying speaker-adapted speech recognition (ASR) systems. We focus on the speaker information of a transformer-based ASR and propose GhostVec, a simple and efficient attack method to extract the speaker information from an encoder-decoder-based ASR system without any external speaker verification system or natural human voice as a reference. To make our results quantitative, we pre-process GhostVec using singular value decomposition (SVD) and synthesize it into waveform. Experiment results show that the synthesized audio of GhostVec reaches 10.83% EER and 0.47 minDCF with target speakers, which suggests the effectiveness of the proposed method. We hope the preliminary discovery in this study to catalyze future speech recognition research on privacy-preserving topics. Sheng Li 0010, Jiyi Li, Yang Cao 0011, Hao Huang 0009, Liang He 0003 |
MMAsia | 2 |
| 2023 | Disordered speech recognition considering low resources and abnormal articulation
Yuqin Lin, Jianwu Dang 0001, Longbiao Wang, Sheng Li 0010, Chenchen Ding |
Speech Commun. | 4 |
| 2022 | Compressing Transformer-Based ASR Model by Task-Driven Loss and Attention-Based Multi-Level Feature DistillationabstractThe current popular knowledge distillation (KD) methods effectively compress the transformer-based end-to-end speech recognition model. However, existing methods fail to utilize complete information of the teacher model, and they distill only a limited number of blocks of the teacher model. In this study, we first integrate a task-driven loss function into the decoder’s intermediate blocks to generate task-related feature representations. Then, we propose an attention-based multi-level feature distillation to automatically learn the feature representation summarized by all blocks of the teacher model. Under the 1.1M parameters model, the experimental results on the Wall Street Journal dataset reveal that our approach achieves a 12.1% WER reduction compared with the baseline system. Yongjie Lv, Longbiao Wang, Meng Ge, Sheng Li 0010, Chenchen Ding, Lixin Pan, Yuguang Wang 0003, Jianwu Dang 0001, Kiyoshi Honda |
ICASSP | 4 |
| 2022 | Mining Hard Samples Locally And Globally For Improved Speech SeparationabstractSpeech separation dataset typically consists of hard and non-hard samples, and the former is minority and latter majority. The data imbalance problem biases the model towards non-hard samples and weakens the generalization capability. Given that the average separation performance is sufficiently good, improving hard samples may contribute more to back-end tasks. In this paper, we propose two methods to alleviate data imbalance in speech separation task, based on local and global hard sample mining. For the local, we propose weighted loss to compensate for hard samples by increasing their weights in each batch. For the global, we perform global hard sample mining and re-sample to increase the proportion of hard samples in the training set. Because hard sample mining using objective loss in dynamic mixing leads to local results, we propose an indirect method using speaker-specific parameters, based on the fact that pitch median difference and x-vector cosine distance of two speakers in a mixture are closely correlated with separation SI-SNRi. Experimental results show that both methods decrease the percentage of hard samples in the test set than using dynamic mixing only while keeping the average SI-SNRi comparable, and the global method shows more promising results than the local one. Yizhou Peng, Hao Huang 0009, Ying Hu 0005, Sheng Li 0010 |
ICASSP | 5 |
| 2022 | GhostVec: Directly Extracting Speaker Embedding from End-to-End Speech Recognition Model Using Adversarial Examples
Sheng Li 0010, Hao Huang 0009 |
ICONIP (6) | 2 |
| 2022 | An End-to-End Chinese and Japanese Bilingual Speech Recognition Systems with Shared Character Decomposition
Sheng Li 0010, Jiyi Li, Qianying Liu, Zhuo Gong |
ICONIP (6) | 1 |
| 2022 | Investigating Effective Domain Adaptation Method for Speaker Verification Task
Guangxing Li, Wangjin Zhou, Sheng Li 0010, Yi Zhao 0006, Hao Huang 0009 |
ICONIP (6) | 3 |
| 2022 | Leveraging Simultaneous Translation for Enhancing Transcription of Low-resource Language via Cross Attention Mechanism
Soky Kak, Sheng Li 0010, Masato Mimura, Chenhui Chu, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2022 | Data Augmentation Using McAdams-Coefficient-Based Speaker Anonymization for Fake Audio DetectionabstractFake audio detection (FAD) is a technique to distinguish synthetic speech from natural speech.In most FAD systems, removing irrelevant features from acoustic speech while keeping only robust discriminative features is essential.Intuitively, speaker information entangled in acoustic speech should be suppressed for the FAD task.Particularly in a deep neural network (DNN)-based FAD system, the learning system may learn speaker information from a training dataset and cannot generalize well on a testing dataset.In this paper, we propose to use the speaker anonymization (SA) technique to suppress speaker information from acoustic speech before inputting it into a DNN-based FAD system.We adopted the McAdamscoefficient-based SA (MC-SA) algorithm, and this is expected that the entangled speaker information will not be involved in the DNN-based FAD learning.Based on this idea, we implemented a light convolutional neural network bidirectional long short-term memory (LCNN-BLSTM)-based FAD system and conducted experiments on the Audio Deep Synthesis Detection Challenge (ADD2022) datasets.The results showed that removing the speaker information from acoustic speech improved the relative performance in the first track of ADD2022 by 17.66%. Kai Li 0018, Sheng Li 0010, Xugang Lu, Masato Akagi, Meng Liu 0017, Lin Zhang 0054, Chang Zeng, Longbiao Wang, Jianwu Dang 0001, Masashi Unoki |
INTERSPEECH | 2 |
| 2022 | Global Signal-to-noise Ratio Estimation Based on Multi-subband Processing Using Convolutional Neural Network
Meng Ge, Longbiao Wang, Masashi Unoki, Sheng Li 0010, Jianwu Dang 0001 |
INTERSPEECH | 5 |
| 2022 | Finer-grained Modeling units-based Meta-Learning for Low-resource Tibetan Speech Recognition
Siqing Qin, Longbiao Wang, Sheng Li 0010, Yuqin Lin, Jianwu Dang 0001 |
INTERSPEECH | 3 |
| 2022 | Monaural Speech Enhancement Based on Spectrogram Decomposition for Convolutional Neural Network-sensitive Feature Extraction
Longbiao Wang, Sheng Li 0010, Jianwu Dang 0001, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2022 | Augmented Adversarial Self-Supervised Learning for Early-Stage Alzheimer's Speech Detection
Longfei Yang, Wenqing Wei, Sheng Li 0010, Jiyi Li, Takahiro Shinozaki |
INTERSPEECH | 3 |
| 2022 | Fusion of Self-supervised Learned Models for MOS PredictionabstractWe participated in the mean opinion score (MOS) prediction challenge, 2022.This challenge aims to predict MOS scores of synthetic speech on two tracks, the main track and a more challenging sub-track: out-of-domain (OOD).To improve the accuracy of the predicted scores, we have explored several model fusion-related strategies and proposed a fused framework in which seven pretrained self-supervised learned (SSL) models have been engaged.These pretrained SSL models are derived from three ASR frameworks, including Wav2Vec, Hubert, and WavLM.For the OOD track, we followed the 7 SSL models selected on the main track and adopted a semi-supervised learning method to exploit the unlabeled data.According to the official analysis results, our system has achieved 1 st rank in 6 out of 16 metrics and is one of the top 3 systems for 13 out of 16 metrics.Specifically, we have achieved the highest LCC, SRCC, and KTAU scores at the system level on main track, as well as the best performance on the LCC, SRCC, and KTAU evaluation metrics at the utterance level on OOD track.Compared with the basic SSL models, the prediction accuracy of the fused system has been largely improved, especially on OOD sub-track. Zhengdong Yang, Wangjin Zhou, Chenhui Chu, Sheng Li 0010, Raj Dabre, Raphaël Rubino, Yi Zhao 0006 |
INTERSPEECH | 4 |
| 2022 | Adversarial Speech Generation and Natural Speech Recovery for Speech Content ProtectionabstractWith the advent of the General Data Protection Regulation (GDPR) and increasing privacy concerns, the sharing of speech data is faced with significant challenges. Protecting the sensitive content of speech is the same important as the voiceprint. This paper proposes an effective speech content protection method by constructing a frame-by-frame adversarial speech generation system. We revisited the adversarial examples generating method in the recent machine learning field and selected the phonetic state sequence of sensitive speech for the adversarial examples generation. We build an adversarial speech collection. Moreover, based on the speech collection, we proposed a neural network-based frame-by-frame mapping method to recover the speech content by converting from the adversarial speech to the human speech. Experiment shows our proposed method can encode and recover any sensitive audio, and our method is easy to be conducted with publicly available resources of speech recognition technology. Sheng Li 0010, Jiyi Li, Qianying Liu, Zhuo Gong |
LREC | 1 |
| 2022 | Multi-Domain Dialogue State Tracking with Top-K Slot Self AttentionabstractAs an important component of task-oriented dialogue systems, dialogue state tracking is designed to track the dialogue state through the conversations between users and systems.Multi-domain dialogue state tracking is a challenging task, in which the correlation among different domains and slots needs to consider.Recently, slot self-attention is proposed to provide a data-driven manner to handle it.However, a full-support slot self-attention may involve redundant information interchange.In this paper, we propose a top-k attention-based slot self-attention for multi-domain dialogue state tracking.In the slot self-attention layers, we force each slot to involve information from the other k prominent slots and mask the rest out.The experimental results on two mainstream multi-domain task-oriented dialogue datasets, MultiWOZ 2.0 and MultiWOZ 2.4, present that our proposed approach is effective to improve the performance of multi-domain dialogue state tracking.We also find that the best result is obtained when each slot interchanges information with only a few slots. Longfei Yang, Jiyi Li, Sheng Li 0010, Takahiro Shinozaki |
SIGDIAL | 3 |
| 2021 | An Investigation of Using Hybrid Modeling Units for Improving End-to-End Speech Recognition SystemabstractThe acoustic modeling unit is crucial for an end-to-end speech recognition system, especially for the Mandarin language. Until now, most of the studies on Mandarin speech recognition focused on individual units, and few of them paid attention to using a combination of these units. This paper uses a hybrid of the syllable, Chinese character, and subword as the modeling units for the end-to-end speech recognition system based on the CTC/attention multi-task learning. In this approach, the character-subword unit is assigned to train the transformer model in the main task learning stage. In contrast, the syllable unit is assigned to enhance the transformer’s shared encoder in the auxiliary task stage with the Connectionist Temporal Classification (CTC) loss function. The recognition experiments were conducted on AISHELL-1 and an open data set of 1200-hour Mandarin speech corpus collected from the OpenSLR, respectively. The experimental results demonstrated that using the syllable-char-subword hybrid modeling unit can achieve better performances than the conventional units of char-subword, and 6.6% relative CER reduction on our 1200-hour data. The substitution error also achieves a considerable reduction. Shunfei Chen, Xinhui Hu, Sheng Li 0010, Xinkang Xu |
ICASSP | 3 |
| 2021 | Encoder-Decoder Based Pitch Tracking and Joint Model Training for Mandarin Tone ClassificationabstractWe pursue an interpretable pitch tracking model and a jointly trained tone model for Mandarin tone classification. For pitch tracking, present deep learning based pitch model structure seldom considers the Viterbi decoding commonly implemented in prevalent manually designed pitch tracking algorithms. We propose RNN based Encoder-Decoder framework with gating mechanism which underlying models both the state cost estimation and Viterbi back-tracing pass implemented in the RAPT algorithm. Then we apply the pitch extractor to a down-stream Mandarin tone classification task. The basic motivation is to combine together the two conventional components in tone classification (i.e., the pitch extractor and tone classifier) and then the whole network are trained simultaneously in an end-to-end fashion. Various cascade methods are evaluated. We carry out pitch extraction and tone classification experiments on Mandarin continuous speech database to show the superiority of the proposed models. Experimental results on pitch extraction show proposed pitch tracking model outperforms the DNN-RNN and bi-directional variants. Tone classification experimental results show the composite model outperforms the traditional cascade tone classification framework which makes use of pitch related feature and a back-end classifier. Hao Huang 0009, Ying Hu 0005, Sheng Li 0010 |
ICASSP | 4 |
| 2021 | Robust Voice Activity Detection Using a Masked Auditory Encoder Based Convolutional Neural NetworkabstractVoice activity detection (VAD) based on deep learning has achieved remarkable success. However, when the traditional features (e.g., raw waveforms and MFCCs) are directly fed to the deep neural network model, the performance decreases because of noise interference. Here, we propose a robust VAD approach using a masked auditory encoder based convolutional neural network (M-AECNN). First, we analyze the effectiveness of using auditory features as deep learning encoder. These features can roughly simulate the transmission of sound to human inner-ear hair cells; thus, they are more robust than the raw waveform and frequency domain features designed as encoders. Second, similar to the human ear’s masking effect for different speech frequencies, the proposed auditory encoder can further improve the robustness of VAD by increasing the gain for cleaner speech frequencies. Extensive experimental results demonstrate that this approach achieves about 10.5% absolute improvement in the area under the curve on the AURORA-2J dataset compared with a VAD method based on a CNN and MFCCs. Longbiao Wang, Masashi Unoki, Sheng Li 0010, Rui Wang 0102, Meng Ge, Jianwu Dang 0001 |
ICASSP | 4 |
| 2021 | Exploring Effective Speech Representation via ASR for High-Quality End-to-End Multispeaker TTS
Longbiao Wang, Sheng Li 0010, Chenchen Ding, Ju Zhang 0001, Jianwu Dang 0001 |
ICONIP (6) | 3 |
| 2021 | Speech Dereverberation Based on Scale-Aware Mean Square Error Loss
Luya Qiang, Meng Ge, Longbiao Wang, Sheng Li 0010, Jianwu Dang 0001 |
ICONIP (5) | 7 |
| 2021 | Simultaneous Progressive Filtering-Based Monaural Speech Enhancement
Longbiao Wang, Luya Qiang, Sheng Li 0010, Meng Ge, Gaoyan Zhang, Jianwu Dang 0001 |
ICONIP (5) | 5 |
| 2021 | End-to-End Speech Separation Using Orthogonal Representation in Complex and Real Time-Frequency Domain
Hao Huang 0009, Ying Hu 0005, Sheng Li 0010 |
Interspeech | 5 |
| 2021 | An End-to-End Dialect Identification System with Transfer Learning from a Multilingual Automatic Speech Recognition Model
Shuaishuai Ye, Xinhui Hu, Sheng Li 0010, Xinkang Xu |
Interspeech | 4 |
| 2020 | Voice-Indistinguishability - Protecting Voiceprint with Differential Privacy under an Untrusted ServerabstractWith the rising adoption of advanced voice-based technology together with increasing consumer demand for smart devices, voice-controlled "virtual assistants" such as Apple's Siri and Google Assistant have been integrated into people's daily lives. However, privacy and security concerns may hinder the development of such voice-based applications since speech data contain the speaker's biometric identifier, i.e., voiceprint (as analogous to fingerprint). To alleviate privacy concerns in speech data collection, we propose a fast speech data de-identification system that allows a user to share her speech data with formal privacy guarantee to an untrusted server. Our open-sourced system can be easily integrated into other speech processing systems for collecting users' voice data in a privacy-preserving way. Experiments on public datasets verify the effectiveness and efficiency of the proposed system. Yaowei Han, Yang Cao 0011, Sheng Li 0010, Qiang Ma 0001, Masatoshi Yoshikawa |
CCS | 3 |
| 2020 | End-to-End Articulatory Modeling for Dysarthric Articulatory Attribute DetectionabstractIn this study, we focus on detecting articulatory attribute errors for dysarthric patients with cerebral palsy (CP) or amyotrophic lateral sclerosis (ALS). There are two major challenges for this task. The pronunciation of dysarthric patients is unclear and inaccurate, which results in poor performances of traditional automatic speech recognition (ASR) systems and traditional automatic speech attribute transcription (ASAT). In addition, the data is limited because of the difficulty of recording. This study proposes an end-to-end automatic speech attribute transcription (E2E-ASAT) method for detecting articulatory attribute errors more precisely. To use the limited data more effectively, the parameters of the acoustic model are refactored into two layers and only one layer is retrained. Our proposed method showed good performances in both ASR and articulatory attribute detection. Our system has a potential as a rehabilitation tool. Yuqin Lin, Longbiao Wang, Jianwu Dang 0001, Sheng Li 0010, Chenchen Ding |
ICASSP | 4 |
| 2020 | Spectrograms Fusion with Minimum Difference Masks Estimation for Monaural Speech DereverberationabstractSpectrograms fusion is an effective method for incorporating complementary speech dereverberation systems. Previous linear spectrograms fusion by averaging multiple spectrograms shows outstanding performance. However, various systems with different features cannot apply this simple method. In this study, we design the minimum difference masks (MDMs) to classify the time-frequency (T-F) bins in spectrograms according to the nearest distances from labels. Then, we propose a two-stage nonlinear spectrograms fusion system for speech dereverberation. First, we conduct a multitarget learning-based speech dereverberation front-end model to get spectrograms simultaneously. Then, MDMs are estimated to take the best parts of different spectrograms. We are using spectrograms in the first stage and MDMs in the second stage to recombine T-F bins. The experiments on the REVERB challenge show that a strong feature complementarity between spectrograms and MDMs. Moreover, the proposed framework can consistently and significantly improve PESQ and SRMR, both real and simulated data, e.g., an average PESQ gain of 0.1 in all simulated data and an average SRMR gain of 1.22 in all real data. Longbiao Wang, Meng Ge, Sheng Li 0010, Jianwu Dang 0001 |
ICASSP | 4 |
| 2020 | Voice-Indistinguishability: Protecting Voiceprint In Privacy-Preserving Speech Data ReleaseabstractWith the development of smart devices, such as the Amazon Echo and Apple's HomePod, speech data have become a new dimension of big data. However, privacy and security concerns may hinder the collection and sharing of real-world speech data, which contain the speaker's identifiable information, i.e., voiceprint, which is considered a type of biometric identifier. Current studies on voiceprint privacy protection do not provide either a meaningful privacy-utility trade-off or a formal and rigorous definition of privacy. In this study, we design a novel and rigorous privacy metric for voiceprint privacy, which is referred to as voice-indistinguishability, by extending differential privacy. We also propose mechanisms and frameworks for privacy-preserving speech data release satisfying voice-indistinguishability. Experiments on public datasets verify the effectiveness and efficiency of the proposed methods. Yaowei Han, Sheng Li 0010, Yang Cao 0011, Qiang Ma 0001, Masatoshi Yoshikawa |
ICME | 2 |
| 2020 | Investigation of Effectively Synthesizing Code-Switched Speech Using Highly Imbalanced Mix-Lingual Data
Shaotong Guo, Longbiao Wang, Sheng Li 0010, Ju Zhang 0001, Yuguang Wang 0003, Jianwu Dang 0001, Kiyoshi Honda |
ICONIP (1) | 3 |
| 2020 | Staged Knowledge Distillation for End-to-End Dysarthric Speech Recognition and Speech Attribute Transcription
Yuqin Lin, Longbiao Wang, Sheng Li 0010, Jianwu Dang 0001, Chenchen Ding |
INTERSPEECH | 3 |
| 2020 | Singing Voice Extraction with Attention-Based Spectrograms Fusion
Longbiao Wang, Sheng Li 0010, Chenchen Ding, Meng Ge, Jianwu Dang 0001, Hiroshi Seki |
INTERSPEECH | 3 |
| 2020 | Knowledge Distillation-Based Representation Learning for Short-Utterance Spoken Language IdentificationabstractWith successful applications of deep feature learning algorithms, spoken language identification (LID) on long utterances obtains satisfactory performance. However, the performance on short utterances is drastically degraded even when the LID system is trained using short utterances. The main reason is due to the large variation of the representation on short utterances which results in high model confusion. To narrow the performance gap between long, and short utterances, we proposed a teacher-student representation learning framework based on a knowledge distillation method to improve LID performance on short utterances. In the proposed framework, in addition to training the student model on short utterances with their true labels, the internal representation from the output of a hidden layer of the student model is supervised with the representation corresponding to their longer utterances. By reducing the distance of internal representations between short, and long utterances, the student model can explore robust discriminative representations for short utterances, which is expected to reduce model confusion. We conducted experiments on our in-house LID dataset, and NIST LRE07 dataset, and showed the effectiveness of the proposed methods for short utterance LID tasks. Xugang Lu, Sheng Li 0010, Hisashi Kawai |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Interactive Learning of Teacher-student Model for Short Utterance Spoken Language IdentificationabstractShort utterance-based spoken language identification (LID) is a challenging task due to the large variation of its feature representation. Improving feature representation of short utterances using a teacher-student method has been shown its effectiveness for LID tasks. However, conventional teacher-student methods use fixed pre-trained teacher models, that makes it difficult to optimize student models. In this paper, rather than using a fixed pre-trained teacher model, we investigate an interactive teacher-student learning by adjusting the teacher model with reference to the performance of the student model when the student model is stuck in a local minimum. Experiments on a 10-language LID task were carried out to test the algorithm. Our results showed its effectiveness of the proposed algorithm on short utterance LID tasks. Xugang Lu, Sheng Li 0010, Hisashi Kawai |
ICASSP | 3 |
| 2019 | Investigation of Sequence-level Knowledge Distillation Methods for CTC Acoustic ModelsabstractThis paper presents knowledge distillation (KD) methods for training connectionist temporal classification (CTC) acoustic models. In a previous study, we proposed a KD method based on the sequence-level cross-entropy, and showed that the conventional KD method based on the frame-level cross-entropy did not work effectively for CTC acoustic models, whereas the proposed method improved the performance of the models. In this paper, we investigate the implementation of sequence-level KD for CTC models and propose a lattice-based sequence-level KD method. Experiments investigating model compression and the training of a noise-robust model using the Wall Street Journal (WSJ) and CHiME4 datasets demonstrate that the sequence-level KD methods improve the performance of CTC acoustic models on both two tasks, and show that the lattice-based method can compute the sequence-level KD more efficiently than the N-best-based method proposed in our previous work. Ryoichi Takashima, Sheng Li 0010, Hisashi Kawai |
ICASSP | 2 |
| 2019 | End-to-End Articulatory Attribute Modeling for Low-Resource Multilingual Speech Recognition
Sheng Li 0010, Chenchen Ding, Xugang Lu, Tatsuya Kawahara, Hisashi Kawai |
INTERSPEECH | 1 |
| 2019 | Investigating Radical-Based End-to-End Speech Recognition Systems for Chinese Dialects and Japanese
Sheng Li 0010, Xugang Lu, Chenchen Ding, Tatsuya Kawahara, Hisashi Kawai |
INTERSPEECH | 1 |
| 2019 | Improving Transformer-Based Speech Recognition Systems with Compressed Structure and Speech Attributes Augmentation
Sheng Li 0010, Raj Dabre, Xugang Lu, Tatsuya Kawahara, Hisashi Kawai |
INTERSPEECH | 1 |
| 2019 | Class-Wise Centroid Distance Metric Learning for Acoustic Event Detection
Xugang Lu, Sheng Li 0010, Yu Tsao 0001, Hisashi Kawai |
INTERSPEECH | 3 |
| 2018 | An Investigation of a Knowledge Distillation Method for CTC Acoustic ModelsabstractEnd-to-end acoustic models, such as connectionist temporal classification (CTC) and the attention model, have been studied, and their speech recognition accuracies come close to those of conventional deep neural network (DNN)-hidden Markov models. However, most high-performance end-to-end models are not suitable for real-time (streaming) speech recognition because they are based on bidirectional recurrent neural networks (RNNs). In this study, to improve the performance of unidirectional RNN-based CTC, which is suitable for real-time processing, we investigate the knowledge distillation (KD)-based model compression method for training a CTC acoustic model. we evaluate a frame-level KD method and a sequence-level KD method for CTC model. The speech recognition experiments on Wall Street Journal tasks demonstrate that, the frame-level KD worsens the WERs ofunidirectional CTC model, whereas sequence-level KD can improve the WERs of the model. Ryoichi Takashima, Sheng Li 0010, Hisashi Kawai |
ICASSP | 2 |
| 2018 | CTC Loss Function with a Unit-Level Ambiguity PenaltyabstractThis paper presents a modified loss function for training connectionist temporal classification (CTC)-based acoustic models. CTC-based acoustic models have been studied as alternatives to conventional hidden Markov models (HMMs), but have often shown worse performance than conventional deep neural network (DNN)-HMM hybrid models. In this paper, we attempt to identify the primary factor preventing CTC-based models from achieving their full potential, and hypothesize this constraint lies in the ambiguity in the identification boundaries among unit-level labels (phonemes or characters). In accordance with this hypothesis, we propose a modified CTC loss function using an ambiguity penalty. This penalty is defined by the conditional entropy and works to increase the separation metrics among unit-level labels. We evaluate the proposed method on the WSJ and CHiME4 tasks, and demonstrate that our modification improves the word error rate compared with that of the conventional CTC-based model when the training dataset is small. Ryoichi Takashima, Sheng Li 0010, Hisashi Kawai |
ICASSP | 2 |
| 2018 | Improving CTC-based Acoustic Model with Very Deep Residual Time-delay Neural Networks
Sheng Li 0010, Xugang Lu, Ryoichi Takashima, Tatsuya Kawahara, Hisashi Kawai |
INTERSPEECH | 1 |
| 2018 | Temporal Attentive Pooling for Acoustic Event Detection
Xugang Lu, Sheng Li 0010, Yu Tsao 0001, Hisashi Kawai |
INTERSPEECH | 3 |
| 2018 | Feature Representation of Short Utterances Based on Knowledge Distillation for Spoken Language Identification
Xugang Lu, Sheng Li 0010, Hisashi Kawai |
INTERSPEECH | 3 |
| 2018 | Improving Very Deep Time-Delay Neural Network With Vertical-Attention For Effectively Training CTC-Based ASR SystemsabstractThe very deep neural network has recently been proposed for speech recognition and achieves significant performance. It has excellent potential for integration with end-to-end (E2E) training. Connectionist temporal classification (CTC) has shown great potential in E2E acoustic modeling. In this study, we investigate deep architectures and techniques which are suitable for CTC-based acoustic modeling. We propose a very deep residual time-delay CTC neural network (VResTD-CTC). How to select a suitable deep architecture optimized with the CTC objective function is crucial for obtaining the state of the art performance. Excellent performances can be obtained by selecting deep architecture for non-E2E ASR systems modeling with tied-triphone states. However, these optimized structures do not guarantee to achieve better or comparable performances on E2E (e.g., CTC-based) systems modeling with dynamic acoustic units. For solving this problem and further leveraging the system performance, we introduce the vertical-attention mechanism to reweight the residual blocks at each time step. Speech recognition experiments show our proposed model significantly outperforms the DNN and LSTM-based (both bidirectional and unidirectional) CTC baseline models. Sheng Li 0010, Xugang Lu, Ryoichi Takashima, Tatsuya Kawahara, Hisashi Kawai |
SLT | 1 |
| 2017 | Incremental training and constructing the very deep convolutional residual network acoustic modelsabstractInspired by the successful applications in image recognition, the very deep convolutional residual network (ResNet) based model has been applied in automatic speech recognition (ASR). However, the computational load is heavy for training the ResNet with a large quantity of data. In this paper, we propose an incremental model training framework to accelerate the training process of the ResNet. The incremental model training framework is based on the unequal importance of each layer and connection in the ResNet. The modules with important layers and connections are regarded as a skeleton model, while those left are regarded as an auxiliary model. The total depth of the skeleton model is quite shallow compared to the very deep full network. In our incremental training, the skeleton model is first trained with the full training data set. Other layers and connections belonging to the auxiliary model are gradually attached to the skeleton model and tuned. Our experiments showed that the proposed incremental training obtained comparable performances and faster training speed compared with the model training as a whole without consideration of the different importance of each layer. Sheng Li 0010, Xugang Lu, Ryoichi Takashima, Tatsuya Kawahara, Hisashi Kawai |
ASRU | 1 |
| 2017 | Semi-supervised ensemble DNN acoustic model trainingabstractIt is very important to exploit abundant unlabeled speech for improving the acoustic model training in automatic speech recognition (ASR). Semi-supervised training methods incorporate unlabeled data in addition to labeled data to enhance the model training, but it encounters the error-prone label problem. The ensemble training scheme trains a set of models and combines them to make the model more general and robust, but it has not been applied to the unlabeled data. In this work, we propose an effective semi-supervised training of deep neural network (DNN) acoustic models by incorporating the diversity among the ensemble of models. The resultant model improved the performance in the lecture transcription task. Moreover, the proposed method has also shown a potential for DNN adaptation. Sheng Li 0010, Xugang Lu, Shinsuke Sakai, Masato Mimura, Tatsuya Kawahara |
ICASSP | 1 |
| 2017 | Conditional Generative Adversarial Nets Classifier for Spoken Language Identification
Xugang Lu, Sheng Li 0010, Hisashi Kawai |
INTERSPEECH | 3 |
| 2016 | Data selection from multiple ASR systems' hypotheses for unsupervised acoustic model trainingabstractThis paper addresses unsupervised training of DNN acoustic model, by exploiting a large amount of unlabeled data with CRF-based classifiers. In the proposed scheme, we obtain ASR hypotheses by complementary GMM and DNN based ASR systems. Then, a set of dedicated classifiers are designed and trained to select the better hypothesis and verify the selected data. It is demonstrated that the classifiers can effectively filter usable data from unlabeled data for acoustic model training. The proposed method achieved significant improvement in the ASR accuracy from the baseline system, and it outperformed the models trained from the data selected based on the confidence measure scores (CMS) and also from the simple ROVER-based system combination. Sheng Li 0010, Yuya Akita, Tatsuya Kawahara |
ICASSP | 1 |
| 2016 | Semi-Supervised Acoustic Model Training by Discriminative Data Selection From Multiple ASR Systems' HypothesesabstractWhile the performance of ASR systems depends on the size of the training data, it is very costly to prepare accurate and faithful transcripts. In this paper, we investigate a semisupervised training scheme, which takes the advantage of huge quantities of unlabeled video lecture archive, particularly for the deep neural network (DNN) acoustic model. In the proposed method, we obtain ASR hypotheses by complementary GMM- and DNN-based ASR systems. Then, a set of CRF-based classifiers is trained to select the correct hypotheses and verify the selected data. The proposed hypothesis combination shows higher quality compared with the conventional system combination method (ROVER). Moreover, compared with the conventional data selection based on confidence measure score, our method is demonstrated more effective for filtering usable data. Significant improvement in the ASR accuracy is achieved over the baseline system and in comparison with the models trained with the conventional system combination and data selection methods. Sheng Li 0010, Yuya Akita, Tatsuya Kawahara |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2015 | Discriminative data selection for lightly supervised training of acoustic model using closed caption textsabstractWe present a novel data selection method for lightly supervised training of acoustic model, which exploits a large amount of data with closed caption texts but not faithful transcripts. In the proposed scheme, a sequence of the closed caption text and that of the ASR hypothesis by the baseline system are aligned. Then, a set of dedicated classifiers is designed and trained to select the correct one among them or reject both. It is demonstrated that the classifiers can effectively filter the usable data for acoustic model training without tuning any threshold parameters. A significant improvement in the ASR accuracy is achieved from the baseline system and also in comparison with the conventional method of lightly supervised training based on simple matching and confidence measure scores. Sheng Li 0010, Yuya Akita, Tatsuya Kawahara |
INTERSPEECH | 1 |
| 2015 | Ensemble speaker modeling using speaker adaptive training deep neural network for speaker adaptationabstractIn this paper, we introduce an ensemble speaker modeling using a speaker adaptive training (SAT) deep neural network (SAT-DNN). We first train a speaker-independent DNN (SIDNN) acoustic model as a universal speaker model (USM). Based on the USM, a SAT-DNN is used to obtain a set of speaker-dependent models by assuming that all other layers except one speaker-dependent (SD) layer are shared among speakers. The speaker ensemble matrix is created by concatenating all of the SD neural weight matrices. With matrix factorization technique, an ensemble speaker subspace is extracted. When testing, an initial model for each target speaker is selected in this ensemble speaker subspace. Then, adaptation is carried out to obtain the final acoustic model for testing. In order to reduce the number of adaptation parameters, low-rank speaker subspace is further explored. We test our algorithm on lecture transcription task. Experimental results showed that our proposed method is effective for unsupervised speaker adaptation. Index Terms: speaker adaptation, deep neural networks, ensemble modeling, lecture transcription Sheng Li 0010, Xugang Lu, Yuya Akita, Tatsuya Kawahara |
INTERSPEECH | 1 |
| 2012 | Cross Linguistic Comparison of Mandarin and English EMA Articulatory Data
Sheng Li 0010 |
INTERSPEECH | 1 |
| 2012 | Phoneme-level articulatory animation in pronunciation training
Hui Chen 0020, Sheng Li 0010, Helen M. Meng |
Speech Commun. | 3 |