EDBT 2026 Demo / reviewers in the wild / expert
Chengyi Wang 0002
dblp:75/533-2
· DBLP profile ↗
17ranked-venue papers
7as first author
12since 2021 · last 2023
0000-0002-6780-9299ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 11 · 6 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | BEATs: Audio Pre-Training with Acoustic TokenizersabstractWe introduce a self-supervised learning (SSL) framework BEATs for general audio representation pre-training, where we optimize an acoustic tokenizer and an audio SSL model by iterations. Unlike the previous audio SSL models that employ reconstruction loss for pre-training, our audio SSL model is trained with the discrete label prediction task, where the labels are generated by a semantic-rich acoustic tokenizer. We propose an iterative pipeline to jointly optimize the tokenizer and the pre-trained model, aiming to abstract high-level semantics and discard the redundant details for audio. The experimental results demonstrate our acoustic tokenizers can generate discrete labels with rich audio semantics and our audio SSL models achieve state-of-the-art (SOTA) results across various audio classification benchmarks, even outperforming previous models that use more training data and model parameters significantly. Specifically, we set a new SOTA mAP 50.6% on AudioSet-2M without using any external data, and 98.1% accuracy on ESC-50. The code and pre-trained models are available at https://aka.ms/beats. Sanyuan Chen, Yu Wu 0012, Chengyi Wang 0002, Shujie Liu 0001, Daniel Tompkins, Zhuo Chen 0006, Wanxiang Che, Xiangzhan Yu, Furu Wei |
ICML | 3 |
| 2023 | GRAVO: Learning to Generate Relevant Audio from Visual Features with Noisy Online Videos
Youngdo Ahn, Chengyi Wang 0002, Yu Wu 0012, Jong Won Shin, Shujie Liu 0001 |
INTERSPEECH | 2 |
| 2022 | SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language ProcessingabstractJunyi Ao, Rui Wang, Long Zhou, Chengyi Wang, Shuo Ren, Yu Wu, Shujie Liu, Tom Ko, Qing Li, Yu Zhang, Zhihua Wei, Yao Qian, Jinyu Li, Furu Wei. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Junyi Ao, Rui Wang 0073, Chengyi Wang 0002, Shuo Ren 0002, Yu Wu 0012, Shujie Liu 0001, Tom Ko, Qing Li 0001, Yu Zhang 0006, Zhihua Wei 0001, Yao Qian, Jinyu Li 0001, Furu Wei |
ACL (1) | 4 |
| 2022 | Large-Scale Self-Supervised Speech Representation Learning for Automatic Speaker VerificationabstractThe speech representations learned from large-scale unlabeled data have shown better generalizability than those from supervised learning and thus attract a lot of interest to be applied for various downstream tasks. In this paper, we explore the limits of speech representations learned by different self-supervised objectives and datasets for automatic speaker verification (ASV), especially with a well-recognized SOTA ASV model, ECAPA-TDNN [1], as a downstream model. The representations from all hidden layers of the pre-trained model are firstly averaged with learnable weights and then fed into the ECAPA-TDNN as input features. The experimental results on Voxceleb dataset show that the weighted average representation is significantly superior to FBank, a conventional handcrafted feature for ASV. Our best single system achieves 0.537%, 0.569%, and 1.180% equal error rate (EER) on the three official trials of VoxCeleb1, separately. Accordingly, the ensemble system with three pre-trained models can further improve the EER to 0.479%, 0.536% and 1.023%. Among the three evaluation trials, our best system outperforms the winner system [2] of the VoxCeleb Speaker Recognition Challenge 2021 (VoxSRC2021) on the VoxCeleb1-E trial. Zhengyang Chen, Sanyuan Chen, Yu Wu 0012, Yao Qian, Chengyi Wang 0002, Shujie Liu 0001, Yanmin Qian, Michael Zeng 0001 |
ICASSP | 5 |
| 2022 | Unispeech-Sat: Universal Speech Representation Learning With Speaker Aware Pre-TrainingabstractSelf-supervised learning (SSL) is a long-standing goal for speech processing, since it utilizes large-scale unlabeled data and avoids extensive human labeling. Recent years have witnessed great successes in applying self-supervised learning in speech recognition, while limited exploration was attempted in applying SSL for modeling speaker characteristics. In this paper, we aim to improve the existing SSL framework for speaker representation learning. Two methods are introduced for enhancing the unsupervised speaker information extraction. First, we apply multi-task learning to the current SSL framework, where we integrate utterance-wise contrastive loss with the SSL objective function. Second, for better speaker discrimination, we propose an utterance mixing strategy for data augmentation, where additional overlapped utterances are created unsupervisely and incorporated during training. We integrate the proposed methods into the HuBERT framework. Experiment results on the SUPERB benchmark show that the proposed system achieves state-of-the-art performance in universal representation learning, especially for speaker identification oriented tasks. An ablation study is performed verifying the efficacy of each proposed method. Finally, we scale up the training dataset to 94 thousand hours of public audio data and achieve further performance improvement in all SUPERB tasks. Sanyuan Chen, Yu Wu 0012, Chengyi Wang 0002, Zhengyang Chen, Zhuo Chen 0006, Shujie Liu 0001, Jian Wu 0027, Yao Qian, Furu Wei, Jinyu Li 0001, Xiangzhan Yu |
ICASSP | 3 |
| 2022 | Wav2vec-Switch: Contrastive Learning from Original-Noisy Speech Pairs for Robust Speech RecognitionabstractThe goal of self-supervised learning (SSL) for automatic speech recognition (ASR) is to learn good speech representations from a large amount of unlabeled speech for the downstream ASR task. However, most SSL frameworks do not consider noise robustness which is crucial for real-world applications. In this paper we propose wav2vec-Switch, a method to encode noise robustness into contextualized representations of speech via contrastive learning. Specifically, we feed original-noisy speech pairs simultaneously into the wav2vec 2.0 network. In addition to the existing contrastive learning task, we switch the quantized representations of the original and noisy speech as additional prediction targets of each other. By doing this, it enforces the network to have consistent predictions for the original and noisy speech, thus allows to learn contextualized representation with noise robustness. Our experiments on synthe-sized and real noisy data show the effectiveness of our method: it achieves 2.9–4.9% relative word error rate (WER) reduction on the synthesized noisy LibriSpeech data without deterioration on the original data, and 5.7% on CHiME-4 real 1-channel noisy data compared to a data augmentation baseline even with a strong language model for decoding. Our results on CHiME-4 can match or even surpass those with well-designed speech enhancement components. Jinyu Li 0001, Heming Wang, Yao Qian, Chengyi Wang 0002, Yu Wu 0012 |
ICASSP | 5 |
| 2022 | Improving Noise Robustness of Contrastive Speech Representation Learning with Speech ReconstructionabstractNoise robustness is essential for deploying automatic speech recognition (ASR) systems in real-world environments. One way to reduce the effect of noise interference is to employ a preprocessing module that conducts speech enhancement, and then feed the enhanced speech to an ASR backend. In this work, instead of suppressing background noise with a conventional cascaded pipeline, we employ a noise-robust representation learned by a refined self-supervised framework for noisy speech recognition. We propose to combine a reconstruction module with contrastive learning and perform multi-task continual pre-training on noisy data. The reconstruction module is used for auxiliary learning to improve the noise robustness of the learned representation and thus is not required during inference. Experiments demonstrate the effectiveness of our proposed method. Our model substantially reduces the word error rate (WER) for the synthesized noisy LibriSpeech test sets, and yields around 4.1/7.5% WER reduction on noisy clean/other test sets compared to data augmentation. For the real-world noisy speech from the CHiME-4 challenge (1-channel track), we have obtained the state of the art ASR performance without any denoising front-end. Moreover, we achieve comparable performance to the best supervised approach reported with only 16% of labeled data. Heming Wang, Yao Qian, Xiaofei Wang 0009, Chengyi Wang 0002, Shujie Liu 0001, Takuya Yoshioka, Jinyu Li 0001, DeLiang Wang |
ICASSP | 5 |
| 2022 | Improving Self-Supervised Learning for Speech Recognition with Intermediate Layer SupervisionabstractRecently, pioneer work finds that self-supervised pre-training methods can improve multiple downstream speech tasks, because the model utilizes bottom layers to learn speaker-related information and top layers to encode content-related information. Since the network capacity is limited, we believe the speech recognition performance could be further improved if the model is dedicated to audio content information learning. To this end, we propose Intermediate Layer Supervision for Self-Supervised Learning (ILS-SSL), which forces the model to concentrate on content information as much as possible by adding an additional SSL loss on the intermediate layers. Experiments on LibriSpeech test-other set show that our method outperforms HuBERT significantly, which achieves a 23.5%/11.6% relative word error rate reduction in the w/o language model setting for Base/Large models. Detailed analysis shows the bottom layers of our model have a better correlation with phonetic units, which is consistent with our intuition and explains the success of our method for ASR. We will release our code and model at https://github.com/microsoft/UniSpeech. Chengyi Wang 0002, Yu Wu 0012, Sanyuan Chen, Shujie Liu 0001, Jinyu Li 0001, Yao Qian, Zhenglu Yang |
ICASSP | 1 |
| 2022 | Why does Self-Supervised Learning for Speech Recognition Benefit Speaker Recognition?abstractRecently, self-supervised learning (SSL) has demonstrated strong performance in speaker recognition, even if the pretraining objective is designed for speech recognition.In this paper, we study which factor leads to the success of selfsupervised learning on speaker-related tasks, e.g.speaker verification (SV), through a series of carefully designed experiments.Our empirical results on the Voxceleb-1 dataset suggest that the benefit of SSL to SV task is from a combination of mask speech prediction loss, data scale, and model size, while the SSL quantizer has a minor impact.We further employ the integrated gradients attribution method and loss landscape visualization to understand the effectiveness of self-supervised learning for speaker recognition performance. Sanyuan Chen, Yu Wu 0012, Chengyi Wang 0002, Shujie Liu 0001, Zhuo Chen 0006, Gang Liu 0001, Jinyu Li 0001, Jian Wu 0027, Xiangzhan Yu, Furu Wei |
INTERSPEECH | 3 |
| 2022 | Supervision-Guided Codebooks for Masked Prediction in Speech Pre-trainingabstractRecently, masked prediction pre-training has seen remarkable progress in self-supervised learning (SSL) for speech recognition.It usually requires a codebook obtained in an unsupervised way, making it less accurate and difficult to interpret.We propose two supervision-guided codebook generation approaches to improve automatic speech recognition (ASR) performance and also the pre-training efficiency, either through decoding with a hybrid ASR system to generate phoneme-level alignments (named PBERT), or performing clustering on the supervised speech features extracted from an end-to-end CTC model (named CTC clustering).Both the hybrid and CTC models are trained on the same small amount of labeled speech as used in fine-tuning.Experiments demonstrate significant superiority of our methods to various SSL and self-training baselines, with up to 17.0% relative WER reduction.Our pre-trained models also show good transferability in a non-ASR speech task. Chengyi Wang 0002, Yu Wu 0012, Sanyuan Chen, Jinyu Li 0001, Shujie Liu 0001, Furu Wei |
INTERSPEECH | 1 |
| 2021 | Continuous Speech Separation with ConformerabstractContinuous speech separation was recently proposed to deal with the overlapped speech in natural conversations. While it was shown to significantly improve the speech recognition performance for multichannel conversation transcription, its effectiveness has yet to be proven for a single-channel recording scenario. This paper examines the use of Conformer architecture in lieu of recurrent neural networks for the separation model. Conformer allows the separation model to efficiently capture both local and global context information, which is helpful for speech separation. Experimental results using the LibriCSS dataset show that the Conformer separation model achieves the state of the art results for both single-channel and multi-channel settings. Results for real meeting recordings are also presented, showing significant performance gains in both word error rate (WER) and speaker-attributed WER. Sanyuan Chen, Yu Wu 0012, Zhuo Chen 0006, Jian Wu 0027, Jinyu Li 0001, Takuya Yoshioka, Chengyi Wang 0002, Shujie Liu 0001, Ming Zhou 0001 |
ICASSP | 7 |
| 2021 | UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled DataabstractIn this paper, we propose a unified pre-training approach called UniSpeech to learn speech representations with both labeled and unlabeled data, in which supervised phonetic CTC learning and phonetically-aware contrastive self-supervised learning are conducted in a multi-task learning manner. The resultant representations can capture information more correlated with phonetic structures and improve the generalization across languages and domains. We evaluate the effectiveness of UniSpeech for cross-lingual representation learning on public CommonVoice corpus. The results show that UniSpeech outperforms self-supervised pretraining and supervised transfer learning for speech recognition by a maximum of 13.4% and 26.9% relative phone error rate reductions respectively (averaged over all testing languages). The transferability of UniSpeech is also verified on a domain-shift speech recognition task, i.e., a relative word error rate reduction of 6% against the previous approach. Chengyi Wang 0002, Yu Wu 0012, Yao Qian, Ken'ichi Kumatani, Shujie Liu 0001, Furu Wei, Michael Zeng 0001, Xuedong Huang 0001 |
ICML | 1 |
| 2020 | Bridging the Gap between Pre-Training and Fine-Tuning for End-to-End Speech TranslationabstractEnd-to-end speech translation, a hot topic in recent years, aims to translate a segment of audio into a specific language with an end-to-end model. Conventional approaches employ multi-task learning and pre-training methods for this task, but they suffer from the huge gap between pre-training and fine-tuning. To address these issues, we propose a Tandem Connectionist Encoding Network (TCEN) which bridges the gap by reusing all subnets in fine-tuning, keeping the roles of subnets consistent, and pre-training the attention module. Furthermore, we propose two simple but effective methods to guarantee the speech encoder outputs and the MT encoder inputs are consistent in terms of semantic representation and sequence length. Experimental results show that our model leads to significant improvements in En-De and En-Fr translation irrespective of the backbones. Chengyi Wang 0002, Yu Wu 0012, Shujie Liu 0001, Zhenglu Yang, Ming Zhou 0001 |
AAAI | 1 |
| 2020 | Curriculum Pre-training for End-to-End Speech TranslationabstractEnd-to-end speech translation poses a heavy burden on the encoder because it has to transcribe, understand, and learn cross-lingual semantics simultaneously.To obtain a powerful encoder, traditional methods pre-train it on ASR data to capture speech features.However, we argue that pre-training the encoder only through simple speech recognition is not enough, and high-level linguistic knowledge should be considered.Inspired by this, we propose a curriculum pre-training method that includes an elementary course for transcription learning and two advanced courses for understanding the utterance and mapping words in two languages.The difficulty of these courses is gradually increasing.Experiments show that our curriculum pre-training method leads to significant improvements on En-De and En-Fr speech translation benchmarks. Chengyi Wang 0002, Yu Wu 0012, Shujie Liu 0001, Ming Zhou 0001, Zhenglu Yang |
ACL | 1 |
| 2020 | Semantic Mask for Transformer Based End-to-End Speech RecognitionabstractAttention-based encoder-decoder model has achieved impressive results for both automatic speech recognition (ASR) and text-to-speech (TTS) tasks.This approach takes advantage of the memorization capacity of neural networks to learn the mapping from the input sequence to the output sequence from scratch, without the assumption of prior knowledge such as the alignments.However, this model is prone to overfitting, especially when the amount of training data is limited.Inspired by SpecAugment and BERT, in this paper, we propose a semantic mask based regularization for training such kind of end-toend (E2E) model.The idea is to mask the input features corresponding to a particular output token, e.g., a word or a wordpiece, in order to encourage the model to fill the token based on the contextual information.While this approach is applicable to the encoder-decoder framework with any type of neural network architecture, we study the transformer-based model for ASR in this work.We perform experiments on Librispeech 960h and TedLium2 data sets, and achieve the state-of-the-art performance on the test set in the scope of E2E models. Chengyi Wang 0002, Yu Wu 0012, Yujiao Du, Jinyu Li 0001, Shujie Liu 0001, Liang Lu 0001, Shuo Ren 0002, Guoli Ye, Sheng Zhao 0002, Ming Zhou 0001 |
INTERSPEECH | 1 |
| 2020 | Low Latency End-to-End Streaming Speech Recognition with a Scout NetworkabstractThe attention-based Transformer model has achieved promising results for speech recognition (SR) in the offline mode.However, in the streaming mode, the Transformer model usually incurs significant latency to maintain its recognition accuracy when applying a fixed-length look-ahead window in each encoder layer.In this paper, we propose a novel low-latency streaming approach for Transformer models, which consists of a scout network and a recognition network.The scout network detects the whole word boundary without seeing any future frames, while the recognition network predicts the next subword by utilizing the information from all the frames before the predicted boundary.Our model achieves the best performance (2.7/6.4WER) with only 639 ms latency on the test-clean and test-other data sets of Librispeech. Chengyi Wang 0002, Yu Wu 0012, Liang Lu 0001, Shujie Liu 0001, Jinyu Li 0001, Guoli Ye, Ming Zhou 0001 |
INTERSPEECH | 1 |
| 2020 | On the Comparison of Popular End-to-End Models for Large Scale Speech RecognitionabstractRecently, there has been a strong push to transition from hybrid models to end-to-end (E2E) models for automatic speech recognition.Currently, there are three promising E2E methods: recurrent neural network transducer (RNN-T), RNN attentionbased encoder-decoder (AED), and Transformer-AED.In this study, we conduct an empirical comparison of RNN-T, RNN-AED, and Transformer-AED models, in both non-streaming and streaming modes.We use 65 thousand hours of Microsoft anonymized training data to train these models.As E2E models are more data hungry, it is better to compare their effectiveness with large amount of training data.To the best of our knowledge, no such comprehensive study has been conducted yet.We show that although AED models are stronger than RNN-T in the non-streaming mode, RNN-T is very competitive in streaming mode if its encoder can be properly initialized.Among all three E2E models, transformer-AED achieved the best accuracy in both streaming and non-streaming mode.We show that both streaming RNN-T and transformer-AED models can obtain better accuracy than a highly-optimized hybrid model. Jinyu Li 0001, Yu Wu 0012, Yashesh Gaur, Chengyi Wang 0002, Rui Zhao 0017, Shujie Liu 0001 |
INTERSPEECH | 4 |