Di Wu 0061

dblp:52/328-61 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
7since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2024 Hydraformer: One Encoder for All Subsampling Rates
abstract
In automatic speech recognition, subsampling is essential for tackling diverse scenarios. However, the inadequacy of a single subsampling rate to address various real-world situations often necessitates training and deploying multiple models, consequently increasing associated costs. To address this issue, we propose HydraFormer, comprising HydraSub, a Conformer-based encoder, and a BiTransformer-based decoder. HydraSub encompasses multiple branches, each representing a distinct subsampling rate, allowing for the flexible selection of any branch during inference based on the specific use case. HydraFormer can efficiently manage different subsampling rates, significantly reducing training and deployment expenses. Experiments on AISHELL-1 and LibriSpeech datasets reveal that HydraFormer effectively adapts to various subsampling rates and languages while maintaining high recognition performance. Additionally, HydraFormer showcases exceptional stability, sustaining consistent performance under various initialization conditions, and exhibits robust transferability by learning from pretrained single subsampling rate automatic speech recognition models1.
Yaoxun Xu, Xingchen Song, Zhiyong Wu 0001, Di Wu 0061, Zhendong Peng
ICME4
2023 Fast-U2++: Fast and Accurate End-to-End Speech Recognition in Joint CTC/Attention Frames
abstract
Recently, the unified streaming and non-streaming two-pass (U2/U2++) end-to-end model for speech recognition has shown great performance in terms of streaming capability, accuracy and latency. In this paper, we present fast-U2++, an enhanced version of U2++ to further reduce partial latency. The core idea of fast-U2++ is to output partial results of the bottom layers in its encoder with a small chunk, while using a large chunk in the top layers of its encoder to compensate the performance degradation caused by the small chunk. More-over, we use knowledge distillation method to reduce the token emission latency. We present extensive experiments on Aishell-1 dataset. Experiments and ablation studies show that compared to U2++, fast-U2++ reduces model latency from 320ms to 80ms, and achieves a character error rate (CER) of 5.06% with a streaming setup.
Chengdong Liang, Xiao-Lei Zhang 0001, Di Wu 0061, Shengqiang Li, Xingchen Song, Zhendong Peng, Fuping Pan
ICASSP4
2023 TrimTail: Low-Latency Streaming ASR with Simple But Effective Spectrogram-Level Length Penalty
abstract
In this paper, we present TrimTail, a simple but effective emission regularization method to improve the latency of streaming ASR models. The core idea of TrimTail is to apply length penalty (i.e., by trimming trailing frames, see Fig. 1-(b)) directly on the spectrogram of input utterances, which does not require any alignment. We demonstrate that TrimTail is computationally cheap and can be applied online and optimized with any training loss or any model architecture on any dataset without any extra effort by applying it on various end-to-end streaming ASR networks either trained with CTC loss [1] or Transducer loss [2]. We achieve 100 ~ 200ms latency reduction with equal or even better accuracy on both Aishell-1 and Librispeech. Moreover, by using TrimTail, we can achieve a 400ms algorithmic improvement of User Sensitive Delay (USD) with an accuracy loss of less than 0.2.
Xingchen Song, Di Wu 0061, Zhiyong Wu 0001, Yuekai Zhang, Zhendong Peng, Wenpeng Li, Fuping Pan, Changbao Zhu
ICASSP2
2023 ZeroPrompt: Streaming Acoustic Encoders are Zero-Shot Masked LMs
abstract
In this paper, we present ZeroPrompt (Figure 1-(a)) and the corresponding Prompt-and-Refine strategy (Figure 3), two simple but effective training-free methods to decrease the Token Display Time (TDT) of streaming ASR models without any accuracy loss.The core idea of ZeroPrompt is to append zeroed content to each chunk during inference, which acts like a prompt to encourage the model to predict future tokens even before they were spoken.We argue that streaming acoustic encoders naturally have the modeling ability of Masked Language Models and our experiments demonstrate that ZeroPrompt is engineering cheap and can be applied to streaming acoustic encoders on any dataset without any accuracy loss.Specifically, compared with our baseline models, we achieve 350 ∼ 700ms reduction on First Token Display Time (TDT-F) and 100 ∼ 400ms reduction on Last Token Display Time (TDT-L), with theoretically and experimentally equal WER on both Aishell-1 and Librispeech datasets.
Xingchen Song, Di Wu 0061, Zhendong Peng, Bo Dang 0004, Fuping Pan, Zhiyong Wu 0001
INTERSPEECH2
2022 WENETSPEECH: A 10000+ Hours Multi-Domain Mandarin Corpus for Speech Recognition
abstract
In this paper, we present WenetSpeech, a multi-domain Mandarin corpus consisting of 10000+ hours high-quality labeled speech, 2400+ hours weakly labeled speech, and about 10000 hours unlabeled speech, with 22400+ hours in total. We collect the data from YouTube and Podcast, which covers a variety of speaking styles, scenarios, domains, topics and noisy conditions. An optical character recognition (OCR) method is introduced to generate the audio/text segmentation candidates for the YouTube data on the corresponding video subtitles, while a high-quality ASR transcription system is used to generate audio/text pair candidates for the Podcast data. Then we propose a novel end-to-end label error detection approach to further validate and filter the candidates. We also provide three manually labelled high-quality test sets along with WenetSpeech for evaluation – Dev for cross-validation purpose in training, Test_Net, collected from Internet for matched test, and Test_Meeting, recorded from real meetings for more challenging mismatched test. Baseline systems trained with WenetSpeech are provided for three popular speech recognition toolkits, namely Kaldi, ESPnet, and WeNet, and recognition results on the three test sets are also provided as benchmarks. To the best of our knowledge, WenetSpeech is the current largest open-source Mandarin speech corpus with transcriptions, which benefits research on production-level speech recognition.
Hang Lv 0001, Qijie Shao, Chao Yang 0031, Lei Xie 0001, Hui Bu, Chenchen Zeng, Di Wu 0061, Zhendong Peng
ICASSP11
2022 WeNet 2.0: More Productive End-to-End Speech Recognition Toolkit
abstract
Recently, we made available WeNet [1], a production-oriented end-to-end speech recognition toolkit, which introduces a unified two-pass (U2) framework and a built-in runtime to address the streaming and non-streaming decoding modes in a single model.To further improve ASR performance and facilitate various production requirements, in this paper, we present WeNet 2.0 with four important updates.(1) We propose U2++, a unified two-pass framework with bidirectional attention decoders, which includes the future contextual information by a right-toleft attention decoder to improve the representative ability of the shared encoder and the performance during the rescoring stage.(2) We introduce an n-gram based language model and a WFSTbased decoder into WeNet 2.0, promoting the use of rich text data in production scenarios.(3) We design a unified contextual biasing framework, which leverages user-specific context (e.g., contact lists) to provide rapid adaptation ability for production and improves ASR accuracy in both with-LM and without-LM scenarios.(4) We design a unified IO to support large-scale data for effective model training.In summary, the brand-new WeNet 2.0 achieves up to 10% relative recognition performance improvement over the original WeNet on various corpora and makes available several important production-oriented features.
Di Wu 0061, Zhendong Peng, Xingchen Song, Zhuoyuan Yao, Hang Lv 0001, Lei Xie 0001, Chao Yang 0031, Fuping Pan, Jianwei Niu 0002
INTERSPEECH2
2021 WeNet: Production Oriented Streaming and Non-Streaming End-to-End Speech Recognition Toolkit
abstract
In this paper, we propose an open source speech recognition toolkit called WeNet, in which a new two-pass approach named U2 is implemented to unify streaming and non-streaming endto-end (E2E) speech recognition in a single model.The main motivation of WeNet is to close the gap between the research and deployment of E2E speech recognition models.WeNet provides an efficient way to ship automatic speech recognition (ASR) applications in real-world scenarios, which is the main difference and advantage to other open source E2E speech recognition toolkits.We develop a hybird connectionist temporal classification (CTC)/attention architecture with transformer or conformer as encoder and an attention decoder to rescore th CTC hypotheses.To achieve streaming and non-streaming in a unified model, we use a dynamic chunk-based attention strategy which allows the self-attention to focus on the right context with random length.Our experiments on the AISHELL-1 dataset show that our model achieves 5.03% relative character error rate (CER) reduction in non-streaming ASR compared to a standard non-streaming transformer.After model quantification, our model achieves reasonable RTF and latency at runtime.The toolkit is publicly available at https://github.com/mobvoi/wenet.
Zhuoyuan Yao, Di Wu 0061, Fan Yu 0002, Chao Yang 0031, Zhendong Peng, Lei Xie 0001
Interspeech2
2019 Design of Gesture Recognition System Based on Multi-Channel Myoelectricity Correlation
abstract
Gesture recognition systems based on myoelectric signal have raised more and more attention from researchers. Traditional gesture recognition methods are susceptible to multiple types of noise and require a large number of features, which increase overhead and decrease recognition efficiency. Fully utilizing the characteristics of the signal to recognize gestures is a big challenge. This paper proposes an improved empirical mode decomposition method (XB-EMD) based on autocorrelation function to denoise myoelectric signal. In addition, a novel deep neural network (CRNet) which combines the CNN and RNN layers together is trained for classifying the gestures based on denoised myoelectric signal. Experimental results show that the proposed gesture recognition system can improve recognition effectiveness at 97.4% with 10 typical gestures, and 72.55% with 16 complicated gestures.
Di Wu 0061, Huiyong Li 0005, Xuefeng Liu 0001, Jianwei Niu 0002
GLOBECOM1