Fuping Pan

dblp:83/2566 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
6since 2021 · last 2023
0000-0001-9171-0726ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2023 LightGrad: Lightweight Diffusion Probabilistic Model for Text-to-Speech
abstract
Recent advances in neural text-to-speech (TTS) models bring thousands of TTS applications into daily life, where models are deployed in cloud to provide services for customs. Among these models are diffusion probabilistic models (DPMs), which can be stably trained and are more parameter-efficient compared with other generative models. As transmitting data between customs and the cloud introduces high latency and the risk of exposing private data, deploying TTS models on edge devices is preferred. When implementing DPMs onto edge devices, there are two practical problems. First, current DPMs are not lightweight enough for resource-constrained devices. Second, DPMs require many denoising steps in inference, which increases latency. In this work, we present LightGrad, a lightweight DPM for TTS. LightGrad is equipped with a lightweight U-Net diffusion decoder and a training-free fast sampling technique, reducing both model parameters and inference latency. Streaming inference is also implemented in LightGrad to reduce latency further. Compared with Grad-TTS, LightGrad achieves 62.2% reduction in paramters, 65.7% reduction in latency, while preserving comparable speech quality on both Chinese Mandarin and English in 4 denoising steps1.
Xingchen Song, Zhendong Peng, Fuping Pan, Zhiyong Wu 0001
ICASSP5
2023 Fast-U2++: Fast and Accurate End-to-End Speech Recognition in Joint CTC/Attention Frames
abstract
Recently, the unified streaming and non-streaming two-pass (U2/U2++) end-to-end model for speech recognition has shown great performance in terms of streaming capability, accuracy and latency. In this paper, we present fast-U2++, an enhanced version of U2++ to further reduce partial latency. The core idea of fast-U2++ is to output partial results of the bottom layers in its encoder with a small chunk, while using a large chunk in the top layers of its encoder to compensate the performance degradation caused by the small chunk. More-over, we use knowledge distillation method to reduce the token emission latency. We present extensive experiments on Aishell-1 dataset. Experiments and ablation studies show that compared to U2++, fast-U2++ reduces model latency from 320ms to 80ms, and achieves a character error rate (CER) of 5.06% with a streaming setup.
Chengdong Liang, Xiao-Lei Zhang 0001, Di Wu 0061, Shengqiang Li, Xingchen Song, Zhendong Peng, Fuping Pan
ICASSP8
2023 TrimTail: Low-Latency Streaming ASR with Simple But Effective Spectrogram-Level Length Penalty
abstract
In this paper, we present TrimTail, a simple but effective emission regularization method to improve the latency of streaming ASR models. The core idea of TrimTail is to apply length penalty (i.e., by trimming trailing frames, see Fig. 1-(b)) directly on the spectrogram of input utterances, which does not require any alignment. We demonstrate that TrimTail is computationally cheap and can be applied online and optimized with any training loss or any model architecture on any dataset without any extra effort by applying it on various end-to-end streaming ASR networks either trained with CTC loss [1] or Transducer loss [2]. We achieve 100 ~ 200ms latency reduction with equal or even better accuracy on both Aishell-1 and Librispeech. Moreover, by using TrimTail, we can achieve a 400ms algorithmic improvement of User Sensitive Delay (USD) with an accuracy loss of less than 0.2.
Xingchen Song, Di Wu 0061, Zhiyong Wu 0001, Yuekai Zhang, Zhendong Peng, Wenpeng Li, Fuping Pan, Changbao Zhu
ICASSP8
2023 Wekws: A Production First Small-Footprint End-to-End Keyword Spotting Toolkit
abstract
Keyword spotting (KWS) enables speech-based user interaction and gradually becomes an indispensable component of smart devices. Recently, end-to-end (E2E) methods have be-come the most popular approach for on-device KWS tasks. However, there is still a gap between the research and deployment of E2E KWS methods. In this paper, we introduce WeKws, a production-quality, easy-to-build, and convenient-to-be-applied E2E KWS toolkit. WeKws contains the implementations of several state-of-the-art backbone networks, making it achieve highly competitive results on three publicly available datasets. To make WeKws a pure E2E toolkit, we utilize a refined max-pooling loss to make the model learn the ending position of the keyword by itself, which significantly simplifies the training pipeline and makes WeKws very efficient to be applied in real-world scenarios. The toolkit is publicly available at https://github.com/wenet-e2e/wekws.
Menglong Xu, Jingyong Hou, Xiao-Lei Zhang 0001, Lei Xie 0001, Fuping Pan
ICASSP7
2023 ZeroPrompt: Streaming Acoustic Encoders are Zero-Shot Masked LMs
abstract
In this paper, we present ZeroPrompt (Figure 1-(a)) and the corresponding Prompt-and-Refine strategy (Figure 3), two simple but effective training-free methods to decrease the Token Display Time (TDT) of streaming ASR models without any accuracy loss.The core idea of ZeroPrompt is to append zeroed content to each chunk during inference, which acts like a prompt to encourage the model to predict future tokens even before they were spoken.We argue that streaming acoustic encoders naturally have the modeling ability of Masked Language Models and our experiments demonstrate that ZeroPrompt is engineering cheap and can be applied to streaming acoustic encoders on any dataset without any accuracy loss.Specifically, compared with our baseline models, we achieve 350 ∼ 700ms reduction on First Token Display Time (TDT-F) and 100 ∼ 400ms reduction on Last Token Display Time (TDT-L), with theoretically and experimentally equal WER on both Aishell-1 and Librispeech datasets.
Xingchen Song, Di Wu 0061, Zhendong Peng, Bo Dang 0004, Fuping Pan, Zhiyong Wu 0001
INTERSPEECH6
2022 WeNet 2.0: More Productive End-to-End Speech Recognition Toolkit
abstract
Recently, we made available WeNet [1], a production-oriented end-to-end speech recognition toolkit, which introduces a unified two-pass (U2) framework and a built-in runtime to address the streaming and non-streaming decoding modes in a single model.To further improve ASR performance and facilitate various production requirements, in this paper, we present WeNet 2.0 with four important updates.(1) We propose U2++, a unified two-pass framework with bidirectional attention decoders, which includes the future contextual information by a right-toleft attention decoder to improve the representative ability of the shared encoder and the performance during the rescoring stage.(2) We introduce an n-gram based language model and a WFSTbased decoder into WeNet 2.0, promoting the use of rich text data in production scenarios.(3) We design a unified contextual biasing framework, which leverages user-specific context (e.g., contact lists) to provide rapid adaptation ability for production and improves ASR accuracy in both with-LM and without-LM scenarios.(4) We design a unified IO to support large-scale data for effective model training.In summary, the brand-new WeNet 2.0 achieves up to 10% relative recognition performance improvement over the original WeNet on various corpora and makes available several important production-oriented features.
Di Wu 0061, Zhendong Peng, Xingchen Song, Zhuoyuan Yao, Hang Lv 0001, Lei Xie 0001, Chao Yang 0031, Fuping Pan, Jianwei Niu 0002
INTERSPEECH9
2013 A novel discriminative method for pronunciation quality assessment
abstract
This paper presents a novel method for automatic pronunciation quality assessment. Unlike the traditional “Goodness of Pronunciation” (GOP) method, we judged utterance's pronunciation quality directly by a discriminative method. Under this novel framework, we also designed an algorithm to calculate the assessment confidence. We decoded the student's utterance for two passes. The first-pass decoding was just for getting the phone time points, and the second-pass decoding was for differentiating the pronunciation quality for each triphone. In the second-pass decoding, we used a specially trained acoustic model (AM), where the triphones in different pronunciation qualities were trained as different units. The confidence of the phone-level scoring was also calculated, and the low confidence phone-level scores were excluded in calculating the word-level score. The experimental results shows that the scoring performance was increased significantly compared to the traditional GOP method.
Fuping Pan, Bin Dong 0003, Yonghong Yan 0002
ICASSP2
2009 A one-step tone recognition approach using MSD-HMM for continuous speech
Changliang Liu, Fengpei Ge, Fuping Pan, Bin Dong 0003, Yonghong Yan 0002
INTERSPEECH3
2009 An SVM-Based Mandarin Pronunciation Quality Assessment System
Fengpei Ge, Fuping Pan, Changliang Liu, Bin Dong 0003, Shui-duen Chan, Yonghong Yan 0002
ISNN (4)2
2009 Dynamic Multiple Pronunciation Incorporation in a Refined Search Space for Reading Miscue Detection
Changliang Liu, Fuping Pan, Fengpei Ge, Bin Dong 0003, Shuiduen Chen, Yonghong Yan 0002
ISNN (4)2
2008 Mandarin vowel pronunciation quality evaluation by a novel formant classification method and its combination with traditional algorithms
abstract
This paper discusses the vowel pronunciation quality assessment of our computer assisted Mandarin Chinese learning system. Under the speech recognition framework, phonetic pronunciation assessment is usually based on the phonetic posterior probability score, which may be computed by normalizing the frame-based posterior probability or be calculated on the phone segment directly. By the first method, we can achieve a human-machine scoring correlation coefficient (CC) of 0.832 for vowel; and by the second, the CC can be up to 0.847. In order to improve the performance, we suggest employing the formant feature of vowel. This paper proposes a novel method to utilize formant: we plot formant candidates of each frame on the time-frequency plane to form a bitmap, and then extract its Gabor feature for pattern classification. When we use the classification probability score for pronunciation assessment, we get a CC of 0.842. Finally we combine the three scores with various linear or nonlinear methods; the best CC of 0.913 is gotten by using neural network.
Fuping Pan, Qingwei Zhao, Yonghong Yan 0002
ICASSP1
2008 Forward optimal modeling of acoustic confusions in Mandarin CALL system
Fengpei Ge, Fuping Pan, Changliang Liu, Bin Dong 0003, Yonghong Yan 0002
INTERSPEECH2
2007 Mandarin vowel pronunciation quality evaluation by using formant pattern recognition
Fuping Pan, Qingwei Zhao, Yonghong Yan 0002
INTERSPEECH1