Wei Liu 0147

dblp:49/3283-147 · DBLP profile ↗
← Back
10ranked-venue papers
7as first author
10since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021
YearPublicationVenuePosition
2024 Sparsely Shared Lora on Whisper for Child Speech Recognition
abstract
Whisper is a powerful automatic speech recognition (ASR) model. Nevertheless, its zero-shot performance on low-resource speech requires further improvement. Child speech, as a representative type of low-resource speech, is leveraged for adaptation. Recently, parameter-efficient fine-tuning (PEFT) in NLP was shown to be comparable and even better than full fine-tuning, while only needing to tune a small set of trainable parameters. However, current PEFT methods have not been well examined for their effectiveness on Whisper. In this paper, only parameter composition types of PEFT approaches such as LoRA and Bitfit are investigated as they do not bring extra inference costs. Different popular PEFT methods are examined. Particularly, we compare LoRA and AdaLoRA and figure out the learnable rank coefficient is a good design. Inspired by the sparse rank distribution allocated by AdaLoRA, a novel PEFT approach Sparsely Shared LoRA (S2-LoRA) is proposed. The two low-rank decomposed matrices are globally shared. Each weight matrix only has to maintain its specific rank coefficients that are constrained to be sparse. Experiments on low-resource Chinese child speech show that with much fewer trainable parameters, S2-LoRA can achieve comparable in-domain adaptation performance to AdaLoRA and exhibit better generalization ability on out-of-domain data. In addition, the rank distribution automatically learned by S2-LoRA is found to have similar patterns to AdaLoRA’s allocation.
Wei Liu 0147, Tan Lee
ICASSP1
2024 A Parameter-efficient Language Extension Framework for Multilingual ASR
Wei Liu 0147, Jingyong Hou, Muyong Cao, Tan Lee
INTERSPEECH1
2024 LUPET: Incorporating Hierarchical Information Path into Multilingual ASR
abstract
Toward high-performance multilingual automatic speech recognition (ASR), various types of linguistic information and model design have demonstrated their effectiveness independently.They include language identity (LID), phoneme information, language-specific processing modules, and crosslingual self-supervised speech representation.It is expected that leveraging their benefits synergistically in a unified solution would further improve the overall system performance.This paper presents a novel design of a hierarchical information path, named LUPET, which sequentially encodes, from the shallow layers to deep layers, multiple aspects of linguistic and acoustic information at diverse granularity scales.The path starts from LID prediction, followed by acoustic unit discovery, phoneme sharing, and finally token recognition routed by a mixture-ofexpert.ASR experiments are carried out on 10 languages in the Common Voice corpus.The results demonstrate the superior performance of LUPET as compared to the baseline systems.Most importantly, LUPET effectively mitigates the issue of performance compromise of high-resource languages with low-resource ones in the multilingual setting.
Wei Liu 0147, Jingyong Hou, Muyong Cao, Tan Lee
INTERSPEECH1
2023 Diffusion-Based Mel-Spectrogram Enhancement for Personalized Speech Synthesis with Found Data
abstract
Creating synthetic voices with found data is challenging, as real-world recordings often contain various types of audio degradation. One way to address this problem is to pre-enhance the speech with an enhancement model and then use the enhanced data for text-to-speech (TTS) model training. This paper investigates the use of conditional diffusion models for generalized speech enhancement, which aims at addressing multiple types of audio degradation simultaneously. The enhancement is performed on the log Mel-spectrogram domain to align with the TTS training objective. Text information is introduced as an additional condition to improve the model robustness. Experiments on real-world recordings demonstrate that the synthetic voice built on data enhanced by the proposed model produces higher-quality synthetic speech, compared to those trained on data enhanced by strong baselines. Code and check-points of the proposed enhancement model are available at https://github.com/dmse4tts/DMSE4TTS.
Yusheng Tian, Wei Liu 0147, Tan Lee
ASRU2
2023 Leveraging Phone-Level Linguistic-Acoustic Similarity For Utterance-Level Pronunciation Scoring
abstract
Recent studies on pronunciation scoring have explored the effect of introducing phone embeddings as reference pronunciation, but mostly in an implicit manner, i.e., addition or concatenation of reference phone embedding and actual pronunciation of the target phone as the phone-level pronunciation quality representation. In this paper, we propose to use linguistic-acoustic similarity to explicitly measure the deviation of non-native production from its native reference for pronunciation assessment. Specifically, the deviation is first estimated by the cosine similarity between reference phone embedding and corresponding acoustic embedding. Next, a phone-level Goodness of pronunciation (GOP) pre-training stage is introduced to guide this similarity-based learning for better initialization of the aforementioned two embeddings. Finally, a transformer-based hierarchical pronunciation scorer is used to map a sequence of phone embeddings, acoustic embeddings along with their similarity measures to predict the final utterance-level score. Experimental results on the non-native databases suggest that the proposed system significantly outperforms the baselines, where the acoustic and phone embeddings are simply added or concatenated. A further examination shows that the phone embeddings learned in the proposed approach are able to capture linguistic-acoustic attributes of native pronunciation as references.
Wei Liu 0147, Kaiqi Fu, Xiaohai Tian, Shuju Shi, Wei Li 0119, Zejun Ma 0001, Tan Lee
ICASSP1
2023 An ASR-Free Fluency Scoring Approach with Self-Supervised Learning
abstract
A typical fluency scoring system generally relies on an automatic speech recognition (ASR) system to obtain time stamps in input speech for the subsequent calculation of fluency-related features or directly modeling speech fluency with an end-to-end approach. This paper describes a novel ASR-free approach for automatic fluency assessment using self-supervised learning (SSL). Specifically, wav2vec2.0 is used to extract frame-level speech features, followed by K-means clustering to assign a pseudo label (cluster index) to each frame. A BLSTM-based model is trained to predict an utterance-level fluency score from frame-level SSL features and the corresponding cluster indexes. Neither speech transcription nor time stamp information is required in the proposed system. It is ASR-free and can potentially avoid the ASR errors effect in practice. Experimental results carried out on non-native English databases show that the proposed approach significantly improves the performance in the "open response" scenario as compared to previous methods and matches the recently reported performance in the "read aloud" scenario.
Wei Liu 0147, Kaiqi Fu, Xiaohai Tian, Shuju Shi, Wei Li 0119, Zejun Ma 0001, Tan Lee
ICASSP1
2023 Model Compression for DNN-based Speaker Verification Using Weight Quantization
Wei Liu 0147, Zhaoyang Zhang 0001, Tan Lee
INTERSPEECH2
2023 CoMFLP: Correlation Measure Based Fast Search on ASR Layer Pruning
abstract
Transformer-based speech recognition (ASR) model with deep layers exhibited significant performance improvement.However, the model is inefficient for deployment on resourceconstrained devices.Layer pruning (LP) is a commonly used compression method to remove redundant layers.Previous studies on LP usually identify the redundant layers according to a task-specific evaluation metric.They are time-consuming for models with a large number of layers, even in a greedy search manner.To address this problem, we propose CoM-FLP, a fast search LP algorithm based on correlation measure.The correlation between layers is computed to generate a correlation matrix, which identifies the redundancy among layers.The search process is carried out in two steps: (1) coarse search: to determine top K candidates by pruning the most redundant layers based on the correlation matrix; (2) fine search: to select the best pruning proposal among K candidates using a task-specific evaluation metric.Experiments on an ASR task show that the pruning proposal determined by CoMFLP outperforms existing LP methods while only requiring constant time complexity.The code is publicly available at https://github.com/louislau1129/CoMFLP.
Wei Liu 0147, Tan Lee
INTERSPEECH1
2022 EDITnet: A Lightweight Network for Unsupervised Domain Adaptation in Speaker Verification
abstract
Performance degradation caused by language mismatch is a common problem when applying a speaker verification system on speech data in different languages.This paper proposes a domain transfer network, named EDITnet, to alleviate the language-mismatch problem on speaker embeddings without requiring speaker labels.The network leverages a conditional variational auto-encoder to transfer embeddings from the target domain into the source domain.A self-supervised learning strategy is imposed on the transferred embeddings so as to increase the cosine distance between embeddings from different speakers.In the training process of the EDITnet, the embedding extraction model is fixed without fine-tuning, which renders the training efficient and low-cost.Experiments on Voxceleb and CN-Celeb show that the embeddings transferred by ED-ITnet outperform the un-transferred ones by around 30% with the ECAPA-TDNN512.Performance improvement can also be achieved with other embedding extraction models, e.g., TDNN, SE-ResNet34.
Wei Liu 0147, Tan Lee
INTERSPEECH2
2021 Utterance-Level Neural Confidence Measure for End-to-End Children Speech Recognition
abstract
Confidence measure is a performance index of particular importance for automatic speech recognition (ASR) systems deployed in real-world scenarios. In the present study, utterance-level neural confidence measure (NCM) in end-to-end automatic speech recognition (E2E ASR) is investigated. The E2E system adopts the joint CTC-attention Transformer architecture. The prediction of NCM is formulated as a task of binary classification, i.e., accept/reject the input utterance, based on a set of predictor features acquired during the ASR decoding process. The investigation is focused on evaluating and comparing the efficacies of predictor features that are derived from different internal and external modules of the E2E system. Experiments are carried out on children speech, for which state-of-the-art ASR systems show less than satisfactory performance and robust confidence measure is particularly useful. It is noted that predictor features related to acoustic information of speech play a more important role in estimating confidence measure than those related to linguistic information. N-best score features show significantly better performance than single-best ones. It has also been shown that the metrics of EER and AUC are not appropriate to evaluate the NCM of a mismatched ASR with significant performance gap.
Wei Liu 0147, Tan Lee
ASRU1