Liyong Guo

dblp:129/3379 · DBLP profile ↗
← Back
22ranked-venue papers
2as first author
19since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 15 since 2021Artificial intelligence and machine learning · 10 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorComputer networks · 1 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2025 ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
abstract
Existing large-scale zero-shot text-to-speech (TTS) models deliver high speech quality but suffer from slow inference speeds due to massive parameters. To address this issue, this paper introduces ZipVoice, a high-quality flow-matching-based zero-shot TTS model with a compact model size and fast inference speed. Key designs include: 1) a Zipformer-based vector field estimator to maintain adequate modeling capabilities under constrained size; 2) Average upsampling-based initial speech-text alignment and Zipformer-based text encoder to improve speech intelligibility; 3) A flow distillation method to reduce sampling steps and eliminate the inference overhead associated with classifier-free guidance. Experiments on 100 k hours multilingual datasets show that ZipVoice matches state-of-the-art models in speech quality, while being 3 times smaller and up to 30 times faster than a DiT-based flow-matching baseline. Codes, model checkpoints and demo samples are publicly available.11https://github.com/k2-fsa/ZipVoice
Zhu Han 0001, Wei Kang 0006, Zengwei Yao, Liyong Guo, Zhaoqing Li, Weiji Zhuang, Long Lin, Daniel Povey
ASRU4
2025 CR-CTC: Consistency regularization on CTC for improved speech recognition
abstract
Connectionist Temporal Classification (CTC) is a widely used method for automatic speech recognition (ASR), renowned for its simplicity and computational efficiency. However, it often falls short in recognition performance. In this work, we propose the Consistency-Regularized CTC (CR-CTC), which enforces consistency between two CTC distributions obtained from different augmented views of the input speech mel-spectrogram. We provide in-depth insights into its essential behaviors from three perspectives: 1) it conducts self-distillation between random pairs of sub-models that process different augmented views; 2) it learns contextual representation through masked prediction for positions within time-masked regions, especially when we increase the amount of time masking; 3) it suppresses the extremely peaky CTC distributions, thereby reducing overfitting and improving the generalization ability. Extensive experiments on LibriSpeech, Aishell-1, and GigaSpeech datasets demonstrate the effectiveness of our CR-CTC. It significantly improves the CTC performance, achieving state-of-the-art results comparable to those attained by transducer or systems combining CTC and attention-based encoder-decoder (CTC/AED). We release our code at \url{https://github.com/k2-fsa/icefall}.
Zengwei Yao, Wei Kang 0006, Xiaoyu Yang 0005, Liyong Guo, Han Zhu 0004, Zengrui Jin, Zhaoqing Li, Long Lin, Daniel Povey
ICLR5
2025 k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning
abstract
Self-supervised learning (SSL) has achieved great success in speech-related tasks. While Transformer and Conformer architectures have dominated SSL backbones, encoders like Zipformer, which excel in automatic speech recognition (ASR), remain unexplored in SSL. Concurrently, inefficiencies in data processing within existing SSL training frameworks, such as fairseq, pose challenges in managing the growing volumes of training data. To address these issues, we propose k2SSL, an open-source framework that offers faster, more memory-efficient, and better-performing self-supervised speech representation learning, focusing on downstream ASR tasks. The optimized HuBERT and proposed Zipformer-based SSL systems exhibit substantial reductions in both training time and memory usage during SSL training. Experiments on LibriSpeech demonstrate that Zipformer Base significantly outperforms HuBERT and WavLM, achieving up to a 34.8% relative WER reduction compared to HuBERT Base after fine-tuning, along with a 3.5x pre-training speedup in GPU hours. When scaled to 60k hours of LibriLight data, Zipformer Large exhibits remarkable efficiency, matching HuBERT Large’s performance while requiring only 5/8 pre-training steps.
Yifan Yang 0005, Jianheng Zhuo, Zengrui Jin, Ziyang Ma 0001, Xiaoyu Yang 0005, Zengwei Yao, Liyong Guo, Wei Kang 0006, Long Lin, Daniel Povey, Xie Chen 0001
ICME7
2025 GLCLAP: A Novel Contrastive Learning Pre-trained Model for Contextual Biasing in ASR
Yuxiang Kong, Fan Cui, Liyong Guo, Heinrich Dinkel, Lichun Fan, Jian Luan 0001
INTERSPEECH3
2024 Libriheavy: A 50, 000 Hours ASR Corpus with Punctuation Casing and Context
abstract
In this paper, we introduce Libriheavy, a large-scale ASR corpus consisting of 50,000 hours of read English speech derived from LibriVox. To the best of our knowledge, Libriheavy is the largest freely-available corpus of speech with supervisions. Different from other open-sourced datasets that only provide normalized transcriptions, Libriheavy contains richer information such as punctuation, casing and text context, which brings more flexibility for system building. Specifically, we propose a general and efficient pipeline to locate, align and segment the audios in previously published Librilight to its corresponding texts. The same as Librilight, Libriheavy also has three training subsets small, medium, large of the sizes 500h, 5000h, 50000h respectively. We also extract the dev and test evaluation sets from the aligned audios and guarantee there is no overlapping speakers and books in training sets. Baseline systems are built on the popular CTC-Attention and transducer models. Additionally, we open-source our dataset creatation pipeline which can also be used to other audio alignment tasks.
Wei Kang 0006, Xiaoyu Yang 0005, Zengwei Yao, Yifan Yang 0005, Liyong Guo, Long Lin, Daniel Povey
ICASSP6
2024 PromptASR for Contextualized ASR with Controllable Style
abstract
Prompts are crucial to large language models as they provide context information such as topic or logical relationships. Inspired by this, we propose PromptASR, a framework that integrates prompts in end-to-end automatic speech recognition (E2E ASR) systems to achieve contextualized ASR with controllable style of transcriptions. Specifically, a dedicated text encoder encodes the text prompts and the encodings are injected into the speech encoder by cross-attending the features from two modalities. When using the ground truth text from preceding utterances as content prompt, the proposed system achieves 21.9% and 6.8% relative word error rate reductions on a book reading dataset and an in-house dataset compared to a baseline ASR system. The system can also take word-level biasing lists as prompt to improve recognition accuracy on rare words. An additional style prompt can be given to the text encoder and guide the ASR system to output different styles of transcriptions. The code is available at icefall1.
Xiaoyu Yang 0005, Wei Kang 0006, Zengwei Yao, Yifan Yang 0005, Liyong Guo, Long Lin, Daniel Povey
ICASSP5
2024 Zipformer: A faster and better encoder for automatic speech recognition
abstract
The Conformer has become the most popular encoder model for automatic speech recognition (ASR). It adds convolution modules to a transformer to learn both local and global dependencies. In this work we describe a faster, more memory-efficient, and better-performing transformer, called Zipformer. Modeling changes include: 1) a U-Net-like encoder structure where middle stacks operate at lower frame rates; 2) reorganized block structure with more modules, within which we re-use attention weights for efficiency; 3) a modified form of LayerNorm called BiasNorm allows us to retain some length information; 4) new activation functions SwooshR and SwooshL work better than Swish. We also propose a new optimizer, called ScaledAdam, which scales the update by each tensor's current scale to keep the relative change about the same, and also explictly learns the parameter scale. It achieves faster converge and better performance than Adam. Extensive experiments on LibriSpeech, Aishell-1, and WenetSpeech datasets demonstrate the effectiveness of our proposed Zipformer over other state-of-the-art ASR models. Our code is publicly available at https://github.com/k2-fsa/icefall.
Zengwei Yao, Liyong Guo, Xiaoyu Yang 0005, Wei Kang 0006, Yifan Yang 0005, Zengrui Jin, Long Lin, Daniel Povey
ICLR2
2024 LibriheavyMix: A 20, 000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
Zengrui Jin, Yifan Yang 0005, Mohan Shi, Wei Kang 0006, Xiaoyu Yang 0005, Zengwei Yao, Liyong Guo, Lingwei Meng, Long Lin, Yong Xu 0004, Shixiong Zhang 0001, Daniel Povey
INTERSPEECH8
2023 Relate Auditory Speech To Eeg By Shallow-Deep Attention-Based Network
abstract
Electroencephalography (EEG) plays a vital role in detecting how brain responses to different stimulus. In this paper, we propose a novel Shallow-Deep Attention-based Network (SDANet) to classify the correct auditory stimulus evoking the EEG signal. It adopts the Attention-based Correlation Module (ACM) to discover the connection between auditory speech and EEG from global aspect, and the Shallow-Deep Similarity Classification Module (SDSCM) to decide the classification result via the embeddings learned from the shallow and deep layers. Moreover, various training strategies and data augmentation are used to boost the model robustness. Experiments are conducted on the dataset provided by Auditory EEG challenge (ICASSP Signal Processing Grand Challenge 2023). Results show that the proposed model has a significant gain over the baseline on the match-mismatch track.
Fan Cui, Liyong Guo, Jiyao Liu, Ercheng Pei, Dongmei Jiang
ICASSP2
2023 Predicting Multi-Codebook Vector Quantization Indexes for Knowledge Distillation
abstract
Knowledge distillation (KD) is a common approach to improve model performance in automatic speech recognition (ASR), where a student model is trained to imitate the output behaviour of a teacher model. However, traditional KD methods suffer from teacher label storage issue, especially when the training corpora are large. Although on-the-fly teacher label generation tackles this issue, the training speed is significantly slower as the teacher model has to be evaluated every batch. In this paper, we reformulate the generation of teacher label as a codec problem. We propose a novel Multi-codebook Vector Quantization (MVQ) approach that compresses teacher embeddings to codebook indexes (CI). Based on this, a KD training framework (MVQ-KD) is proposed where a student model predicts the CI generated from the embeddings of a self-supervised pre-trained teacher model. Experiments on the LibriSpeech clean-100 hour show that MVQ-KD framework achieves comparable performance as traditional KD methods (11, 12), while requiring 256 times less storage. When the full LibriSpeech dataset is used, MVQ-KD framework results in 13.8% and 8.2% relative word error rate reductions (WERRs) for non -streaming transducer on test-clean and test-other and 4.0% and 4.9% for streaming transducer. The implementation of this work is already released as a part of the open-source project icefall1.
Liyong Guo, Xiaoyu Yang 0005, Quandong Wang, Yuxiang Kong, Zengwei Yao, Fan Cui, Wei Kang 0006, Long Lin, Mingshuang Luo, Piotr Zelasko, Daniel Povey
ICASSP1
2023 Fast and Parallel Decoding for Transducer
abstract
The transducer architecture is becoming increasingly popular in the field of speech recognition, because it is naturally streaming as well as high in accuracy. One of the drawbacks of transducer is that it is difficult to decode in a fast and parallel way due to an unconstrained number of symbols that can be emitted per time step.In this work, we introduce a constrained version of transducer loss to learn strictly monotonic alignments between the sequences; we also improve the standard greedy search and beam search algorithms by limiting the number of symbols that can be emitted per time step in transducer decoding, making it more efficient to decode in parallel with batches. Furthermore, we propose an finite state automaton-based (FSA) parallel beam search algorithm that can run with graphs on GPU efficiently. The experiment results show that we achieve slight word error rate (WER) improvement as well as significant speedup in decoding. Our work is open-sourced and publicly available1.
Wei Kang 0006, Liyong Guo, Long Lin, Mingshuang Luo, Zengwei Yao, Xiaoyu Yang 0005, Piotr Zelasko, Daniel Povey
ICASSP2
2023 Delay-Penalized Transducer for Low-Latency Streaming ASR
abstract
In streaming automatic speech recognition (ASR), it is desirable to reduce latency as much as possible while having minimum impact on recognition accuracy. Although a few existing methods are able to achieve this goal, they are difficult to implement due to their dependency on external alignments. In this paper, we propose a simple way to penalize symbol delay in transducer model, so that we can balance the trade-off between symbol delay and accuracy for streaming models without external alignments. Specifically, our method adds a small constant times (T/2 - t), where T is the number of frames and t is the current frame, to all the non-blank log-probabilities (after normalization) that are fed into the two dimensional transducer recursion. For both streaming Conformer models and unidirectional long short-term memory (LSTM) models, experimental results show that it can significantly reduce the symbol delay with an acceptable performance degradation. Our method achieves similar delay-accuracy trade-off to the previously published FastEmit, but we believe our method is preferable because it has a better justification: it is equivalent to penalizing the average symbol delay. Our work is open-sourced and publicly available1.
Wei Kang 0006, Zengwei Yao, Liyong Guo, Xiaoyu Yang 0005, Long Lin, Piotr Zelasko, Daniel Povey
ICASSP4
2023 Blank-regularized CTC for Frame Skipping in Neural Transducer
Yifan Yang 0005, Xiaoyu Yang 0005, Liyong Guo, Zengwei Yao, Wei Kang 0006, Long Lin, Xie Chen 0001, Daniel Povey
INTERSPEECH3
2023 Delay-penalized CTC Implemented Based on Finite State Transducer
Zengwei Yao, Wei Kang 0006, Liyong Guo, Xiaoyu Yang 0005, Yifan Yang 0005, Long Lin, Daniel Povey
INTERSPEECH4
2022 Multi-Scale Refinement Network Based Acoustic Echo Cancellation
abstract
Recently, deep encoder-decoder networks have shown outstanding performance in acoustic echo cancellation (AEC). However, the subsampling operations like convolution striding in the encoder layers significantly decrease the feature resolution lead to fine-grained information loss. This paper proposes an encoder-decoder network for acoustic echo cancellation with mutli-scale refinement paths to exploit the information at different feature scales. In the encoder stage, high-level features are obtained to get a coarse result. Then, the decoder layers with multiple refinement paths can directly refine the result with fine-grained features. Refinement paths with different feature scales are combined by learnable weights. The experimental results show that using the proposed multi-scale refinement structure can significantly improve the objective criteria. In the ICASSP 2022 Acoustic echo cancellation Challenge, our submitted system achieves an overall MOS score of 4.439 with 4.37 million parameters at a system latency of 40ms.
Fan Cui, Liyong Guo, Peng Gao 0013
ICASSP2
2022 Exploring representation learning for small-footprint keyword spotting
abstract
In this paper, we investigate representation learning for low-resource keyword spotting (KWS). The main challenges of KWS are limited labeled data and limited available device resources. To address those challenges, we explore representation learning for KWS by self-supervised contrastive learning and self-training with pretrained model. First, local-global contrastive siamese networks (LGCSiam) are designed to learn similar utterance-level representations for similar audio samplers by proposed local-global contrastive loss without requiring ground-truth. Second, a self-supervised pretrained Wav2Vec 2.0 model is applied as a constraint module (WVC) to force the KWS model to learn frame-level acoustic representations. By the LGCSiam and WVC modules, the proposed small-footprint KWS model can be pretrained with unlabeled data. Experiments on speech commands dataset show that the self-training WVC module and the self-supervised LGCSiam module significantly improve accuracy, especially in the case of training on a small labeled dataset.
Fan Cui, Liyong Guo, Quandong Wang, Peng Gao 0013
INTERSPEECH2
2022 Pruned RNN-T for fast, memory-efficient ASR training
abstract
The RNN-Transducer (RNN-T) framework for speech recognition has been growing in popularity, particularly for deployed real-time ASR systems, because it combines high accuracy with naturally streaming recognition.One of the drawbacks of RNN-T is that its loss function is relatively slow to compute, and can use a lot of memory.Excessive GPU memory usage can make it impractical to use RNN-T loss in cases where the vocabulary size is large: for example, for Chinese character-based ASR.We introduce a method for faster and more memoryefficient RNN-T loss computation.We first obtain pruning bounds for the RNN-T recursion using a simple joiner network that is linear in the encoder and decoder embeddings; we can evaluate this without using much memory.We then use those pruning bounds to evaluate the full, non-linear joiner network.The code is open-sourced and publicly available.
Liyong Guo, Wei Kang 0006, Long Lin, Mingshuang Luo, Zengwei Yao, Daniel Povey
INTERSPEECH2
2022 Reducing noisy annotations for depression estimation from facial images
abstract
Depression has been considered the most dominant mental disorder over the past few years. To help clinicians effectively and efficiently estimate the severity scale of depression, various automated systems based on deep learning have been proposed. To estimate the severity of depression, i.e., the depression severity score (Beck Depression Inventory-II), various deep architectures have been designed to perform regression using the Euclidean loss. However, they do not consider the label distribution, and they do not learn the relationships between the facial images and BDI-II scores, which can be resulting in the noisy labeling for automatic depression estimation (ADE). To mitigate this problem, we propose an automated deep architecture, namely the self-adaptation network (SAN), to improve this uncertain labeling for ADE. Specifically, the architecture consists of four modules: (1) ResNet-18 and ResNet-50 are adopted in the deep feature extraction module (DFEM) to extract informative deep features; (2) a self-attention module (SAM) is adopted to learn the weights from the mini-batch; (3) a square ranking regularization module (SRRM) to create high partitions and low partitions is proposed; and (4) a re-label module (RM) is used to re-label the uncertain annotations for ADE in the low partitions. We conduct extensive experiments on depression databases (i.e., AVEC2013 and AVEC2014) and obtain a performance comparable to the performances of other ADE methods in assessing the severity of depression. More importantly, the proposed method can learn valuable depression patterns from facial videos and obtain a performance comparable to the performances of other methods for depression recognition.
Prayag Tiwari, Chonghua Lv, Wenshuai Wu, Liyong Guo
Neural Networks5
2021 Performance Analysis of Delay Distribution and Packet Loss Ratio for Body-to-Body Networks
abstract
With the increasing wide applications of wearable wireless networks, body-to-body networks (BBNs) have become significantly important to provide timely and reliable data delivery services. For a specific BBN, assessing its theoretically achievable Quality of Service (QoS) is necessary, especially on the key performance metrics of end-to-end delay distribution and packet loss ratio. The existing analysis models in the literature mainly focused on 1-D space scenarios. In this article, BBN in a 2-D area is considered, where mobile nodes freely and stochastically move along lanes. By introducing two new definitions: 1) node entrance probability and 2) network entrance probability, a systematically analytical framework for end-to-end delay distribution and packet loss ratio is presented. The proposed analytical framework is built on three critical techniques: 1) the Markov chain to model node behaviors; 2) the first passage theory to calculate node entrance probability and network entrance probability; and 3) the central limit theory to decrease the computation time for summing up per-hop delay. Simulation results demonstrate the effectiveness and accuracy of our proposed analysis model.
Xiaolong Li 0004, Jun Cai 0001, Liyong Guo, Shaonian Huang, Yunfei Yi
IEEE Internet Things J.4
2019 A Novel High-Accuracy Phase-Derived Velocity Measurement Method for Wideband LFM Radar
abstract
A novel high-accuracy phase-derived velocity measurement (PDVM) method for fast-moving space targets is presented in this letter. First, a wideband linear frequency-modulated signal model that considers the effect of radial acceleration was developed. To obtain the unambiguous phase difference between two adjacent echo pulses, coarse velocity and acceleration measurements derived from range profile cross correlation were used to resolve the phase ambiguity. Then, the derived accurate and unambiguous phase difference and the phase error induced by the discrete Fourier transform were analyzed. The PDVM technique was applied after the phase error was compensated for, and all required parameters were calculated. Under low signal-to-noise ratio (SNR) conditions, a correction for the phase unwrapping error was developed. The simulation results showed that the proposed PDVM technique was highly accurate. The root-mean-square error of the PDVM results was less than 0.025 m/s when the SNR was greater than 15 dB.
Liyong Guo, Huayu Fan, Quanhua Liu 0002, Xiaopeng Yang 0002
IEEE Geosci. Remote. Sens. Lett.1
2018 Attribute-Based Traceable Anonymous Proxy Signature Strategy for Mobile Healthcare
Guojun Wang 0001, Quanyou Zhao, Liyong Guo
ISPEC6
2013 Specificity and affinity quantification of protein-protein interactions
abstract
MOTIVATION: Most biological processes are mediated by the protein-protein interactions. Determination of the protein-protein structures and insight into their interactions are vital to understand the mechanisms of protein functions. Currently, compared with the isolated protein structures, only a small fraction of protein-protein structures are experimentally solved. Therefore, the computational docking methods play an increasing role in predicting the structures and interactions of protein-protein complexes. The scoring function of protein-protein interactions is the key responsible for the accuracy of the computational docking. Previous scoring functions were mostly developed by optimizing the binding affinity which determines the stability of the protein-protein complex, but they are often lack of the consideration of specificity which determines the discrimination of native protein-protein complex against competitive ones. RESULTS: We developed a scoring function (named as SPA-PP, specificity and affinity of the protein-protein interactions) by incorporating both the specificity and affinity into the optimization strategy. The testing results and comparisons with other scoring functions show that SPA-PP performs remarkably on both predictions of binding pose and binding affinity. Thus, SPA-PP is a promising quantification of protein-protein interactions, which can be implemented into the protein docking tools and applied for the predictions of protein-protein structure and affinity. AVAILABILITY: The algorithm is implemented in C language, and the code can be downloaded from http://dl.dropbox.com/u/1865642/Optimization.cpp.
Liyong Guo
Bioinform.2