Yulan Liu

dblp:03/7673 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 3 since 2021Theory of computation · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Towards Accurate Identification of Anti-Hepatitis C Peptides Using Stack-AHCP
abstract
Hepatitis C virus (HCV) infection remains a significant global health burden, contributing to progressive hepatic pathologies including chronic hepatitis, cirrhosis, and hepatocellular carcinoma. While anti-hepatitis C peptides (AHCPs) have emerged as promising therapeutic candidates with distinct antiviral mechanisms, conventional wet-lab approaches for AHCP discovery face critical limitations in throughput and scalability. To overcome these constraints, we present Stack-Ahcp, an innovative stacked ensemble learning framework that synergistically integrates multiple machine learning algorithms through a meta-classification strategy. Our model achieves unprecedented predictive performance with 93.1% accuracy and an MCC of 0.863, substantially outperforming existing computational methods. Through comprehensive Shapley additive explanations (SHAP) analysis, we further delineate critical key intrinsic features and determinants governing AHCP bioactivity, enhancing the mechanistic interpretability of the prediction system. To facilitate translational applications, we have implemented an intuitive web interface (accessible at https://awi.cuhk.edu.cn/~biosequence/StackAHCP/index.php) that enables rapid screening and prioritization of candidate peptides. This resource is anticipated to streamline the identification of next-generation peptide therapeutics against HCV while reducing experimental validation costs. Beyond virology applications, our methodological framework establishes a paradigm for interpretable machine learning in biological sequence analysis, with potential adaptability to diverse multi-omics investigation scenarios.
Lantian Yao, Yen-Peng Chiu, Jiahui Guan, Peilin Xie, Yulan Liu, Yunlu Peng, Ying-Chih Chiang, Tzong-Yi Lee
CIBCB9
2025 Toward high-efficiency, low-resource, and explainable neuropeptide prediction with MSKDNP
abstract
Neuropeptides are essential signaling molecules produced in the nervous system that regulate diverse physiological processes and are closely implicated in the pathogenesis of neurodegenerative and neuropsychiatric disorders. Investigating neuropeptides contributes to a better understanding of their regulatory mechanisms and offers new insights into therapeutic strategies for related diseases. Therefore, accurate identification of neuropeptides is crucial for advancing biomedical research and drug development. Due to the high cost of experimental validation, various artificial intelligence methods have been developed for rapid neuropeptide identification. However, existing approaches often suffer from high computational resource consumption, slow processing speed, and poor deploy ability. Moreover, a user-friendly web server for practical application is still lacking. To this end, we propose MSKDNP, a neuropeptide prediction model based on a multi-stage knowledge distillation framework. With only 1.2% of the parameters, MSKDNP attains performance comparable to a fully fine-tuned protein language model while achieving state-of-the-art results in neuropeptide recognition. Moreover, MSKDNP provides favorable interpretability, facilitating biological understanding. A freely accessible web server is available at https://awi.cuhk.edu.cn/∼biosequence/MSKDNP/index.php.
Peilin Xie, Jiahui Guan, Yulan Liu, Zhang Cheng, Xuxin He, Zhenglong Sun 0001, Tzong-Yi Lee, Lantian Yao, Ying-Chih Chiang
Briefings Bioinform.4
2024 Kernel support vector machine classifiers with ℓ0-norm hinge loss
Rongrong Lin, Yingjia Yao, Yulan Liu
Neurocomputing3
2024 The Proximal Operator of the Piece-Wise Exponential Function
abstract
This paper characterizes the proximal operator of the piece-wise exponential function$1\!-\!e^{-|x|/\sigma }$with a given shape parameter$\sigma \!>\!0$, which is a popular non-convex surrogate of the$\ell _{0}$-norm in support vector machines, compressed sensing, neural networks, etc. Although Malek-Mohammadi et al. [IEEE Transactions on Signal Processing, 64(21):5657–5671, 2016] once worked on this problem, the expressions they derived were regrettably inaccurate. In a sense, it was lacking a case. Using the Lambert W function and an extensive study of the piece-wise exponential function, we have rectified the formulation of the proximal operator of the piece-wise exponential function in light of their work. We have also undertaken a thorough analysis of this operator. A comparative analysis of eleven sparse-promoting functions in compressed sensing demonstrates the effectiveness and efficiency of the piece-wise exponential function.
Yulan Liu, Rongrong Lin
IEEE Signal Process. Lett.1
2023 Calmness of partial perturbation to composite rank constraint systems and its applications
Yitian Qian, Shaohua Pan 0001, Yulan Liu
J. Glob. Optim.3
2022 Multi-Turn RNN-T for Streaming Recognition of Multi-Party Speech
abstract
Automatic speech recognition (ASR) of single channel far-field recordings with an unknown number of speakers is traditionally tackled by cascaded modules. Recent research shows that end-to-end (E2E) multi-speaker ASR models can achieve superior recognition accuracy compared to modular systems. However, these models do not ensure real-time applicability due to their dependency on full audio context. This work takes real-time applicability as the first priority in model design and addresses a few challenges in previous work on multi-speaker recurrent neural network transducer (MS-RNN-T). First, we introduce on-the-fly overlapping speech simulation during training, yielding 14% relative word error rate (WER) improvement on LibriSpeechMix test set. Second, we propose a novel multi-turn RNN-T (MT-RNN-T) model with an overlap-based target arrangement strategy that generalizes to an arbitrary number of speakers without changes in the model architecture. We investigate the impact of the maximum number of speakers seen during training on MT-RNN-T performance on LibriCSS test set, and report 28% relative WER improvement over the two-speaker MS-RNN-T. Third, we experiment with a rich transcription strategy for joint recognition and segmentation of multi-party speech. Through an in-depth analysis, we discuss potential pitfalls of the proposed system as well as promising future research directions.
Ilya Sklyar, Anna Piunova, Xianrui Zheng, Yulan Liu
ICASSP4
2021 Streaming Multi-Speaker ASR with RNN-T
abstract
Recent research shows end-to-end ASR systems can recognize overlapped speech from multiple speakers. However, all published works have assumed no latency constraints during inference, which does not hold for most voice assistant interactions. This work focuses on multi-speaker speech recognition based on a recurrent neural network transducer (RNN-T) that has been shown to provide high recognition accuracy at a low latency online recognition regime. We investigate two approaches to multi-speaker model training of the RNN-T: deterministic output-target assignment and permutation invariant training. We show that guiding separation with speaker order labels in the former case enhances the high-level speaker tracking capability of RNN-T. Apart from that, with multi-style training on single- and multi-speaker utterances, the resulting models gain robustness against ambiguous numbers of speakers during inference. Our best model achieves a WER of 10.2% on simulated 2-speaker LibriSpeech data, which is competitive with the previously reported state-of-the-art non-streaming model (10.3%), while the proposed model could be directly applied for streaming applications.
Ilya Sklyar, Anna Piunova, Yulan Liu
ICASSP3
2021 Using Synthetic Audio to Improve the Recognition of Out-of-Vocabulary Words in End-to-End Asr Systems
abstract
Today, many state-of-the-art automatic speech recognition (ASR) systems apply all-neural models that map audio to word sequences trained end-to-end along one global optimisation criterion in a fully data driven fashion. These models allow high precision ASR for domains and words represented in the training material but have difficulties recognising words that are rarely or not at all represented during training, i.e. trending words and new named entities. In this paper, we use a text-to-speech (TTS) engine to provide synthetic audio for out-of-vocabulary (OOV) words. We aim to boost the recognition accuracy of a recurrent neural network transducer (RNN-T) on OOV words by using the extra audio-text pairs, while maintaining the performance on the non-OOV words. Different regularisation techniques are explored and the best performance is achieved by fine-tuning the RNN-T on both original training data and extra synthetic data with elastic weight consolidation (EWC) applied on the encoder. This yields a 57% relative word error rate (WER) reduction on utterances containing OOV words without any degradation on the whole test set.
Xianrui Zheng, Yulan Liu, Deniz Gunceler, Daniel Willett
ICASSP2
2021 SynthASR: Unlocking Synthetic Data for Speech Recognition
abstract
End-to-end (E2E) automatic speech recognition (ASR) models have recently demonstrated superior performance over the traditional hybrid ASR models. Training an E2E ASR model requires a large amount of data which is not only expensive but may also raise dependency on production data. At the same time, synthetic speech generated by the state-of-the-art text-to-speech (TTS) engines has advanced to near-human naturalness. In this work, we propose to utilize synthetic speech for ASR training (SynthASR) in applications where data is sparse or hard to get for ASR model training. In addition, we apply continual learning with a novel multi-stage training strategy to address catastrophic forgetting, achieved by a mix of weighted multi-style training, data augmentation, encoder freezing, and parameter regularization. In our experiments conducted on in-house datasets for a new application of recognizing medication names, training ASR RNN-T models with synthetic audio via the proposed multi-stage training improved the recognition performance on new application by more than 65% relative, without degradation on existing general applications. Our observations show that SynthASR holds great promise in training the state-of-the-art large-scale E2E ASR models for new applications while reducing the costs and dependency on production data.
Amin Fazel, Yulan Liu, Roberto Barra-Chicote, Yixiong Meng, Roland Maas, Jasha Droppo
Interspeech3
2021 Bootstrap an End-to-End ASR System by Multilingual Training, Transfer Learning, Text-to-Text Mapping and Synthetic Audio
abstract
Bootstrapping speech recognition on limited data resources has been an area of active research for long. The recent transition to all-neural models and end-to-end (E2E) training brought along particular challenges as these models are known to be data hungry, but also came with opportunities around language-agnostic representations derived from multilingual data as well as shared word-piece output representations across languages that share script and roots. We investigate here the effectiveness of different strategies to bootstrap an RNN-Transducer (RNN-T) based automatic speech recognition (ASR) system in the low resource regime, while exploiting the abundant resources available in other languages as well as the synthetic audio from a text-to-speech (TTS) engine. Our experiments demonstrate that transfer learning from a multilingual model, using a post-ASR text-to-text mapping and synthetic audio deliver additive improvements, allowing us to bootstrap a model for a new language with a fraction of the data that would otherwise be needed. The best system achieved a 46% relative word error rate (WER) reduction compared to the monolingual baseline, among which 25% relative WER improvement is attributed to the post-ASR text-to-text mappings and the TTS synthetic data.
Manuel Giollo, Deniz Gunceler, Yulan Liu, Daniel Willett
Interspeech3
2018 Equivalent Lipschitz surrogates for zero-norm and rank optimization problems
Yulan Liu, Shujun Bi, Shaohua Pan 0001
J. Glob. Optim.1
2016 Image edge extraction based on fuzzy theory and Sobel operator
abstract
Image edge detection is a hot topic in the study of image processing and computer vision. Accurate edge information is helpful to the study of high-level image features. However, the image edge information is vague and uncertain. Fuzzy theory is expressed with the imprecise knowledge, and it is one of the effective methods to solve the problem with fuzzy phenomenon. The fuzzy theory and Sobel operator are combined in this research, and the main idea of the method is summarized as three steps, firstly, the image's gray matrix is converted into fuzzy feature matrix. Secondly, edge information is extruded by fuzzy enhancement. Finally, the image edge is detected by Sobel operator. The method of extracting image's edge with fuzzy theory has low computational complexity and simple procedure and precise edge information.
Yulan Liu
CSCWD1
2016 webASR 2 - Improved Cloud Based Speech Technology
abstract
This paper presents the most recent developments of the \nwebASR service (www.webasr.org), the world’s first web– \nbased fully functioning automatic speech recognition platform \nfor scientific use. Initially released in 2008, the functionalities \nof webASR have recently been expanded with 3 main goals in \nmind: Facilitate access through a RESTful architecture, that allows \nfor easy use through either the web interface or an API; allow \nthe use of input metadata when available by the user to improve \nsystem performance; and increase the coverage of available \nsystems beyond speech recognition. Several new systems \nfor transcription, diarisation, lightly supervised alignment and \ntranslation are currently available through webASR. The results \nin a series of well–known benchmarks (RT’09, IWSLT’12 and \nMGB’15 evaluations) show how these webASR systems provides \nstate–of–the–art performances across these tasks
Thomas Hain, Jeremy Christian, Oscar Saz-Torralba, Salil Deena, Madina Hasan, Raymond W. M. Ng, Rosanna Milner, Mortaza Doulaty, Yulan Liu
INTERSPEECH9
2016 The Sheffield Wargame Corpus - Day Two and Day Three
abstract
Improving the performance of distant speech recognition is of considerable current interest, driven by a desire to bring speech recognition into people’s homes. Standard approaches to this task aim to enhance the signal prior to recognition, typically using beamforming techniques on multiple channels. Only few real-world recordings are available that allow experimentation with such techniques. This has become even more pertinent with recent works with deep neural networks aiming to learn beamforming from data. Such approaches require large multi-channel training sets, ideally with location annotation for moving speakers, which is scarce in existing corpora. This paper presents a freely available and new extended corpus of English speech recordings in a natural setting, with moving speakers. The data is recorded with diverse microphone arrays, and uniquely, with ground truth location tracking. It extends the 8.0 hour Sheffield Wargames Corpus released in Interspeech 2013, with a further 16.6 hours of fully annotated data, including 6.1 hours of female speech to improve gender bias. Additional blog-based language model data is provided alongside, as well as a Kaldi baseline system. Results are reported with a standard Kaldi configuration, and a baseline meeting recognition system.
Yulan Liu, Charles Fox, Madina Hasan, Thomas Hain
INTERSPEECH1
2015 The 2015 sheffield system for transcription of Multi-Genre Broadcast media
abstract
We describe the University of Sheffield system for participation in the 2015 Multi-Genre Broadcast (MGB) challenge task of transcribing multi-genre broadcast shows. Transcription was one of four tasks proposed in the MGB challenge, with the aim of advancing the state of the art of automatic speech recognition, speaker diarisation and automatic alignment of subtitles for broadcast media. Four topics are investigated in this work: Data selection techniques for training with unreliable data, automatic speech segmentation of broadcast media shows, acoustic modelling and adaptation in highly variable environments, and language modelling of multi-genre shows. The final system operates in multiple passes, using an initial unadapted decoding stage to refine segmentation, followed by three adapted passes: a hybrid DNN pass with input features normalised by speaker-based cepstral normalisation, another hybrid stage with input features normalised by speaker feature-MLLR transformations, and finally a bottleneck-based tandem stage with noise and speaker factorisation. The combination of these three system outputs provides a final error rate of 27.5% on the official development set, consisting of 47 multi-genre shows.
Oscar Saz-Torralba, Mortaza Doulaty, Salil Deena, Rosanna Milner, Raymond W. M. Ng, Madina Hasan, Yulan Liu, Thomas Hain
ASRU7
2015 An investigation into speaker informed DNN front-end for LVCSR
abstract
Deep Neural Network (DNN) has become a standard method in many ASR tasks. Recently there is considerable interest in “informed training” of DNNs, where DNN input is augmented with auxiliary codes, such as i-vectors, speaker codes, speaker separation bottleneck (SSBN) features, etc. This paper compares different speaker informed DNN training methods in LVCSR task. We discuss mathematical equivalence between speaker informed DNN training and “bias adaptation” which uses speaker dependent biases, and give detailed analysis on influential factors such as dimension, discrimination and stability of auxiliary codes. The analysis is supported by experiments on a meeting recognition task using bottleneck feature based system. Results show that i-vector based adaptation is also effective in bottleneck feature based system (not just hybrid systems). However all tested methods show poor generalisation to unseen speakers. We introduce a system based on speaker classification followed by speaker adaptation of biases, which yields equivalent performance to an i-vector based system with 10.4% relative improvement over baseline on seen speakers. The new approach can serve as a fast alternative especially for short utterances.
Yulan Liu, Panagiota Karanasou, Thomas Hain
ICASSP1
2014 Using neural network front-ends on far field multiple microphones based speech recognition
abstract
This paper presents an investigation of far field speech recognition using beamforming and channel concatenation in the context of Deep Neural Network (DNN) based feature extraction. While speech enhancement with beamforming is attractive, the algorithms are typically signal-based with no information about the special properties of speech. A simple alternative to beamforming is concatenating multiple channel features. Results presented in this paper indicate that channel concatenation gives similar or better results. On average the DNN front-end yields a 25% relative reduction in Word Error Rate (WER). Further experiments aim at including relevant information in training adapted DNN features. Augmenting the standard DNN input with the bottleneck feature from a Speaker Aware Deep Neural Network (SADNN) shows a general advantage over the standard DNN based recognition system, and yields additional improvements for far field speech recognition.
Yulan Liu, Pengyuan Zhang, Thomas Hain
ICASSP1
2014 Semi-supervised DNN training in meeting recognition
abstract
Training acoustic models for ASR requires large amounts of labelled data which is costly to obtain. Hence it is desirable to make use of unlabelled data. While unsupervised training can give gains for standard HMM training, it is more difficult to make use of unlabelled data for discriminative models. This paper explores semi-supervised training of Deep Neural Networks (DNN) in a meeting recognition task. We first analyse the impact of imperfect transcription on the DNN and the ASR performance. As labelling error is the source of the problem, we investigate two options available to reduce that: selecting data with fewer errors, and changing the dependence on noise by reducing label precision. Both confidence based data selection and label resolution change are explored in the context of two scenarios of matched and unmatched unlabelled data. We introduce improved DNN based confidence score estimators and show their performance on data selection for both scenarios. Confidence score based data selection was found to yield up to 14.6% relative WER reduction, while better balance between label resolution and recognition hypothesis accuracy allowed further WER reductions by 16.6% relative in the mismatched scenario.
Pengyuan Zhang, Yulan Liu, Thomas Hain
SLT2
2013 The sheffield wargames corpus
abstract
Recognition of speech in natural environments is a challenging task, even more so if this involves conversations between sev-eral speakers. Work on meeting recognition has addressed some of the significant challenges, mostly targeting formal, business style meetings where people are mostly in a static position in a room. Only limited data is available that contains high qual-ity near and far field data from real interactions between par-ticipants. In this paper we present a new corpus for research on speech recognition, speaker tracking and diarisation, based on recordings of native speakers of English playing a table-top wargame. The Sheffield Wargames Corpus comprises 7 hours of data from 10 recording sessions, obtained from 96 micro-phones, 3 video cameras and, most importantly, 3D location data provided by a sensor tracking system. The corpus repre-sents a unique resource, that provides for the first time location tracks (1.3Hz) of speakers that are constantly moving and talk-ing. The corpus is available for research purposes, and includes annotated development and evaluation test sets. Baseline results for close-talking and far field sets are included in this paper. 1.
Charles Fox, Yulan Liu, Erich Zwyssig, Thomas Hain
INTERSPEECH2