Leyuan Qu

dblp:180/2820 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
12since 2021 · last 2026
0000-0001-6694-5355ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 7 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Think-Before-Draw: Decomposing emotion semantics for fine-grained controllable generation of expressive talking heads
Hanlei Shi, Leyuan Qu, Yu Liu 0132, Linlin Gong, Yuhua Zheng, Taihao Li
Pattern Recognit.2
2025 Label Semantic-Driven Contrastive Learning for Speech Emotion Recognition
Jiaxi Hu, Leyuan Qu, Haoxun Li, Taihao Li
INTERSPEECH2
2025 EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis
Haoxun Li, Leyuan Qu, Jiaxi Hu, Taihao Li
INTERSPEECH2
2025 HOPE: Hierarchical Fusion for Optimized and Personality-Aware Estimation of Depression
abstract
Depression detection remains challenged by generalized modeling approaches that fail to account for individual heterogeneity. To address this, the Multimodal Personality-aware Depression Detection (MPDD) Challenge introduced personalized features into the modeling process, aiming to better capture individual variability. However, the baseline models still exhibit two critical limitations: the neglect of textual semantics embedded in audio, and inconsistent predictions for the same subject across tasks and samples. Motivated by these limitations, we introduce HOPE (Hierarchical fusion for Optimized and Personality-aware Estimation of Depression), a unified framework for consistent, subject-level depression estimation. HOPE first employs a Latent Semantic Projection (LSP) module to reconstruct textual semantics from audio features when transcripts are unavailable. It then introduces a consistency-aware integration mechanism that hierarchically fuses multi-branch predictions to resolve inter-task and inter-sample contradictions. HOPE achieved first place in the MPDD Challenge Young Track, demonstrating strong cross-modal learning capabilities and consistent, subject-level depression prediction.
Hanlei Shi, Yu Liu 0132, Haoxun Li, Jiaxi Hu, Leyuan Qu, Taihao Li
ACM Multimedia6
2025 Disentanglement of Prosody Representations via Diffusion Models and Scheduled Gradient Reversal
abstract
Prosody plays a fundamental role in human speech and communication, facilitating intelligibility and conveying emotional and cognitive states. Extracting accurate prosodic information from speech is vital for building assistive technology, such as controllable speech synthesis, speaking style transfer, and speech emotion recognition (SER). However, it is challenging to disentangle speaker-independent prosody representations since prosodic attributes, such as intonation, excessively entangle with speaker-specific attributes, e.g., pitch. In this article, we propose a novel model, called Diffsody, to disentangle and refine prosody representations: 1) to disentangle prosody representations, we leverage the expressive generative ability of a diffusion model by conditioning it on quantified semantic information and pretrained speaker embeddings. Additionally, a prosody encoder automatically learns prosody representations used for spectrogram reconstruction in an unsupervised fashion; and 2) to refine and learn speaker-invariant prosody representations, a scheduled gradient reversal layer (sGRL) is proposed and integrated into the prosody encoder of Diffsody. We extensively evaluate Diffsody through qualitative and quantitative means. t-SNE visualization and speaker verification experiments demonstrate the efficacy of the sGRL method in preventing speaker-specific information leakage. Experimental results on speaker-independent SER and automatic depression detection (ADD) tasks demonstrate that Diffsody can efficiently factorize speaker-independent prosody representations, resulting in a significant boost in SER and ADD. In addition, Diffsody synergistically integrates with the semantic representation model WavLM, which leads to a discernibly elevated performance, outperforming contemporary methods in both SER and ADD tasks. Furthermore, the Diffsody model exhibits promising potential for various practical applications, such as voice or style conversion. Some audio samples can be found on our https://leyuanqu.github.io/Diffsody/demo website.
Leyuan Qu, Cornelius Weber, Wei Wang 0310, Jia Jin, Yingming Gao, Taihao Li, Stefan Wermter
IEEE Trans. Neural Networks Learn. Syst.1
2024 Improving Speech Emotion Recognition with Unsupervised Speaking Style Transfer
abstract
Humans can effortlessly modify various prosodic attributes, such as the placement of stress and the intensity of sentiment, to convey a specific emotion while maintaining consistent linguistic content. Motivated by this capability, we propose EmoAug, a novel style transfer model designed to enhance emotional expression and tackle the data scarcity issue in speech emotion recognition tasks. EmoAug consists of a semantic encoder and a paralinguistic encoder that represent verbal and non-verbal information respectively. Additionally, a decoder reconstructs speech signals by conditioning on the aforementioned two information flows in an unsupervised fashion. Once training is completed, EmoAug enriches expressions of emotional speech with different prosodic attributes, such as stress, rhythm and intensity, by feeding different styles into the paralinguistic encoder. EmoAug enables us to generate similar numbers of samples for each class to tackle the data imbalance issue as well. Experimental results on the IEMOCAP dataset demonstrate that EmoAug can successfully transfer different speaking styles while retaining the speaker identity and semantic content. Furthermore, we train a SER model with data augmented by EmoAug and show that the augmented model not only surpasses the state-of-the-art supervised and self-supervised methods but also overcomes overfitting problems caused by data imbalance. Some audio samples can be found on our demo website1.
Leyuan Qu, Wei Wang 0310, Cornelius Weber, Pengcheng Yue, Taihao Li, Stefan Wermter
ICASSP1
2024 Multi-Modal Emotion Recognition Using Multiple Acoustic Features and Dual Cross-Modal Transformer
abstract
Multi-modal emotion recognition (MER) using speech and text has attracted extensive attention because of the easy availability of data for these two modalities. Recently, the self-surprised learning (SSL) pre-trained model has become the state-of-the-art (SOTA) method for the extraction of acoustic and textual features. However, the SSL speech representation may lose some important paralinguistic information, resulting in limited speech knowledge for MER. In this paper, we propose to adopt two kinds of acoustic features (i.e., the SSL representation and the spectral feature) as inputs to comprehensively extract speech characteristics. In addition, a dual cross-modal Transformer module is presented to model the interaction on the unaligned sequences between the textual feature and two acoustic features. Moreover, we introduce a blended loss including two uni-modal losses to better extract the uni-modal information. Experiments conducted on the widely used IEMOCAP dataset indicate that our proposed method achieves the SOTA performance compared with previous methods.
Pengcheng Yue, Leyuan Qu, Taihao Li, Yu-Ping Ruan
ICASSP3
2024 Disentangling Prosody Representations With Unsupervised Speech Reconstruction
abstract
Human speech can be characterized by different components, including semantic content, speaker identity and prosodic information. Significant progress has been made in disentangling representations for semantic content and speaker identity in Automatic Speech Recognition (ASR) and speaker verification tasks respectively. However, it is still an open challenging research question to extract prosodic information because of the intrinsic association of different attributes, such as timbre and rhythm, and because of the need for supervised training schemes to achieve robust large-scale and speaker-independent ASR. The aim of this paper is to address the disentanglement of emotional prosody from speech based on unsupervised reconstruction. Specifically, we identify, design, implement and integrate three crucial components in our proposed speech reconstruction model Prosody2Vec: (1) a unit encoder that transforms speech signals into discrete units for semantic content, (2) a pretrained speaker verification model to generate speaker identity embeddings, and (3) a trainable prosody encoder to learn prosody representations. We first pretrain the Prosody2Vec representations on unlabelled emotional speech corpora, then fine-tune the model on specific datasets to perform Speech Emotion Recognition (SER) and Emotional Voice Conversion (EVC) tasks. Both objective (weighted and unweighted accuracies) and subjective (mean opinion score) evaluations on the EVC task suggest that Prosody2Vec effectively captures general prosodic features that can be smoothly transferred to other emotional speech. In addition, our SER experiments on the IEMOCAP dataset reveal that the prosody features learned by Prosody2Vec are complementary and beneficial for the performance of widely used speech pretraining models and surpass the state-of-the-art methods when combining Prosody2Vec with HuBERT representations. Some audio samples can be found on our demo website
Leyuan Qu, Taihao Li, Cornelius Weber, Theresa Pekarek-Rosin, Fuji Ren, Stefan Wermter
IEEE ACM Trans. Audio Speech Lang. Process.1
2024 LipSound2: Self-Supervised Pre-Training for Lip-to-Speech Reconstruction and Lip Reading
abstract
The aim of this work is to investigate the impact of crossmodal self-supervised pre-training for speech reconstruction (video-to-audio) by leveraging the natural co-occurrence of audio and visual streams in videos. We propose LipSound2 that consists of an encoder-decoder architecture and location-aware attention mechanism to map face image sequences to mel-scale spectrograms directly without requiring any human annotations. The proposed LipSound2 model is first pre-trained on ∼ 2400 -h multilingual (e.g., English and German) audio-visual data (VoxCeleb2). To verify the generalizability of the proposed method, we then fine-tune the pre-trained model on domain-specific datasets (GRID and TCD-TIMIT) for English speech reconstruction and achieve a significant improvement on speech quality and intelligibility compared to previous approaches in speaker-dependent and speaker-independent settings. In addition to English, we conduct Chinese speech reconstruction on the Chinese Mandarin Lip Reading (CMLR) dataset to verify the impact on transferability. Finally, we train the cascaded lip reading (video-to-text) system by fine-tuning the generated audios on a pre-trained speech recognition system and achieve the state-of-the-art performance on both English and Chinese benchmark datasets.
Leyuan Qu, Cornelius Weber, Stefan Wermter
IEEE Trans. Neural Networks Learn. Syst.1
2023 Emphasizing unseen words: New vocabulary acquisition for end-to-end speech recognition
abstract
Due to the dynamic nature of human language, automatic speech recognition (ASR) systems need to continuously acquire new vocabulary. Out-Of-Vocabulary (OOV) words, such as trending words and new named entities, pose problems to modern ASR systems that require long training times to adapt their large numbers of parameters. Different from most previous research focusing on language model post-processing, we tackle this problem on an earlier processing level and eliminate the bias in acoustic modeling to recognize OOV words acoustically. We propose to generate OOV words using text-to-speech systems and to rescale losses to encourage neural networks to pay more attention to OOV words. Specifically, we enlarge the classification loss used for training neural networks' parameters of utterances containing OOV words (sentence-level), or rescale the gradient used for back-propagation for OOV words (word-level), when fine-tuning a previously trained model on synthetic audio. To overcome catastrophic forgetting, we also explore the combination of loss rescaling and model regularization, i.e. L2 regularization and elastic weight consolidation (EWC). Compared with previous methods that just fine-tune synthetic audio with EWC, the experimental results on the LibriSpeech benchmark reveal that our proposed loss rescaling approach can achieve significant improvement on the recall rate with only a slight decrease on word error rate. Moreover, word-level rescaling is more stable than utterance-level rescaling and leads to higher recall rates and precision rates on OOV word recognition. Furthermore, our proposed combined loss rescaling and weight consolidation methods can support continual learning of an ASR system.
Leyuan Qu, Cornelius Weber, Stefan Wermter
Neural Networks1
2022 A Multimodal German Dataset for Automatic Lip Reading Systems and Transfer Learning
abstract
Large datasets as required for deep learning of lip reading do not exist in many languages. In this paper we present the dataset GLips (German Lips) consisting of 250,000 publicly available videos of the faces of speakers of the Hessian Parliament, which was processed for word-level lip reading using an automatic pipeline. The format is similar to that of the English language LRW (Lip Reading in the Wild) dataset, with each video encoding one word of interest in a context of 1.16 seconds duration, which yields compatibility for studying transfer learning between both datasets. By training a deep neural network, we investigate whether lip reading has language-independent features, so that datasets of different languages can be used to improve lip reading models. We demonstrate learning from scratch and show that transfer learning from LRW to GLips and vice versa improves learning speed and performance, in particular for the validation set.
Gerald Schwiebert, Cornelius Weber, Leyuan Qu, Henrique Siqueira, Stefan Wermter
LREC3
2021 Hearing Faces: Target Speaker Text-to-Speech Synthesis from a Face
abstract
The existence of a learnable cross-modal association between a person's face and their voice is recently becoming more and more evident. This provides the basis for the task of target speaker text-to-speech (TTS) synthesis from face ref-erence. In this paper, we approach this task by proposing a cross-modal model architecture combining existing unimodal models. We use Tacotron 2 multi-speaker TTS with auditory speaker embeddings based on Global Style Tokens. We trans-fer learn a FaceNet face encoder to predict these embeddings from a static face image reference instead of a voice reference and thus predict a speaker's voice and speaking characteristics from their face. Compared to Face2Speech, the only existing work on this task, we use a more modular architecture that allows the use of openly available and pretrained model components. This approach enables high-quality speech synthesis and allows for an easily extensible model architecture. Ex-perimental results show good matching ability while retaining better voice naturalness than Face2Speech. We examine the limitations of our model and discuss multiple possible av-enues of improvement for future work.
Björn Plüster, Cornelius Weber, Leyuan Qu, Stefan Wermter
ASRU3
2020 Variational Autoencoder with Global- and Medium Timescale Auxiliaries for Emotion Recognition from Speech
Hussam Almotlak, Cornelius Weber, Leyuan Qu, Stefan Wermter
ICANN (1)3
2020 Multimodal Target Speech Separation with Voice and Face References
abstract
Target speech separation refers to isolating target speech from a multi-speaker mixture signal by conditioning on auxiliary information about the target speaker. Different from the mainstream audio-visual approaches which usually require simultaneous visual streams as additional input, e.g. the corresponding lip movement sequences, in our approach we propose the novel use of a single face profile of the target speaker to separate expected clean speech. We exploit the fact that the image of a face contains information about the person's speech sound. Compared to using a simultaneous visual sequence, a face image is easier to obtain by pre-enrollment or on websites, which enables the system to generalize to devices without cameras. To this end, we incorporate face embeddings extracted from a pretrained model for face recognition into the speech separation, which guide the system in predicting a target speaker mask in the time-frequency domain. The experimental results show that a pre-enrolled face image is able to benefit separating expected speech signals. Additionally, face information is complementary to voice reference and we show that further improvement can be achieved when combing both face and voice embeddings.
Leyuan Qu, Cornelius Weber, Stefan Wermter
INTERSPEECH1
2019 LipSound: Neural Mel-Spectrogram Reconstruction for Lip Reading
Leyuan Qu, Cornelius Weber, Stefan Wermter
INTERSPEECH1
2018 Combining Articulatory Features with End-to-End Learning in Speech Recognition
Leyuan Qu, Cornelius Weber, Egor Lakomkin, Johannes Twiefel, Stefan Wermter
ICANN (3)1
2016 Landmark of Mandarin nasal codas and its application in pronunciation error detection
abstract
L2 learners of Mandarin have difficulty learning native-like pronunciation of nasal codas. In order to help them learn native-like pronunciation, we propose to develop targeted classifiers for automatic pronunciation error detection. In this paper, perceptual experiments with modified speech are designed to analyze the exact position of the landmark of a nasal coda. Based on perceptual results from isolated words, we propose that information about nasal coda place of articulation is most dense near a landmark at the center of the nasalized vowel. Landmarks detected in a database of Japanese learners of Mandarin, and classified as correct vs. incorrect using an SVM. The result shows that the detection performance of the SVM+Landmark system is similar to that of a DNN-HMM+MFCC system. When the two systems are combined, an FRR of 4.6% is achieved at DA of 83.9%. This performance is comparable to that of previously developed classifiers for 16 common Mandarin pronunciation errors.
Yanlu Xie, Mark Hasegawa-Johnson, Leyuan Qu, Jinsong Zhang 0001
ICASSP3