Rui Liu 0008

dblp:42/469-8 · DBLP profile ↗
← Back
51ranked-venue papers
30as first author
45since 2021 · last 2026
0000-0003-4524-7413ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 22 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 18 first-author · 26 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards Authentic Movie Dubbing with Retrieve-Augmented Director-Actor Interaction Learning
abstract
The automatic movie dubbing model generates vivid speech from given scripts, replicating a speaker's timbre from a brief timbre prompt while ensuring lip-sync with the silent video. Existing approaches simulate a simplified workflow where actors dub directly without preparation, overlooking the critical director–actor interaction. In contrast, authentic workflows involve a dynamic collaboration: directors actively engage with actors, guiding them to internalize the context cues, specifically emotion, before performance. To address this issue, we propose a new Retrieve-Augmented Director-Actor Interaction Learning scheme to achieve authentic movie dubbing, termed Authentic-Dubber, which contains three novel mechanisms: (1) We construct a multimodal Reference Footage library to simulate the learning footage provided by directors. Note that we integrate Large Language Models (LLMs) to achieve deep comprehension of emotional representations across multimodal signals. (2) To emulate how actors efficiently and comprehensively internalize director-provided footage during dubbing, we propose an Emotion-Similarity-based Retrieval-Augmentation strategy. This strategy retrieves the most relevant multimodal information that aligns with the target silent video. (3) We develop a Progressive Graph-based speech generation approach that incrementally incorporates the retrieved multimodal emotional knowledge, thereby simulating the actor's final dubbing process. The above mechanisms enable the Authentic-Dubber to faithfully replicate the authentic dubbing workflow, achieving comprehensive improvements in emotional expressiveness. Both subjective and objective evaluations on the V2C-Animation benchmark dataset validate the effectiveness.
Rui Liu 0008, Yuan Zhao 0016, Zhenqi Jia
AAAI1
2026 Emphasis rendering for conversational text-to-speech with multi-modal multi-scale context modeling
Rui Liu 0008, Zhenqi Jia, Yifan Hu 0004, Haizhou Li 0001
Speech Commun.1
2025 Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech
abstract
Visual Text-to-Speech (VTTS) aims to take the environmental image as the prompt to synthesize the reverberant speech for the spoken content. The challenge of this task lies in understanding the spatial environment from the image. Many attempts have been made to extract global spatial visual information from the RGB space of an spatial image. However, local and depth image information are crucial for understanding the spatial environment, which previous works have ignored. To address the issues, we propose a novel multi-modal and multi-scale spatial environment understanding scheme to achieve immersive VTTS, termed M2SE-VTTS. The multi-modal aims to take both the RGB and Depth spaces of the spatial image to learn more comprehensive spatial information, and the multi-scale seeks to model the local and global spatial knowledge simultaneously. Specifically, we first split the RGB and Depth images into patches and adopt the Gemini-generated environment captions to guide the local spatial understanding. After that, the multi-modal and multi-scale features are integrated by the local-aware global spatial understanding. In this way, M2SE-VTTS effectively models the interactions between local and global spatial contexts in the multi-modal spatial environment. Objective and subjective evaluations suggest that our model outperforms the advanced baselines in environmental speech generation.
Rui Liu 0008, Shuwei He, Yifan Hu 0004, Haizhou Li 0001
AAAI1
2025 Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis
abstract
Conversational Speech Synthesis (CSS) aims to generate speech with natural prosody by understanding the multimodal dialogue history (MDH).The latest work predicts the accurate prosody expression of the target utterance by modeling the utterance-level interaction characteristics of MDH and the target utterance.However, MDH contains fine-grained semantic and prosody knowledge at the word level.Existing methods overlook the fine-grained semantic and prosodic interaction modeling.To address this gap, we propose MFCIG-CSS, a novel Multimodal Fine-grained Context Interaction Graph-based CSS system.Our approach constructs two specialized multimodal fine-grained dialogue interaction graphs: a semantic interaction graph and a prosody interaction graph.These two interaction graphs effectively encode interactions between word-level semantics, prosody, and their influence on subsequent utterances in MDH.The encoded interaction features are then leveraged to enhance synthesized speech with natural conversational prosody.Experiments on the DailyTalk dataset demonstrate that MFCIG-CSS outperforms all baseline models in terms of prosodic expressiveness.Code and speech samples are available at https://github.com/AI-S2-Lab/MFCIG-CSS.
Zhenqi Jia, Rui Liu 0008, Berrak Sisman, Haizhou Li 0001
EMNLP2
2025 MEIJU - The 1st Multimodal Emotion and Intent Joint Understanding Challenge
abstract
Multimodal Emotion and Intent Joint Understanding (MEIJU) aims to decode the semantic information expressed in the multimodal dialogues while inferring the emotions and intents, providing users with a more humanized human-machine interaction experience. However, challenges such as difficulties in data acquisition and imbalance annotations have made it difficult for current methods to meet the demands of practical applications. Therefore, we have organized two tracks focusing on the key themes of semi-supervised learning and class imbalance. Additionally, we have prepared data in two different languages (English and Mandarin) for each track, treating each language as a sub-track, to encourage participants to explore solutions in more diverse linguistic environments. Our code can be found at https://github.com/AI-S2-Lab/MEIJU2025-baseline.
Rui Liu 0008, Xiaofen Xing, Zheng Lian 0004, Haizhou Li 0001, Björn W. Schuller, Haolin Zuo
ICASSP1
2025 Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech
abstract
Visual Text-to-Speech (VTTS) aims to take the environmental image as the prompt to synthesize reverberant speech for the spoken content. Previous works focus on the RGB modality for global environmental modeling, overlooking the potential of multi-source spatial knowledge like depth, speaker position, and environmental semantics. To address these issues, we propose a novel multi-source spatial knowledge understanding scheme for immersive VTTS, termed MS2KU-VTTS. Specifically, we first prioritize RGB image as the dominant source and consider depth image, speaker position knowledge from object detection, and Gemini-generated semantic captions as supplementary sources. Afterwards, we propose a serial interaction mechanism to effectively integrate both dominant and supplementary sources. The resulting multi-source knowledge is dynamically integrated based on the respective contributions of each source. This enriched interaction and integration of multi-source spatial knowledge guides the speech generation model, enhancing the immersive speech experience. Experimental results demonstrate that the MS2KU-VTTS surpasses existing baselines in generating immersive speech. Demos and code are available at: https://github.com/AI-S2-Lab/MS2KU-VTTS.
Shuwei He, Rui Liu 0008
ICASSP2
2025 Intra- and Inter-modal Context Interaction Modeling for Conversational Speech Synthesis
abstract
Conversational Speech Synthesis (CSS) aims to effectively take the multimodal dialogue history (MDH) to generate speech with appropriate conversational prosody for target utterance. The key challenge of CSS is to model the interaction between the MDH and the target utterance. Note that text and speech modalities in MDH have their own unique influences, and they complement each other to produce a comprehensive impact on the target utterance. Previous works did not explicitly model such intra-modal and inter-modal interactions. To address this issue, we propose a new intra-modal and inter-modal context interaction scheme-based CSS system, termed I3-CSS. Specifically, in the training phase, we combine the MDH with the text and speech modalities in the target utterance to obtain four modal combinations, including Historical Text-Next Text, Historical Speech-Next Speech, Historical Text-Next Speech, and Historical Speech-Next Text. Then, we design two contrastive learning-based intra-modal and two inter-modal interaction modules to deeply learn the intra-modal and inter-modal context interaction. In the inference phase, we take MDH and adopt trained interaction modules to fully infer the speech prosody of the target utterance’s text content. Subjective and objective experiments on the DailyTalk dataset show that I3-CSS outperforms the advanced baselines in terms of prosody expressiveness. Code and speech samples are available at https://github.com/AI-S2-Lab/I3CSS.
Zhenqi Jia, Rui Liu 0008
ICASSP2
2025 Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction
abstract
Automatic Video Dubbing (AVD) generates speech aligned with lip motion and facial emotion from scripts. Recent research focuses on modeling multimodal context to enhance prosody expressiveness but overlooks two key issues: 1) Mul-tiscale prosody expression attributes in the context influence the current sentence’s prosody. 2) Prosody cues in context interact with the current sentence, impacting the final prosody expressiveness. To tackle these challenges, we propose M2CI-Dubber, a Multiscale Multimodal Context Interaction scheme for AVD. This scheme includes two shared M2CI encoders to model the multiscale multimodal context and facilitate its deep interaction with the current sentence. By extracting global and local features for each modality in the context, utilizing attention-based mechanisms for aggregation and interaction, and employing an interaction-based graph attention network for fusion, the proposed approach enhances the prosody expressiveness of synthesized speech for the current sentence. Experiments on the Chem dataset show our model outperforms baselines in dubbing expressiveness. The code and demos are available at https://github.com/AI-S2-Lab/M2CI-Dubber.
Yuan Zhao 0016, Rui Liu 0008, Gaoxiang Cong 0001
ICASSP2
2025 AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models
abstract
The emergence of multimodal large language models (MLLMs) advances multimodal emotion recognition (MER) to the next level—from naive discriminative tasks to complex emotion understanding with advanced video understanding abilities and natural language description. However, the current community suffers from a lack of large-scale datasets with intensive, descriptive emotion annotations, as well as a multimodal-centric framework to maximize the potential of MLLMs for emotion understanding. To address this, we establish a new benchmark for MLLM-based emotion understanding with a novel dataset (MER-Caption) and a new model (AffectGPT). Utilizing our model-based crowd-sourcing data collection strategy, we construct the largest descriptive emotion dataset to date (by far), featuring over 2K fine-grained emotion categories across 115K samples. We also introduce the AffectGPT model, designed with pre-fusion operations to enhance multimodal integration. Finally, we present MER-UniBench, a unified benchmark with evaluation metrics tailored for typical MER tasks and the free-form, natural language output style of MLLMs. Extensive experimental results show AffectGPT's robust performance across various MER tasks. We have released both the code and the dataset to advance research and development in emotion understanding: https://github.com/zeroQiaoba/AffectGPT.
Zheng Lian 0004, Haoyu Chen 0001, Lan Chen 0005, Haiyang Sun 0004, Licai Sun, Yong Ren 0006, Zebang Cheng, Bin Liu 0041, Rui Liu 0008, Xiaojiang Peng, Jiangyan Yi, Jianhua Tao 0001
ICML9
2025 OV-MER: Towards Open-Vocabulary Multimodal Emotion Recognition
abstract
Multimodal Emotion Recognition (MER) is a critical research area that seeks to decode human emotions from diverse data modalities. However, existing machine learning methods predominantly rely on predefined emotion taxonomies, which fail to capture the inherent complexity, subtlety, and multi-appraisal nature of human emotional experiences, as demonstrated by studies in psychology and cognitive science. To overcome this limitation, we advocate for introducing the concept of open vocabulary into MER. This paradigm shift aims to enable models to predict emotions beyond a fixed label space, accommodating a flexible set of categories to better reflect the nuanced spectrum of human emotions. To achieve this, we propose a novel paradigm: Open-Vocabulary MER (OV-MER), which enables emotion prediction without being confined to predefined spaces. However, constructing a dataset that encompasses the full range of emotions for OV-MER is practically infeasible; hence, we present a comprehensive solution including a newly curated database, novel evaluation metrics, and a preliminary benchmark. By advancing MER from basic emotions to more nuanced and diverse emotional states, we hope this work can inspire the next generation of MER, enhancing its generalizability and applicability in real-world scenarios. Code and dataset are available at: https://github.com/zeroQiaoba/AffectGPT.
Zheng Lian 0004, Haiyang Sun 0004, Licai Sun, Haoyu Chen 0001, Lan Chen 0005, Zhuofan Wen 0001, Hailiang Yao, Bin Liu 0041, Rui Liu 0008, Shan Liang 0007, Ya Li 0001, Jiangyan Yi, Jianhua Tao 0001
ICML12
2025 Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset
Rui Liu 0008, Pu Gao, Jiatian Xi, Berrak Sisman, Carlos Busso, Haizhou Li 0001
INTERSPEECH1
2025 UniTalker: Conversational Speech-Visual Synthesis
abstract
Conversational Speech Synthesis (CSS) is a key task in the user-agent interaction area, aiming to generate more expressive and empathetic speech for users. However, it is well-known that ''listening'' and ''eye contact'' play crucial roles in conveying emotions during real-world interpersonal communication. Existing CSS research is limited to perceiving only text and speech within the dialogue context, which restricts its effectiveness. Moreover, speech-only responses further constrain the interactive experience. To address these limitations, we introduce a Conversational Speech-Visual Synthesis (CSVS) task as an extension of traditional CSS. By leveraging multimodal dialogue context, it provides users with coherent audiovisual responses. To this end, we develop a CSVS system named UniTalker, which is a unified model that seamlessly integrates multimodal perception and multimodal rendering capabilities. Specifically, it leverages a large-scale language model to comprehensively understand multimodal cues in the dialogue context, including speaker, text, speech, and the talking-face animations. After that, it employs multi-task sequence prediction to first infer the target utterance's emotion and then generate empathetic speech and natural talking-face animations. To ensure that the generated speech-visual content remains consistent in terms of emotion, content, and duration, we introduce three key optimizations: 1) Designing a specialized neural landmark codec to tokenize and reconstruct facial expression sequences. 2) Proposing a bimodal speech-visual hard alignment decoding strategy. 3) Applying emotion-guided rendering during the generation stage. Comprehensive objective and subjective experiments demonstrate that our model synthesizes more empathetic speech and provides users with more natural and emotionally consistent talking-face animations. The source code and generated samples are available at: https://github.com/AI-S2-Lab/UniTalker.
Yifan Hu 0004, Rui Liu 0008, Yi Ren 0006, Xiang Yin 0006, Haizhou Li 0001
ACM Multimedia2
2025 MER 2025: When Affective Computing Meets Large Language Models
abstract
MER2025 is the third year of our MER series of challenges. Previously, MER2023 (http://merchallenge.cn/mer2023) focused on multi-label learning, noise robustness, and semi-supervised learning, while MER2024 (https://zeroqiaoba.github.io/MER2024-website) introduced a new track dedicated to open-vocabulary emotion recognition. This year, MER2025 centers on the theme ''When Affective Computing Meets Large Language Models (LLMs)''. We aim to shift the paradigm from traditional categorical frameworks reliant on predefined emotion taxonomies to LLM-driven generative methods, offering innovative solutions for more accurate and reliable emotion understanding. The challenge contains four tracks: MER-SEMI focuses on fixed categorical emotion recognition enhanced by semi-supervised learning; MER-FG explores fine-grained emotions, expanding recognition from basic to nuanced emotional states; MER-DES incorporates multimodal cues (beyond emotion words) into predictions to enhance model interpretability; MER-PR reveals whether emotion prediction results can improve personality recognition performance. For the first three tracks, the baseline code is available at MERTools (https://github.com/zeroQiaoba/MERTools) and datasets can be accessed via Hugging Face (https://huggingface.co/datasets/MERChallenge/MER2025). For the last track, the dataset and baseline code are available on GitHub (https://github.com/cai-cong/MER25_personality).
Zheng Lian 0004, Rui Liu 0008, Kele Xu, Bin Liu 0041, Xuefei Liu, Yazhou Zhang 0001, Xin Liu 0012, Yong Li 0032, Zebang Cheng, Haolin Zuo, Ziyang Ma 0001, Xiaojiang Peng, Xie Chen 0001, Ya Li 0001, Erik Cambria, Guoying Zhao 0001, Björn W. Schuller, Jianhua Tao 0001
ACM Multimedia2
2025 Hardness-Aware Dynamic Curriculum Learning for Robust Multimodal Emotion Recognition with Missing Modalities
abstract
Missing modalities have recently emerged as a critical research direction in multimodal emotion recognition (MER). Conventional approaches typically address this issue through missing modality reconstruction. However, these methods fail to account for variations in reconstruction difficulty across different samples, consequently limiting the model's ability to handle hard samples effectively. To overcome this limitation, we propose a novel Hardness-Aware Dynamic Curriculum Learning framework, termed HARDY-MER. Our framework operates in two key stages: first, it estimates the hardness level of each sample, and second, it strategically emphasizes hard samples during training to enhance model performance on these challenging instances. Specifically, we first introduce a Multi-view Hardness Evaluation mechanism that quantifies reconstruction difficulty by considering both Direct Hardness (modality reconstruction errors) and Indirect Hardness (cross-modal mutual information). Meanwhile, we introduce a Retrieval-based Dynamic Curriculum Learning strategy that dynamically adjusts the training curriculum by retrieving samples with similar semantic information and balancing the learning focus between easy and hard instances. Extensive experiments on benchmark datasets demonstrate that HARDY-MER consistently outperforms existing methods in missing-modality scenarios. Our code will be made publicly available at https://github.com/HARDY-MER/HARDY-MER.
Rui Liu 0008, Haolin Zuo, Zheng Lian 0004, Hongyu Yuan
ACM Multimedia1
2025 MSD-YOLO: An Efficient Algorithm for Small Target Detection
Dongyu Liu, Rui Liu 0008, Zhecong Xing, Weiyang Geng
MMM (3)3
2025 Robust Self-Localization of Wireless Acoustic Sensor Networks
abstract
Wireless acoustic sensor networks (WASNs), or the so-called Internet of Audio Things (IoAuT), have attracted increasing attention in the Internet of Things community. As the geometric structure of WASNs is required in audio/speech processing tasks like source localization or acoustic beamforming, automatic self-localization of sensors is necessary. However, most of the existing approaches suffer from poor stability, as their constructed cost functions involve nonconvex programming. To address this issue, we investigate the robust self-localization (or geometry calibration) of WASNs in this article. Specifically, a rough self-localization (RSL) method is first presented based on measurements including Time-Difference-of-Arrivals (TDoAs), direction-of-arrivals (DoA), and energy-rates (ERs), and its closed-form solution is further derived. As ER estimates are sensitive to acoustic environments, the performance of the RSL method is somewhat limited. Therefore, a precise self-localization (PSL) method is then developed by building a weighted (and nonconvex) TDoA-DoA cost function, after regarding the RSL approach as an initialization step. As the RSL offers better initial values compared with existing initialization strategies, the combination of RSL and PSL methods (named as RSL-PSL method) shows stronger robustness and stability. In addition, computational complexity of both RSL and PSL methods is analyzed in detail. Finally, the Cramér-Rao Bound (CRB) of the PSL method is derived to show its theoretical lower bound. The proposed RSL-PSL method outperforms the state-of-the-arts in terms of stability and accuracy, which is confirmed by numerical real-world and simulation experiments.
Xu Wang 0058, De Hu, Rui Liu 0008, Feilong Bao
IEEE Internet Things J.3
2025 Connecting Cross-Modal Representations for Compact and Robust Multimodal Sentiment Analysis With Sentiment Word Substitution Error
abstract
Multimodal Sentiment Analysis (MSA) seeks to fuse textual, acoustic, and visual information to predict a speaker’s sentiment states effectively. However, in real-world scenarios, the text modality received by MSA systems is often obtained through automatic speech recognition (ASR) models. Unfortunately, ASR may erroneously recognize sentiment words as phonetically similar neutral alternatives, leading to sentiment degradation in text and impacting MSA accuracy. Recent attempts aim to first identify the sentiment word substitution (SWS) error in ASR results and then refine the corrupted word embeddings using multimodal information for final multimodal fusion. However, such a method includes a burdensome and ambiguous detection operation and ignores the inherent correlations and heterogeneity among different modalities. To address these issues, we propose a more compact system, termedARF-MSAconsisting of three key components to achieving robust MSA with SWS errors: 1)Alignment: we establish connections between the “text-acoustic’ and “text-visual” representations to effectively map the “text-acoustic-visual” data into a unified sentiment space by leveraging their multimodal correlation knowledge; 2)Refinement: we perform fine-grained comparisons between the text modality and the other two modalities in the unified sentiment space, enabling refinement of the sentiment expression within the text modality more concisely; 3)Fusion: Finally, we hierarchically fuse the dominant and non-dominant representation from three heterogeneity modalities to obtain the multimodal feature for MSA. We conduct extensive experiments on the real-world datasets and the results demonstrate the effectiveness of our model.
Qiyuan Sun, Haolin Zuo, Rui Liu 0008, Haizhou Li 0001
IEEE Trans. Affect. Comput.3
2024 Emotion Rendering for Conversational Speech Synthesis with Heterogeneous Graph-Based Context Modeling
abstract
Conversational Speech Synthesis (CSS) aims to accurately express an utterance with the appropriate prosody and emotional inflection within a conversational setting. While recognising the significance of CSS task, the prior studies have not thoroughly investigated the emotional expressiveness problems due to the scarcity of emotional conversational datasets and the difficulty of stateful emotion modeling. In this paper, we propose a novel emotional CSS model, termed ECSS, that includes two main components: 1) to enhance emotion understanding, we introduce a heterogeneous graph-based emotional context modeling mechanism, which takes the multi-source dialogue history as input to model the dialogue context and learn the emotion cues from the context; 2) to achieve emotion rendering, we employ a contrastive learning-based emotion renderer module to infer the accurate emotion style for the target utterance. To address the issue of data scarcity, we meticulously create emotional labels in terms of category and intensity, and annotate additional emotional information on the existing conversational dataset (DailyTalk). Both objective and subjective evaluations suggest that our model outperforms the baseline models in understanding and rendering emotions. These evaluations also underscore the importance of comprehensive emotional annotations. Code and audio samples can be found at: https://github.com/walker-hyf/ECSS.
Rui Liu 0008, Yifan Hu 0004, Yi Ren 0006, Xiang Yin 0006, Haizhou Li 0001
AAAI1
2024 DCM-YOLOv8: An Improved YOLOv8-Based Small Target Detection Model for UAV Images
Zhecong Xing, Rui Liu 0008
ICIC (6)3
2024 OB-YOLO: A UAV Image Detection Model for Reducing Computational Resource Consumption
abstract
In contemporary society, the pervasive integration of Unmanned Aerial Vehicles (UAVs) in everyday activities is notable. Object detection emerges as a pivotal task within the UAV operational context. However, challenges such as the presence of expansive backgrounds in UAV images, insufficient target pixel resolution, and the prevalence of image interferences contribute to the diminished accuracy observed in existing object detection models tailored for UAV aerial imagery. Conventional strategies employed to enhance accuracy often incur exorbitant computational costs, failing to strike a harmonious balance between precision improvement and computational resource utilization. To address these challenges, this paper introduces an optimized variant of YOLOv8, denoted as OB-YOLO, specifically tailored for UAV aerial photography scenarios. The proposed model exhibits enhanced accuracy while concurrently mitigating parameter and floating-point operation costs. Particularly, the BiFPN concept is incorporated to fortify the feature fusion process, enabling comprehensive consideration and reuse of multi-scale feature fusion within the model. Additionally, the study integrates the full-dimensional dynamic convolution (ODConv) structure to replace the ordinary convolution within residual networks in the C2f module of the backbone network. This augmentation not only enhances the model’s feature extraction capabilities but also significantly reduces both the number of model parameters and computational workload through the parallel implementation of ODConv, coupled with the simultaneous introduction of the multidimensional attention mechanism. Furthermore, InnerIoU is employed for computing Intersection over Union (IoU) loss using auxiliary edges, and MPDIoU is integrated to expedite convergence speed and enhance accuracy. The confluence of these methodologies, enriched by the incorporation of the minimum point distance in MPDIoU, collectively contributes to superior detection performance. The proposed algorithm is systematically compared and evaluated on the extensively utilized VisDrone2019 dataset. The results demonstrate that OB-YOLO surpasses the YOLOv8 baseline model by 4.8% on the VisDrone2019-DET dataset, showcasing improved performance while concurrently reducing both network parameter and floating-point calculations.
Rui Liu 0008, Zhecong Xing
IJCNN1
2024 Pre-training Language Model for Mongolian with Agglutinative Linguistic Knowledge Injection
abstract
BERT based Pre-training Language Model (PLM) has become a crucial step in achieving the best results in various natural language processing (NLP) tasks. However, the current progress, which mainly focuses on major languages such as English and Chinese, has not thoroughly investigated the low-resource languages, particularly agglutinative languages like Mongolian, due to the scarcity of large-scale data resources and the difficulty of understanding agglutinative knowledge. In this paper, we propose a novel PLM for the Mongolian language, that incorporates a novel three-stage agglutinative knowledge injection strategy. Specifically, early-stage injection aims to convert the Mongolia word sequence to the fine-grained sub-word token that comprises a stem and some suffixes; Middle-stage injection designed a morphological knowledge-based masking strategy to enhance the model's ability to learn agglutinative knowledge; Late-stage injection not only involves the model restoring the masked tokens but also predicting the order of suffixes. To address the issue of data scarcity, we create a large-scale Mongolian PLM dataset and three datasets for three downstream tasks, that are News Classification, Name Entity Recognition (NER), and Part-of-Speech (POS) prediction, etc. The experimental results on three downstream tasks demonstrate that our method surpasses the traditional BERT approach and successfully learns agglutinative language knowledge in Mongolian.
Muhan Na, Rui Liu 0008, Feilong Bao, Guanglai Gao
IJCNN2
2024 FluentEditor: Text-based Speech Editing by Considering Acoustic and Prosody Consistency
Rui Liu 0008, Jiatian Xi, Ziyue Jiang 0001, Haizhou Li 0001
INTERSPEECH1
2024 Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge
Rui Liu 0008, Zening Ma
INTERSPEECH1
2024 Generative Expressive Conversational Speech Synthesis
abstract
Conversational Speech Synthesis (CSS) aims to express a target utterance with the proper speaking style in a user-agent conversation setting. Existing CSS methods employ effective multi-modal context modeling techniques to achieve empathy understanding and expression. However, they often need to design complex network architectures and meticulously optimize the modules within them. In addition, due to the limitations of small-scale datasets containing scripted recording styles, they often fail to simulate real natural conversational styles. To address the above issues, we propose a novel generative expressive CSS system, termed GPT-Talker.We transform the multimodal information of the multi-turn dialogue history into discrete token sequences and seamlessly integrate them to form a comprehensive user-agent dialogue context. Leveraging the power of GPT, we predict the token sequence, that includes both semantic and style knowledge, of response for the agent. After that, the expressive conversational speech is synthesized by the conversation-enriched VITS to deliver feedback to the user.Furthermore, we propose a large-scale Natural CSS Dataset called NCSSD, that includes both naturally recorded conversational speech in improvised styles and dialogues extracted from TV shows. It encompasses both Chinese and English languages, with a total duration of 236 hours. We conducted comprehensive experiments on the reliability of the NCSSD and the effectiveness of our GPT-Talker. Both subjective and objective evaluations demonstrate that our model outperforms other state-of-the-art CSS systems significantly in terms of naturalness and expressiveness. The Code, Dataset, and Pre-trained Model are available at: https://github.com/AI-S2-Lab/GPT-Talker.
Rui Liu 0008, Yifan Hu 0004, Yi Ren 0006, Xiang Yin 0006, Haizhou Li 0001
ACM Multimedia1
2024 Contrastive Learning Based Modality-Invariant Feature Acquisition for Robust Multimodal Emotion Recognition With Missing Modalities
abstract
Multimodal emotion recognition (MER) aims to understand the way that humans express their emotions by exploring complementary information across modalities. However, it is hard to guarantee that full-modality data is always available in real-world scenarios. To deal with missing modalities, researchers focused on meaningful joint multimodal representation learning during cross-modal missing modality imagination. However, the cross-modal imagination mechanism is highly susceptible to errors due to the “modality gap” issue, which affects the imagination accuracy, thus, the final recognition performance. To this end, we introduce the concept of a modality-invariant feature into the missing modality imagination network, which contains two key modules: 1) a novel contrastive learning-based module to extract modality-invariant features under full modalities; 2) a robust imagination module based on imagined invariant features to reconstruct missing information under missing conditions. Finally, we incorporate imagined and available modalities for emotion recognition. Experimental results on benchmark datasets demonstrate that our proposed method outperforms existing state-of-the-art strategies. Compared with our previous work, our extended version is more effective on multimodal emotion recognition with missing modalities. The code is released athttps://github.com/ZhuoYulang/CIF-MMIN.
Rui Liu 0008, Haolin Zuo, Zheng Lian 0004, Björn W. Schuller, Haizhou Li 0001
IEEE Trans. Affect. Comput.1
2024 Text-to-Speech for Low-Resource Agglutinative Language With Morphology-Aware Language Model Pre-Training
abstract
Text-to-Speech (TTS) aims to convert the input text to a human-like voice. With the development of deep learning, encoder-decoder based TTS models perform superior performance, in terms of naturalness, in mainstream languages such as Chinese, English, etc. Note that the linguistic information learning capability of the text encoder is the key. However, for TTS of low-resource agglutinative languages, the scale of the$< $text, speech$>$paired data is limited. Therefore, how to extract rich linguistic information from small-scale text data to enhance the naturalness of the synthesized speech, is an urgent issue that needs to be addressed. In this paper, we first collect a large unsupervised text data for BERT-like language model pre-training, and then adopt the trained language model to extract deep linguistic information for the input text of the TTS model to improve the naturalness of the final synthesized speech. It should be emphasized that in order to fully exploit the prosody-related linguistic information in agglutinative languages, we incorporated morphological information into the language model training and constructed a morphology-aware masking based BERT model (MAM-BERT). Experimental results based on various advanced TTS models validate the effectiveness of our approach. Further comparison of the various data scales also validates the effectiveness of our approach in low-resource scenarios.
Rui Liu 0008, Yifan Hu 0004, Haolin Zuo, Zhaojie Luo, Longbiao Wang, Guanglai Gao
IEEE ACM Trans. Audio Speech Lang. Process.1
2024 Controllable Accented Text-to-Speech Synthesis With Fine and Coarse-Grained Intensity Rendering
abstract
Accented text-to-speech (TTS) synthesis seeks to generate speech with an accent (L2) as a variant of the standard version (L1), which is challenging as L2 is different from L1 in terms of phonetic rendering and prosody pattern (pitch, energy, and duration variance, etc.). Accented TTS has several significant real-world applications, such as language learning, preserving and documenting endangered languages and dialects, etc. that make it an important area of research and development. Moreover, changing the accent intensity of any conversational AI system has the potential to allow specific users to understand its produced speech better. However, there is no intuitive solution for the control of the accent intensity for an utterance at both fine and coarse-grained levels, that are phoneme and utterance levels respectively. In this work, we propose a neural TTS architecture that allows us to control the accent style and its intensity. This is achieved through two novel mechanisms: 1) the front-end and back-end accent knowledge injection mechanism to enhance the accent interpretability of TTS modeling; and 2) an automatic speech recognition (ASR) based accent intensity modeling strategy to quantify the accent intensity in both L2 phoneme and utterance levels. In the front-end, a newaccent variation adaptorseeks to project the accent-aware pitch, energy and duration features at a phoneme level, with the help of the fine-grained accent intensity information; In the back-end, a consistency constraint module that ensures the synthesized L2 speech manifests the expected accent intensity, is injected in the front-end, precisely. Experiments show that the proposed system attains superior performance to the baseline models in terms of accent rendering and intensity control. To our knowledge, this is the first study of accented TTS with explicit intensity control at both fine and coarse-grained levels.
Rui Liu 0008, Berrak Sisman, Guanglai Gao, Haizhou Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Exploiting Modality-Invariant Feature for Robust Multimodal Emotion Recognition with Missing Modalities
abstract
Multimodal emotion recognition leverages complementary information across modalities to gain performance. However, we cannot guarantee that the data of all modalities are always present in practice. In the studies to predict the missing data across modalities, the inherent difference between heterogeneous modalities, namely the modality gap, presents a challenge. To address this, we propose to use invariant features for a missing modality imagination network (IF-MMIN) which includes two novel mechanisms: 1) an invariant feature learning strategy that is based on the central moment discrepancy (CMD) distance under the full-modality scenario; 2) an invariant feature based imagination module (IF-IM) to alleviate the modality gap during the missing modalities prediction, thus improving the robustness of multimodal joint representation. Comprehensive experiments on the benchmark dataset IEMOCAP demonstrate that the proposed model outperforms all baselines and invariantly improves the overall emotion recognition performance under uncertain missing-modality conditions. We release the code at: https://github.com/ZhuoYulang/IF-MMIN.
Haolin Zuo, Rui Liu 0008, Jinming Zhao, Guanglai Gao, Haizhou Li 0001
ICASSP2
2023 Betray Oneself: A Novel Audio DeepFake Detection Model via Mono-to-Stereo Conversion
Rui Liu 0008, Guanglai Gao, Haizhou Li 0001
INTERSPEECH1
2023 Explicit Intensity Control for Accented Text-to-speech
Rui Liu 0008, Haolin Zuo, De Hu, Guanglai Gao, Haizhou Li 0001
INTERSPEECH1
2023 Distributed Sensor Selection for Speech Enhancement With Acoustic Sensor Networks
abstract
In distributed acoustic sensor networks, only a few nodes make a significant contribution to speech enhancement tasks. Using these most informative nodes instead of the entire network not only avoids unnecessary energy consumption but also prolongs the lifetime of sensors. To this end, a sensor selection method for distributed speech enhancement is proposed. The best subset of microphone nodes is determined by maximizing the signal-to-noise ratio (SNR), while keeping the activated nodes connected with each other. The above criterion involves an integer and non-linear programming, which is linearized with multiple base-3 sub-optimization problems, and each of them is solved by a state-of-the-art steepest descent (SD) algorithm. In addition, a greedy searching strategy is presented to select sensors rapidly. Finally, a distributed SD algorithm is further derived, which is more suitable for distributed sensor networks. The proposed method can obtain the optimal subnetwork in noisy and reverberant environments. Unlike the existing approaches, it can select nodes from a microphone network with arbitrary communication graphs. Moreover, it requires only local communications among nodes without an external central processor. Experimental results confirm the validity of the proposed method.
De Hu, Qintuya Si, Rui Liu 0008, Feilong Bao
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Decoupling Speaker-Independent Emotions for Voice Conversion via Source-Filter Networks
abstract
Emotional voice conversion (VC) aims to convert a neutral voice to an emotional one while retaining the linguistic information and speaker identity. We note that the decoupling of emotional features from other speech information (such as content, speaker identity, etc.) is the key to achieving promising performance. Some recent attempts of speech representation decoupling on the neutral speech cannot work well on the emotional speech, due to the more complex entanglement of acoustic properties in the latter. To address this problem, here we propose a novel Source-Filter-based Emotional VC model (SFEVC) to achieve proper filtering of speaker-independent emotion cues from both the timbre and pitch features. Our SFEVC model consists of multi-channel encoders, emotion separate encoders, pre-trained speaker-dependent encoders, and the corresponding decoder. Note that all encoder modules adopt a designed information bottleneck auto-encoder. Additionally, to further improve the conversion quality for various emotions, a novel training strategy based on the 2D Valence-Arousal (VA) space is proposed. Experimental results show that the proposed SFEVC along with a VA training strategy outperforms all baselines and achieves the state-of-the-art performance in speaker-independent emotional VC with nonparallel data.
Zhaojie Luo, Shoufeng Lin, Rui Liu 0008, Jun Baba, Yuichiro Yoshikawa, Hiroshi Ishiguro
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Visualtts: TTS with Accurate Lip-Speech Synchronization for Automatic Voice Over
abstract
In this paper, we formulate a novel task to synthesize speech in sync with a silent pre-recorded video, denoted as automatic voice over (AVO). Unlike traditional speech synthesis, AVO seeks to generate not only human-sounding speech, but also perfect lip-speech synchronization. A natural solution to AVO is to condition the speech rendering on the temporal progression of lip sequence in the video. We propose a novel text-to-speech model that is conditioned on visual input, named VisualTTS, for accurate lip-speech synchronization. The proposed VisualTTS adopts two novel mechanisms that are 1) textual-visual attention, and 2) visual fusion strategy during acoustic decoding, which both contribute to forming accurate alignment between the input text content and lip motion in input lip sequence. Experimental results show that VisualTTS achieves accurate lip-speech synchronization and outperforms all baseline systems.
Junchen Lu, Berrak Sisman, Rui Liu 0008, Mingyang Zhang 0003, Haizhou Li 0001
ICASSP3
2022 Alignment-Learning Based Single-Step Decoding for Accurate and Fast Non-Autoregressive Speech Recognition
abstract
Non-autoregressive transformer (NAT) based speech recognition models have gained more and more attention since they perform faster inference speed compared with autoregressive counterparts, especially when the single-step decoding is applied. However, the single-step decoding process with length prediction will suffer from the decoding stability problem and limited improvement for inference speed. To address this, in this paper, we propose an alignment learning based NAT model, named AL-NAT. Our idea is inspired by the fact that the encoder CTC output and the target sequence are monotonically related. Specifically, we design an alignment cost matrix between the CTC output tokens and the target tokens and define a novel alignment loss to minimize the distance between the alignment cost matrix and the ground truth monotonic alignment path. By eliminating the length prediction mechanism, our AL-NAT model achieves remarkable improvements in recognition accuracy and decoding speed. To learn the contextual knowledge to improve the decoding accuracy, we further add lightweight language model on both the encoder and decoder side. Our proposed method achieves WERs of 2.8%/6.3% and RTF of 0.011 on Librispeech test clean/other sets with a lightweight 3-gram LM, and a CER of 5.3% and RTF of 0.005 on Aishell1 without LM, respectively.
Yonghe Wang, Rui Liu 0008, Feilong Bao, Hui Zhang 0031, Guanglai Gao
ICASSP2
2022 A Deep Investigation of RNN and Self-attention for the Cyrillic-Traditional Mongolian Bidirectional Conversion
Muhan Na, Rui Liu 0008, Feilong Bao, Guanglai Gao
ICONIP (6)2
2022 Accurate Emotion Strength Assessment for Seen and Unseen Speech Based on Data-Driven Deep Learning
abstract
Emotion classification of speech and assessment of the emotion strength are required in applications such as emotional text-to-speech and voice conversion. The emotion attribute ranking function based on Support Vector Machine (SVM) was proposed to predict emotion strength for emotional speech corpus. However, the trained ranking function doesn't generalize to new domains, which limits the scope of applications, especially for out-of-domain or unseen speech. In this paper, we propose a data-driven deep learning model, i.e. StrengthNet, to improve the generalization of emotion strength assessment for seen and unseen speech. This is achieved by the fusion of emotional data from various domains. We follow a multi-task learning network architecture that includes an acoustic encoder, a strength predictor, and an auxiliary emotion predictor. Experiments show that the predicted emotion strength of the proposed StrengthNet is highly correlated with ground truth scores for both seen and unseen speech. We release the source codes at: https://github.com/ttslr/StrengthNet.
Rui Liu 0008, Berrak Sisman, Björn W. Schuller, Guanglai Gao, Haizhou Li 0001
INTERSPEECH1
2022 Multistage Deep Transfer Learning for EmIoT-Enabled Human-Computer Interaction
abstract
Emotional Internet of Things (EmIoT), which provides Internet of Things (IoT) devices cognitive and socialization capabilities, has been regarded as a future direction to improve users’ experiences. With the development of intelligent techniques, the requirement of EmIoT is not only sensing the users’ emotional states but also providing emotional feedbacks. Human–computer interaction has been studied to achieve speech interaction with IoT devices. The recent advances in neural text-to-speech (TTS) have made “human parity” synthesized speech possible for IoT-enabled human–computer interaction. Furthermore, emotion control can be achieved by using the emotional codes in a unified model, referred to as emotional TTS (or ETTS for short). Such ETTS models have achieved promising emotional expressiveness using large-scale emotion-annotated English data set; however, they are not practical in IoT environments with other mainstream languages, especially for Chinese. In fact, the limited available large-scale emotion-annotated data set is challenging the development of Chinese ETTS. To address that we propose a multistage deep transfer learning scheme to design a high-quality Chinese ETTS system under a small-scale training corpus to achieve EmIoT in Mandarin environments. In this scheme, the pretrained knowledge from the former stages corresponding to a large-scale neutral English and a medium-scale emotional English corpora is transferred to a Mandarin ETTS model. Thereby, the trained model can achieve high-quality emotional speech with limited available emotional corpus, which is able to serve various EmIoT-oriented applications. The experiments have been conducted to demonstrate the effectiveness and superiority of the proposed model as compared to other counterparts in terms of naturalness and emotional expressiveness. We refer readers to visit our demo Webpage1enjoy the synthesized speech samples.
Rui Liu 0008, Qi Liu 0005, Hongxu Zhu, Hui Cao 0004
IEEE Internet Things J.1
2022 Emotional voice conversion: Theory, databases and ESD
abstract
In this paper, we first provide a review of the state-of-the-art emotional voice conversion research, and the existing emotional speech databases. We then motivate the development of a novel emotional speech database (ESD) that addresses the increasing research need. With this paper, the ESD database1 is now made available to the research community. The ESD database consists of 350 parallel utterances spoken by 10 native English and 10 native Chinese speakers and covers 5 emotion categories (neutral, happy, angry, sad and surprise). More than 29 h of speech data were recorded in a controlled acoustic environment. The database is suitable for multi-speaker and cross-lingual emotional voice conversion studies. As case studies, we implement several state-of-the-art emotional voice conversion systems on the ESD database. This paper provides a reference study on ESD in conjunction with its release.
Kun Zhou 0003, Berrak Sisman, Rui Liu 0008, Haizhou Li 0001
Speech Commun.3
2022 Decoding Knowledge Transfer for Neural Text-to-Speech Training
abstract
Neural end-to-end text-to-speech (TTS) is superior to conventional statistical methods in many ways. However, the exposure bias problem, that arises from the mismatch between the training and inference process in autoregressive models, remains an issue. It often leads to performance degradation in face of out-of-domain test data. To address this problem, we study a novel decoding knowledge transfer strategy, and propose a multi-teacher knowledge distillation (MT-KD) network for Tacotron2 TTS model. The idea is to pre-train two Tacotron2 TTS teacher models in teacher forcing and scheduled sampling modes, and transfer the pre-trained knowledge to a student model that performs free running decoding. We show that the MT-KD network provides an adequate platform for neural TTS training, where the student model learns to emulate the behaviors of the two teachers, at the same time, minimizing the mismatch between training and run-time inference. Experiments on both Chinese and English data show that MT-KD system consistently outperforms the competitive baselines in terms of naturalness, robustness and expressiveness for in-domain and out-of-domain test data. Furthermore, we show that knowledge distillation outperforms adversarial learning and data augmentation in addressing the exposure bias problem.
Rui Liu 0008, Berrak Sisman, Guanglai Gao, Haizhou Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Graphspeech: Syntax-Aware Graph Attention Network for Neural Speech Synthesis
abstract
Attention-based end-to-end text-to-speech synthesis (TTS) is superior to conventional statistical methods in many ways. Transformer-based TTS is one of such successful implementations. While Transformer TTS models the speech frame sequence well with a self-attention mechanism, it does not associate input text with output utterances from a syntactic point of view at sentence level. We propose a novel neural TTS model, denoted as GraphSpeech, that is formulated under graph neural network framework. GraphSpeech encodes explicitly the syntactic relation of input lexical tokens in a sentence, and incorporates such information to derive syntactically motivated character embeddings for TTS attention mechanism. Experiments show that GraphSpeech consistently outperforms the Transformer TTS baseline in terms of spectrum and prosody rendering of utterances.
Rui Liu 0008, Berrak Sisman, Haizhou Li 0001
ICASSP1
2021 Seen and Unseen Emotional Style Transfer for Voice Conversion with A New Emotional Speech Dataset
abstract
Emotional voice conversion aims to transform emotional prosody in speech while preserving the linguistic content and speaker identity. Prior studies show that it is possible to disentangle emotional prosody using an encoder-decoder network conditioned on discrete representation, such as one-hot emotion labels. Such networks learn to remember a fixed set of emotional styles. In this paper, we propose a novel framework based on variational auto-encoding Wasserstein generative adversarial network (VAW-GAN), which makes use of a pre-trained speech emotion recognition (SER) model to transfer emotional style during training and at run-time inference. In this way, the network is able to transfer both seen and unseen emotional style to a new utterance. We show that the proposed framework achieves remarkable performance by consistently outperforming the baseline framework. This paper also marks the release of an emotional speech dataset (ESD) for voice conversion, which has multiple speakers and languages.
Kun Zhou 0003, Berrak Sisman, Rui Liu 0008, Haizhou Li 0001
ICASSP3
2021 Reinforcement Learning for Emotional Text-to-Speech Synthesis with Improved Emotion Discriminability
abstract
Emotional text-to-speech synthesis (ETTS) has seen much progress in recent years.However, the generated voice is often not perceptually identifiable by its intended emotion category.To address this problem, we propose a new interactive training paradigm for ETTS, denoted as i-ETTS, which seeks to directly improve the emotion discriminability by interacting with a speech emotion recognition (SER) model.Moreover, we formulate an iterative training strategy with reinforcement learning to ensure the quality of i-ETTS optimization.Experimental results demonstrate that the proposed i-ETTS outperforms the state-of-the-art baselines by rendering speech with more accurate emotion style.To our best knowledge, this is the first study of reinforcement learning in emotional text-to-speech synthesis.
Rui Liu 0008, Berrak Sisman, Haizhou Li 0001
Interspeech1
2021 FastTalker: A neural text-to-speech architecture with shallow and group autoregression
Rui Liu 0008, Berrak Sisman, Yixing Lin, Haizhou Li 0001
Neural Networks1
2021 Exploiting Morphological and Phonological Features to Improve Prosodic Phrasing for Mongolian Speech Synthesis
abstract
Prosodic phrasing is an important factor that affects naturalness and intelligibility in text-to-speech synthesis. Studies show that deep learning techniques improve prosodic phrasing when large text and speech corpus are available. However, for low-resource languages, such as Mongolian, prosodic phrasing remains a challenge for various reasons. First, the database suitable for system training is limited. Second, word composition knowledge that is prosody-informing has not been used in prosodic phrase modeling. To address these problems, in this article, we propose a feature augmentation method in conjunction with a self-attention neural classifier. We augment input text with morphological and phonological decompositions of words to enhance the text encoder. We study the use of self-attention classifier, that makes use of global context of a sentence, as a decoder for phrase break prediction. Both objective and subjective evaluations validate the effectiveness of the proposed phrase break prediction framework, that consistently improves voice quality in a Mongolian text-to-speech synthesis system.
Rui Liu 0008, Berrak Sisman, Feilong Bao, Guanglai Gao, Haizhou Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Expressive TTS Training With Frame and Style Reconstruction Loss
abstract
We propose a novel training strategy for Tacotron-based text-to-speech (TTS) system that improves the speech styling at utterance level. One of the key challenges in prosody modeling is the lack of reference that makes explicit modeling difficult. The proposed technique doesn’t require prosody annotations from training data. It doesn’t attempt to model prosody explicitly either, but rather encodes the association between input text and its prosody styles using a Tacotron-based TTS framework. This study marks a departure from the style token paradigm where prosody is explicitly modeled by a bank of prosody embeddings. It adopts a combination of two objective functions: 1) frame level reconstruction loss, that is calculated between the synthesized and target spectral features; 2) utterance level style reconstruction loss, that is calculated between the deep style features of synthesized and target speech. The style reconstruction loss is formulated as a perceptual loss to ensure that utterance level speech style is taken into consideration during training. Experiments show that the proposed training strategy achieves remarkable performance and outperforms the state-of-the-art baseline in both naturalness and expressiveness. To our best knowledge, this is the first study to incorporate utterance level perceptual quality as a loss function into Tacotron training for improved expressiveness.
Rui Liu 0008, Berrak Sisman, Guanglai Gao, Haizhou Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2020 Teacher-Student Training For Robust Tacotron-Based TTS
abstract
While neural end-to-end text-to-speech (TTS) is superior to conventional statistical methods in many ways, the exposure bias problem in the autoregressive models remains an issue to be resolved. The exposure bias problem arises from the mismatch between the training and inference process, that results in unpredictable performance for out-of-domain test data at run-time. To overcome this, we propose a teacher-student training scheme for Tacotron-based TTS by introducing a distillation loss function in addition to the feature loss function. We first train a Tacotron2-based TTS model by always providing natural speech frames to the decoder, that serves as a teacher model. We then train another Tacotron2-based model as a student model, of which the decoder takes the predicted speech frames as input, similar to how the decoder works during run-time inference. With the distillation loss, the student model learns the output probabilities from the teacher model, that is called knowledge distillation. Experiments show that our proposed training scheme consistently improves the voice quality for out-of-domain test data both in Chinese and English systems.
Rui Liu 0008, Berrak Sisman, Jingdong Li, Feilong Bao, Guanglai Gao, Haizhou Li 0001
ICASSP1
2020 Modeling Prosodic Phrasing With Multi-Task Learning in Tacotron-Based TTS
abstract
Tacotron-based end-to-end speech synthesis has shown remarkable voice quality. However, the rendering of prosody in the synthesized speech remains to be improved, especially for long sentences, where prosodic phrasing errors can occur frequently. In this letter, we extend the Tacotron-based speech synthesis framework to explicitly model the prosodic phrase breaks. We propose a multi-task learning scheme for Tacotron training, that optimizes the system to predict both Mel spectrum and phrase breaks. To our best knowledge, this is the first implementation of multi-task learning for Tacotron based TTS with a prosodic phrasing model. Experiments show that our proposed training scheme consistently improves the voice quality for both Chinese and Mongolian systems.
Rui Liu 0008, Berrak Sisman, Feilong Bao, Guanglai Gao, Haizhou Li 0001
IEEE Signal Process. Lett.1
2019 Building Mongolian TTS Front-End with Encoder-Decoder Model by Using Bridge Method and Multi-view Features
Rui Liu 0008, Feilong Bao, Guanglai Gao
ICONIP (5)1
2018 A LSTM Approach with Sub-Word Embeddings for Mongolian Phrase Break Prediction
abstract
In this paper, we first utilize the word embedding that focuses on sub-word units to the Mongolian Phrase Break (PB) prediction task by using Long-Short-Term-Memory (LSTM) model. Mongolian is an agglutinative language. Each root can be followed by several suffixes to form probably millions of words, but the existing Mongolian corpus is not enough to build a robust entire word embedding, thus it suffers a serious data sparse problem and brings a great difficulty for Mongolian PB prediction. To solve this problem, we look at sub-word units in Mongolian word, and encode their information to a meaningful representation, then fed it to LSTM to decode the best corresponding PB label. Experimental results show that the proposed model significantly outperforms traditional CRF model using manually features and obtains 7.49% F-Measure gain.
Rui Liu 0008, Feilong Bao, Guanglai Gao, Hui Zhang 0031, Yonghe Wang
COLING1
2018 Improving Mongolian Phrase Break Prediction by Using Syllable and Morphological Embeddings with BiLSTM Model
Rui Liu 0008, Feilong Bao, Guanglai Gao, Hui Zhang 0031, Yonghe Wang
INTERSPEECH1
2018 Phonologically Aware BiLSTM Model for Mongolian Phrase Break Prediction with Attention Mechanism
Rui Liu 0008, Feilong Bao, Guanglai Gao, Hui Zhang 0031, Yonghe Wang
PRICAI (1)1